Context-aware unmanned aerial vehicle flight shooting scheme generation method and system

CN122593366APending Publication Date: 2026-08-18BEIJING JUNYU AEROSPACE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610828107.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

现有智能拍摄技术在处理此类请求时面临巨大挑战:无法理解动感或安全在特定上下文中的具体含义;系统仅能被动响应实时移动,无法预判事件发展结构,如提前识别进球动作并调整相机位置;在复杂任务中,用户必须人工规划多个航点以实现行为序列(例如先大范围环绕取景、再低空跟随主体、最后悬停特写),导致操作繁琐且难以适应环境突变

Benefits of technology

通过大语言模型对模糊的自然语言用户请求进行深度的意图解析与任务分解,结合以摄影学为基础的知识图谱检索拍摄规则,实现了从语义约束到飞行参数的自动化转换,解决了现有技术无法理解情境化用户需求的问题;通过生成—验证—修正的闭环机制,在飞行方案执行前对其进行运动学一致性、环境约束强制及法规遵从性的多维可行性检查,并将失败原因反馈至大语言模型进行重规划,有效保证了生成拍摄方案的安全性和可执行性;通过将拍摄任务分解为环境取景任务、氛围营造任务和多个优先级任务的层次化结构,使系统能够在无需人工设置航点的情况下自动串联多个飞行行为,支持复杂拍摄场景的全自动规划。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122593366A_ABST
    Figure CN122593366A_ABST
Patent Text Reader

Abstract

The application provides a kind of unmanned aerial vehicle flight shooting scheme generation method and system based on situation awareness, it is related to unmanned aerial vehicle technical field.Through the deep intention analysis and task decomposition of fuzzy natural language user request to large language model, combined with the knowledge graph retrieval shooting rule based on photography, the automatic conversion from semantic constraint to flight parameter is realized;Through the closed-loop mechanism of generation-verification-correction, multidimensional feasibility check of kinematics consistency, environmental constraint enforcement and regulatory compliance is carried out before the flight scheme is executed, and the failure reason is fed back to the large language model for re-planning.By decomposing the shooting task into the hierarchical structure of environment framing task, atmosphere creating task and multiple priority tasks, the system can automatically link multiple flight behaviors without manually setting the waypoint, supporting full-automatic planning of complex shooting scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and more specifically, to a method and system for generating UAV flight photography schemes based on context awareness. Background Technology

[0002] Current drone intelligent shooting technology mainly relies on preset fixed modes to operate. Users need to manually set specific parameters, such as selecting the orbit mode and entering the radius value. The system can only mechanically execute commands and lacks the ability to understand the user's true intentions.

[0003] In real-world applications, user needs are often highly contextualized and ambiguous, such as: recording one's daughter playing soccer; maintaining safety while capturing a dynamic moment, and capturing the instant she scores. Existing intelligent shooting technologies face significant challenges in handling such requests: they cannot understand the specific meanings of dynamism or safety within a given context; the system can only passively respond to real-time movement and cannot predict the structure of events, such as recognizing a goal in advance and adjusting the camera position; in complex tasks, users must manually plan multiple waypoints to achieve a sequence of actions (e.g., first framing a wide-angle shot, then following the subject at low altitude, and finally hovering for a close-up), resulting in cumbersome operation and difficulty adapting to sudden environmental changes.

[0004] Due to a lack of comprehensive understanding of contextual factors, existing systems are unable to automatically convert ambiguous natural language commands into precise flight control parameters, which severely restricts the intelligent shooting performance of drones in dynamic scenarios. Summary of the Invention

[0005] The problem this invention aims to solve is how to translate ambiguous natural language into precise flight control.

[0006] To address the aforementioned problems, in a first aspect, the present invention provides a method for generating drone flight photography schemes based on context awareness, comprising: A large language model is used to parse the intent of user requests, decompose the task, extract the subject, atmosphere and scene, and transform the extracted content into semantic constraints. A search is performed in the knowledge graph to return shooting rules that match the semantic constraints. The shooting rules include shooting mode and shooting parameters. In the knowledge graph, atmosphere and subject correspond to shooting rules, respectively. Based on the retrieved shooting rules and semantic constraints, a large language model is used to generate a shooting plan for the drone. Based on preset rules, a feasibility check is performed on the shooting plan. If the shooting plan fails the overall check, the shooting plan is rejected, a reason for failure is generated, a prompt is given to the large language model, and the shooting plan is regenerated until the overall shooting plan passes the check. The approved shooting plan is converted into flight control commands and transmitted to the drone flight controller for execution.

[0007] Optionally, the step of using a large language model to parse the user request's intent, decompose the task, extract the subject, atmosphere, and scene, and transform the extracted content into semantic constraints, including: The user request is parsed using a large language model, which first extracts the subject, atmosphere, scene, idea, and implicit intent; Generate the main task based on the concept image, theme, atmosphere, and scene; Based on the main task and implicit intent, phased tasks are generated, including environmental framing tasks, atmosphere creation tasks, and multiple priority tasks. Break down the phased tasks into sub-tasks.

[0008] Optionally, the knowledge graph organizes knowledge entities based on photography, and the stored mapping and matching rules include: the correspondence between atmosphere tags and flight speed, flight altitude, motion trajectory type and camera parameters; and the correspondence between scene type and safety distance requirements, flight noise constraints, recommended shooting angle and aerial photography motion mode.

[0009] Optionally, the step of searching in the knowledge graph and returning shooting rules that match the semantic constraints includes shooting modes and shooting parameters, including: For environmental framing tasks, the knowledge graph is searched based on the scene to obtain shooting rules. The shooting mode adopts the aerial photography motion mode corresponding to the scene, and the shooting parameters include the safety distance requirements, flight noise constraints and recommended shooting angles corresponding to the scene. The environmental framing task obtains the entire environmental layout and constructs a 3D environmental model. For the atmosphere creation task, the shooting rules are obtained by searching the knowledge graph based on the atmosphere. The shooting parameters include the flight speed, flight altitude and camera parameters corresponding to the atmosphere. The atmosphere creation task runs through the entire shooting process. For multiple priority tasks, identify the subject in each priority task, retrieve the motion trajectory type corresponding to the atmosphere from the knowledge graph based on the atmosphere, and set the object of motion trajectory tracking as the subject.

[0010] Optionally, the preset rules include kinematic consistency, environmental constraints, and regulatory compliance. The kinematic consistency includes the maximum thrust of the UAV, battery life, wind resistance, and maximum tilt angle. The environmental constraints include that the path cannot intersect with obstacles and the minimum distance between the UAV and people and moving objects. The regulatory compliance includes the permitted flight range and altitude limits.

[0011] Optionally, the feasibility check of the shooting plan based on preset rules includes: Based on the 3D environmental model, the range of personnel movement, and the maximum height of personnel, a safe flight zone is formed in the 3D environmental model; Based on the preset action path of the main body, the motion trajectory type and tracking relationship retrieved from the priority tasks, the corresponding drone operation trajectory is generated; Analyze whether the drone's flight path is entirely within the safe flight zone. If yes, the test is considered passed; otherwise, the test is considered failed. Further analysis of the drone's flight trajectory and planned flight time is conducted to determine whether the drone's motion parameters meet the requirements of kinematic consistency and regulatory compliance. If yes, the test is deemed passed; otherwise, the test is deemed failed. If all preset rules are deemed acceptable, the shooting plan is considered acceptable overall; otherwise, the shooting plan is considered unacceptable overall.

[0012] Optionally, the reasons for failure include: the drone's flight path intersecting with the safe flight area, flight parameters exceeding limits, and flight altitude or flight range exceeding limits.

[0013] Optionally, providing prompts to the large language model and regenerating the shooting plan includes: If the failure is due to the drone's trajectory intersecting with the safe flight area, the path segment outside the safe flight area will be removed, and the severed trajectory will be reconnected within the safe flight area; if reconnection is not possible, the filming plan will be rejected.

[0014] Secondly, the present invention also provides a context-aware UAV flight photography scheme generation system, comprising: The semantic reasoning module is used to parse the intent of user requests using a large language model, decompose the task, extract the subject, atmosphere and scene, and transform the extracted content into semantic constraints. The knowledge graph retrieval module is used to perform retrieval in the knowledge graph and return shooting rules that match the semantic constraints. The shooting rules include shooting mode and shooting parameters. In the knowledge graph, atmosphere and subject correspond to the shooting rules respectively. The flight plan generation module is used to generate drone shooting plans based on retrieved shooting rules and semantic constraints using a large language model. The feasibility check module is used to check the feasibility of the shooting plan based on preset rules. If the shooting plan fails the overall check, the shooting plan is rejected, the reason for failure is generated, the large language model is given a prompt, and the shooting plan is regenerated until the shooting plan passes the overall check. The scheme conversion module is used to convert the approved shooting scheme into flight control commands and transmit them to the UAV flight controller for execution.

[0015] This invention provides a method and system for generating drone flight photography schemes based on context awareness. Compared with existing technologies, it has the following advantages: By employing a large language model to perform deep intent parsing and task decomposition of fuzzy natural language user requests, and combining this with knowledge graph retrieval of shooting rules based on photography, the system achieves automated conversion from semantic constraints to flight parameters, solving the problem that existing technologies cannot understand contextualized user needs. Through a closed-loop mechanism of generation, verification, and correction, the system performs multi-dimensional feasibility checks on the flight plan before execution, including kinematic consistency, environmental constraint enforcement, and regulatory compliance. Failure reasons are fed back to the large language model for replanning, effectively ensuring the safety and executability of the generated shooting plan. By decomposing the shooting task into a hierarchical structure of environmental framing tasks, atmosphere creation tasks, and multiple priority tasks, the system can automatically connect multiple flight behaviors without manual waypoint setting, supporting fully automated planning for complex shooting scenarios. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating a method for generating a drone flight photography scheme based on context awareness, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a context-aware UAV flight photography scheme generation system provided in an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application are described clearly and completely. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0020] like Figure 1 As shown in the figure, an embodiment of this application provides a method for generating a drone flight photography scheme based on context awareness, including: S1: Use a large language model to parse the intent of user requests, decompose the task, extract the subject, atmosphere and scene, and transform the extracted content into semantic constraints.

[0021] Specifically, user requests can be submitted directly to the large language model via a text input box. For example, the user could input: "Shoot a wedding reception, focusing on the bride, creating a romantic atmosphere." The large language model is trained to identify key information in the user's language, such as identifying elements like "who," "what," and "where" through keyword matching or simple syntactic analysis, thus initially extracting the subject, atmosphere, and scene. For instance, when a user inputs "Shoot my daughter playing in the park, with a warm and cozy feel," the large language model can identify "daughter" as the subject, "park" as the scene, and "warm and cozy" as the atmosphere. This extracted information is then formatted into structured semantic constraints.

[0022] S2: Search the knowledge graph and return the shooting rules that match the semantic constraints. The shooting rules include shooting mode and shooting parameters. In the knowledge graph, atmosphere and subject correspond to the shooting rules respectively.

[0023] Specifically, the knowledge graph can be constructed as a simple database storing predefined shooting rules. For example, a "warm" atmosphere can be directly mapped to shooting parameters such as "slow flight," "low altitude," and "soft light," while a "person" subject can be mapped to "follow mode." When semantic constraints are received, the system directly queries this database to find entries that perfectly match the extracted subject and atmosphere tags, thereby obtaining the corresponding shooting mode and shooting parameters.

[0024] S3: Based on the retrieved shooting rules and semantic constraints, generate a drone shooting plan using a large language model.

[0025] Specifically, the large language model can receive retrieved shooting modes and parameters, along with the original semantic constraints, and then attempt to combine this information into a preliminary shooting plan. For example, the large language model can generate a text description of how the drone follows the subject's movement based on parameters such as "follow mode" and "slow flight," and specify how the camera should be adjusted. This plan can be a high-level description, such as "the drone follows the subject slowly and maintains low-altitude flight."

[0026] S4: Based on preset rules, perform a feasibility check on the shooting plan. If the overall shooting plan fails the check, the shooting plan is rejected, the reason for failure is generated, a prompt is given to the large language model, and the shooting plan is regenerated until the overall shooting plan passes the check.

[0027] Specifically, once a shooting plan is generated, the system compares each parameter and path in the plan with these preset rules. If any non-compliance is found, such as the specified flight altitude exceeding the maximum allowable altitude, the plan is deemed infeasible. In this case, the system generates a simple reason for failure, such as "flight altitude exceeded," and feeds it as a text prompt to the large language model, prompting it to adjust the plan.

[0028] S5: Convert the approved shooting plan into flight control commands and transmit them to the UAV flight controller for execution.

[0029] Specifically, the high-level description of the shooting plan is translated into low-level instructions that the drone flight controller can directly understand and execute. For example, if the shooting plan is described as "the drone follows the subject at a speed of 2 meters per second and maintains an altitude of 5 meters," this description will be translated into a series of specific waypoint coordinates, velocity vectors, and camera gimbal angle instructions. These instructions are then sent to the drone flight controller via a wireless communication link, which drives the drone to fly and shoot according to the instructions.

[0030] In this optional embodiment, a large language model is used to perform deep intent parsing and task decomposition on fuzzy natural language user requests. Combined with knowledge graph retrieval of shooting rules based on photography, automated conversion from semantic constraints to flight parameters is achieved, solving the problem that existing technologies cannot understand contextualized user needs. Through a closed-loop mechanism of "generation-verification-correction," a multi-dimensional feasibility check is performed on the flight plan before execution, considering kinematic consistency, environmental constraint enforcement, and regulatory compliance. Failure reasons are fed back to the large language model for replanning, effectively ensuring the safety and executability of the generated shooting plan. By decomposing the shooting task into a hierarchical structure of environmental framing tasks, atmosphere creation tasks, and multiple priority tasks, the system can automatically connect multiple flight behaviors without manual waypoint setting, supporting fully automated planning for complex shooting scenarios. This effectively transforms fuzzy natural language requests into precise, safe, and intelligent UAV flight control commands, significantly improving the automation, intelligence, and user experience of UAV intelligent shooting.

[0031] The following is a detailed description of each step.

[0032] S1: Employ a large language model to parse the user's intent, decompose the task, extract the subject, atmosphere, and scenario, and transform the extracted content into semantic constraints. This step specifically includes the following:

[0033] The user request is parsed using a large language model, which first extracts the subject, atmosphere, scene, idea, and implicit intent.

[0034] Specifically, the large language model receives natural language requests from users (in text or voice form), performs deep parsing of the requests, and extracts the following elements: Subject: The core object the user wants to capture, such as a specific person or object, which needs to be tracked in conjunction with the visual recognition module; Atmosphere: The emotional tone the user expects for the shot, such as romantic, dynamic, or tranquil, used to drive the retrieval of flight parameters in the knowledge graph; Scene: The environmental background of the shot, such as a football field or a wedding scene, used to determine flight constraints; Main Image: The core shooting target in the user's request; Implicit Intent: Shooting expectations that the user does not express directly but can be inferred from the context, such as "If she scores, I must capture that moment," which implies a conditional trigger monitoring requirement for a specific event.

[0035] Large language models, with their powerful semantic understanding and reasoning capabilities, can analyze user-input text using natural language processing techniques to identify key information. For example, they can utilize named entity recognition (NER) to identify subjects and scenes, sentiment analysis or keyword matching to identify atmosphere, and intent recognition models to distinguish between intentional and implicit intents.

[0036] Generate the main task based on the main idea, theme, atmosphere, and scene.

[0037] Specifically, the main task is a general description of the user's overall shooting goal, integrating the user's scattered intentions into a clear and actionable shooting direction. For example, if the main idea is "shooting a wedding video," the subjects are "the bride and groom," the atmosphere is "happy and romantic," and the scene is "outdoor lawn," then the main task can be generated as "shooting a happy and romantic wedding video for the bride and groom on an outdoor lawn." This step can be generated by a large language model based on preset rules or templates, or by using the parsed elements as input to allow the large language model to directly generate the task description.

[0038] Based on the main task and implicit intent, phased tasks are generated. These phased tasks include environmental framing tasks, atmosphere creation tasks, and multiple priority tasks. Multiple priority tasks refer to shooting tasks that are set according to the importance of the subject being filmed or a specific event, and that require priority completion or special handling, based on user preferences.

[0039] Break down the phased tasks into sub-tasks.

[0040] Subtasks are the smallest executable units of a phased task, directly corresponding to specific drone flight maneuvers, camera parameter settings, or shooting mode selections. For example, an "environmental framing task" can be broken down into subtasks such as "drone high-altitude panoramic shooting" and "wide-angle lens capturing landmark buildings"; an "atmosphere creation task" can be broken down into subtasks such as "stable low-speed flight" and "adjusting camera white balance to warm tones"; a "priority task" such as "shooting close-up shots of couples interacting" can be broken down into subtasks such as "tracking the subject," "close-up shots," and "switching to slow-motion mode." This progressive decomposition process ensures a complete mapping from vague user requests to specific drone flight commands.

[0041] For example, consider the task of "filming my daughter playing soccer in a dynamic style, capturing the moment a goal is scored": the subject is the target person, the atmosphere is dynamic, the scene is a soccer field, the main image is a close-up of the person's movement, and the implicit intention is to trigger a close-up shot when a goal is scored; the environmental framing task is a panoramic aerial shot of the field, the atmosphere creation task is to continuously use high-speed, wide-angle parameters, and one of the priority tasks is to monitor the goal and trigger zoom in. This multi-layered and refined intent analysis and task decomposition mechanism can comprehensively and accurately capture all of the user's explicit and implicit shooting needs, thereby generating more accurate, detailed, and comprehensive semantic constraints.

[0042] S2: Search the knowledge graph and return the shooting rules that match the semantic constraints. The shooting rules include shooting mode and shooting parameters. In the knowledge graph, atmosphere and subject correspond to the shooting rules respectively.

[0043] Specifically, the knowledge graph used in this application organizes knowledge entities based on photography, storing mapping and matching rules between semantic contextual elements and engineering flight parameters. These mapping and matching rules include: the correspondence between atmosphere labels and flight speed, flight altitude, motion trajectory type, and camera parameters (including shutter speed, equivalent focal length, and aperture range); and the correspondence between scene types and safety distance requirements, flight noise constraints, recommended shooting angles, and aerial photography motion modes. The role of the knowledge graph is to anchor the semantic output of the large language model to the range of engineering parameters that conform to the aesthetic standards of the photographic industry, preventing the large language model from generating meaningless, illusory outputs.

[0044] Fundamental theories in photography, such as composition principles (like the rule of thirds and the golden ratio), lighting techniques (like front lighting, backlighting, and side lighting), depth-of-field control (like large depth of field and shallow depth of field), and color theory, are treated as entities in a knowledge graph, with hierarchical relationships, associations, or attributes defined between them. For example, the entity "large depth of field" can be associated with attributes or entities such as "small aperture" and "long-distance shooting." Another approach is to treat the characteristics of photographic equipment (such as sensor size, lens focal length range, and aperture range of different types of drone cameras), shooting techniques (such as time-lapse photography, surround shooting, and tracking shooting), and the photographic requirements corresponding to different shooting themes (such as portraits, landscapes, and architecture) as knowledge entities. These entities are connected through relationships such as "applies to," "needs," and "contains."

[0045] This establishes a mapping between atmosphere labels and flight speed, flight altitude, motion trajectory type, and camera parameters. The goal is to create a direct link between the user's desired shooting atmosphere and the drone's specific flight attitude and camera settings. For example, these mappings can be stored in the form of tables or databases, where an "Atmosphere Label" field (e.g., "Romantic," "Epic," "Dynamic") corresponds to multiple "Flight Speed" ranges (e.g., "Slow," "Medium," "Fast"), "Flight Altitude" ranges (e.g., "Low," "Medium," "High"), "Motion Trajectory Types" (e.g., "Transverse," "Circle," "Ascend"), and "Camera Parameters" (e.g., "Large Aperture," "Small Aperture," "High ISO").

[0046] This document outlines the correspondence between scene types and safety distance requirements, flight noise constraints, recommended shooting angles, and aerial photography modes. This rule ensures the safety and compliance of drone photography in different scenarios and provides optimal shooting recommendations specific to each scenario. For example, it can be stored in a table or database, where a "Scene Type" field (e.g., "City Park," "Mountainous Area," "Seaside," "Indoor") corresponds to "Safety Distance Requirements" (e.g., "More than 5 meters from people," "More than 3 meters from buildings"), "Flight Noise Constraints" (e.g., "Low Noise Mode," "Unrestricted"), "Recommended Shooting Angles" (e.g., "Overhead View," "Level View," "Upward View"), and "Aerial Photography Modes" (e.g., "Hovering," "Straight Flight," "Circling Flight").

[0047] Step S2 specifically includes the following:

[0048] For environmental framing tasks, the knowledge graph is used to retrieve shooting rules based on the scene. The shooting mode adopts the aerial photography motion mode corresponding to the scene, and the shooting parameters include the safety distance requirements, flight noise constraints and recommended shooting angles corresponding to the scene. The environmental framing task obtains the entire environmental layout and constructs a 3D environmental model.

[0049] Specifically, within the knowledge graph, scene information parsed from user requests is retrieved to obtain shooting rules matching the scene. For example, for a "mountainous" scene, suitable aerial photography movement modes, safety distance requirements, flight noise constraints, and recommended shooting angles can be retrieved. Another approach is for the system to automatically trigger corresponding environmental framing tasks based on preset scene types (such as city, countryside, or water), and obtain corresponding shooting strategies from the knowledge graph. The shooting mode, which corresponds to the scene, refers to the specific flight trajectory and attitude followed by the drone when shooting in the air. For example, it can include a circling mode, where the drone flies in a circular or elliptical pattern around the target; a gradually receding mode, where the drone gradually moves away from the target and ascends to shoot; or a straight-line flight mode, where the drone flies along a straight path and shoots. The choice of these modes is closely related to the characteristics of the scene.

[0050] For the atmosphere creation task, the shooting rules are obtained by searching the knowledge graph based on the atmosphere. The shooting parameters include the flight speed, flight altitude and camera parameters corresponding to the atmosphere. The atmosphere creation task runs through the entire shooting process.

[0051] Specifically, within the knowledge graph, the system retrieves atmosphere information parsed from user requests to obtain flight speed, altitude, and camera parameters matching the desired atmosphere. For example, for a "romantic" atmosphere, it retrieves information on low-speed, low-altitude flight, coupled with a large aperture and low ISO camera settings. Based on preset atmosphere types (such as cheerful, solemn, or mysterious), the system automatically triggers corresponding atmosphere-creating tasks and retrieves appropriate shooting strategies from the knowledge graph. Camera parameters, including aperture, shutter speed, ISO, and white balance, collectively determine the image's brightness, depth of field, motion blur, and color reproduction, directly impacting the final visual atmosphere. During shooting, the drone can dynamically adjust camera parameters based on real-time changes in ambient light or the subject's emotional expression to maintain or enhance the desired atmosphere.

[0052] For multiple priority tasks, identify the subject in each priority task, retrieve the motion trajectory type corresponding to the atmosphere from the knowledge graph based on the atmosphere, and set the object of motion trajectory tracking as the subject.

[0053] Specifically, within the knowledge graph, the subject in each priority task is first identified, for example, through image recognition or user specification. Then, based on the atmosphere information parsed from the user request, the knowledge graph retrieves the motion trajectory type corresponding to that atmosphere, and the identified subject is set as the object of the motion trajectory tracking. For example, in a "wedding" scenario, the bride and groom might be the subjects of a priority task; the system would retrieve a soft circling or following trajectory based on the "romantic" atmosphere and have the drone track the bride and groom. Identifying the subject in each priority task can be achieved through various technologies. For example, a deep learning-based object detection algorithm can be used to identify predefined subject categories in real-time drone footage. Motion trajectory types can include following trajectories, circling trajectories, or ascending / descending trajectories.

[0054] Through this task-based, contextualized retrieval and rule-based application mechanism, the solution proposed in this application can transform the general photography knowledge stored in the knowledge graph into specific and executable shooting instructions for specific shooting situations and tasks. This effectively solves the problem of how to transform abstract shooting intentions into refined aerial shooting plans, making the generated shooting plans more targeted, artistic, and operable.

[0055] S3: Based on the retrieved shooting rules and semantic constraints, generate a drone shooting plan using a large language model.

[0056] The scheme is expressed as a high-level target sequence, including the flight action types, camera operation commands, event monitoring conditions, and action triggering logic for each stage. Each command in the shooting scheme corresponds to one or more sub-tasks and is bound to the range of flight parameters returned by the knowledge graph.

[0057] S4: Based on preset rules, perform a feasibility check on the shooting plan. If the overall shooting plan fails the check, the shooting plan is rejected, the reason for failure is generated, a prompt is given to the large language model, and the shooting plan is regenerated until the overall shooting plan passes the check.

[0058] Specifically, the feasibility check verifies the generated shooting plan item by item according to preset rules. These preset rules include kinematic consistency, environmental constraints, and regulatory compliance. The kinematic consistency rule requires that the planned flight trajectory and motion parameters in the shooting plan must not exceed the physical capability boundaries of the UAV, including the UAV's maximum thrust, battery life, wind resistance, and maximum tilt angle. The environmental constraints include the path not intersecting with obstacles and the minimum distances between the UAV and people and moving objects. Regulatory compliance includes permissible flight range and altitude restrictions.

[0059] In step S4, a feasibility check is performed on the shooting plan based on preset rules, including: Based on the 3D environmental model, the range of personnel movement, and the maximum height of personnel, a safe flight zone is formed within the 3D environmental model. This constructs a precise and dynamic safe flight space to ensure that the drone does not pose a hazard to personnel or the environment when performing its filming mission.

[0060] Specifically, a safe flight zone can be formed by acquiring precise 3D point cloud data of the environment using technologies such as LiDAR scanning, stereo vision systems, or structured light sensors. This data, combined with pre-defined safety margins, incorporates information on the range of movement and maximum height of personnel, and is constructed using voxelization or meshing methods. Alternatively, pre-drawn site CAD drawings or BIM models can be used as the basis for a 3D environmental model. This model can be combined with real-time or estimated personnel location information (e.g., obtained through UWB positioning systems or visual tracking systems) to dynamically exclude the space occupied by personnel in the 3D model, thereby generating a safe flight zone that changes with personnel movement.

[0061] Based on the preset action path of the main body, the motion trajectory type and tracking relationship retrieved from the priority tasks, the corresponding drone operation trajectory is generated.

[0062] Specifically, the drone's trajectory can be generated using path planning algorithms, such as sampling-based path planning methods, combined with the subject's pre-set movement path (e.g., a pre-set performance route or travel route) and specific motion trajectory types retrieved from a knowledge graph (e.g., circling, following, diving, etc.), to generate a smooth drone trajectory that meets the shooting requirements. The tracking relationship guides the algorithm to adjust the drone's position and attitude as the subject moves. Alternatively, a model predictive control (MPC) approach can be used, taking the subject's movement path, motion trajectory type, and tracking relationship as inputs to the controller to calculate and generate the drone's optimal trajectory in real time.

[0063] Analyze whether the drone's flight path is entirely within the safe flight zone. If yes, the test is considered passed; otherwise, the test is considered failed.

[0064] Specifically, by comparing the planned drone trajectory with the previously established safe flight zone in terms of spatial geometry, the risk of collision with obstacles or people can be quickly determined. This analysis uses collision detection algorithms to represent the drone trajectory as a series of discrete points or a continuous curve, and the safe flight zone as a three-dimensional voxel or polygonal mesh. The safety of the trajectory is determined by checking whether each point or segment on the trajectory intersects with or exceeds the boundary of the safe flight zone.

[0065] Further analysis of the drone's flight trajectory and planned flight time is conducted to determine whether the drone's motion parameters meet the requirements of kinematic consistency and regulatory compliance. If so, the test is deemed passed; otherwise, the test is deemed failed.

[0066] Specifically, the analysis not only focuses on the safety of the flight path but also ensures that the drone's kinematic parameters, such as speed, acceleration, and attitude, do not exceed its performance limits and comply with all relevant flight regulations while executing the path. This analysis involves time-parameterizing the drone's trajectory to calculate kinematic parameters such as speed, acceleration, and angular velocity at each time point. These calculated parameters are then compared with kinematic consistency indicators such as the drone's maximum thrust, battery life, wind resistance, and maximum tilt angle. Simultaneously, the analysis checks whether the trajectory exceeds regulatory compliance requirements such as permissible flight range and altitude restrictions.

[0067] If all preset rules are deemed acceptable, the shooting plan is considered acceptable overall; otherwise, the shooting plan is considered unacceptable overall.

[0068] This multi-dimensional, hierarchical inspection mechanism significantly improves the safety, reliability, and feasibility of the generated shooting plan, reduces flight risks, and ensures the smooth completion of the shooting mission. This avoids flight accidents or shooting failures caused by unreasonable plans, and improves the intelligence level and user experience of drone flight shooting.

[0069] The reasons for failure include: the drone's flight path intersecting with the safe flight area, flight parameters exceeding limits, and flight altitude or range exceeding limits. Based on this, in step S4, a large language model is provided with prompts to regenerate the shooting plan, including: If the failure is due to the drone's trajectory intersecting with the safe flight area, the path segment outside the safe flight area will be removed, and the severed trajectory will be reconnected within the safe flight area; if reconnection is not possible, the filming plan will be rejected.

[0070] Specifically, after removing unsafe path segments, if the original trajectory is divided into multiple discontinuous segments, a new path entirely within the safe flight area needs to be found to connect these truncated trajectory segments. Various path planning algorithms can be used to solve this problem. For example, the RRT (Fast Random Tree) algorithm or the potential field method can be used to plan a collision-free path from one truncation point to another within a space constrained by the safe flight area. Alternatively, a series of intermediate waypoints can be generated within the safe flight area, and then a smooth connecting path can be generated using curve interpolation (such as Bézier curves or B-spline curves).

[0071] When the drone's flight path intersects with a safe flight area, the system can perform intelligent local repair instead of simply regenerating the entire system. This significantly improves the efficiency and success rate of shooting scheme generation and reduces unnecessary computational resource consumption. Simultaneously, this mechanism ensures the safety of the drone's flight path, avoiding potential safety hazards that might recur due to blind regeneration. Even if local repair fails, it provides a clearer reason for the failure, offering more precise guidance for the subsequent optimization and generation of the large language model, thus making the entire drone flight shooting scheme generation process more robust and efficient.

[0072] S5: Convert the approved shooting plan into flight control commands and transmit them to the UAV flight controller for execution.

[0073] like Figure 2 As shown in the figure, an embodiment of this application provides a context-aware UAV flight photography scheme generation system, comprising: The semantic reasoning module is used to parse the intent of user requests using a large language model, decompose the task, extract the subject, atmosphere and scene, and transform the extracted content into semantic constraints.

[0074] Specifically, this module typically runs on the cloud (for heavy-duty inference) or on high-performance edge computing devices such as drones. The semantic inference module is used to perform intent parsing and task decomposition on the content requested by the user (voice or text). For example, if the user requests "shoot a wedding reception, focusing on the bride and creating a romantic atmosphere," the LLM parser decomposes it into: Subject: Bride (requires the cooperation of the visual recognition module); Scene: Wedding reception (implying constraints of low noise, slow movement, and maintaining a safe distance); Atmosphere: Romantic (implying soft lighting, slow panning, and the color temperature of the golden hour).

[0075] The knowledge graph retrieval module is used to perform retrieval in the knowledge graph and return shooting rules that match the semantic constraints. The shooting rules include shooting modes and shooting parameters. In the knowledge graph, atmosphere and subject correspond to the shooting rules, respectively.

[0076] Specifically, the system references a knowledge graph that stores the ontology of photography. This serves as a bridge connecting natural language and engineering parameters. It effectively utilizes prior knowledge, preventing the meaningless illusions generated by LLM and ensuring that the generated solutions conform to the aesthetic standards of the film industry.

[0077] The flight plan generation module is used to generate drone shooting plans based on retrieved shooting rules and semantic constraints using a large language model.

[0078] The feasibility check module is used to check the feasibility of the shooting plan based on preset rules. If the shooting plan fails the overall check, the shooting plan is rejected, the reason for failure is generated, the large language model is given a prompt, and the shooting plan is regenerated until the shooting plan passes the overall check.

[0079] Specifically, an LLM might generate code instructions that require the drone to pass through walls or fly at accelerations exceeding the motor's limits.

[0080] Kinematic consistency check: The generated trajectory will be compared with the drone's physical constraints model (maximum thrust, battery life, wind resistance). If the LLM requires a 50mph sharp turn within 1 meter, the feasibility check module will mark it as "infeasible".

[0081] Environmental constraints are enforced: Utilizing real-time sensing data (LiDAR / stereo depth maps), the feasibility check module constructs a safe flight corridor, i.e., a safe flight zone. If the generative plan intersects with an obstacle (such as a tree), the filter modifies the path using the obstacle control function, or rejects the plan entirely, prompting the LLM to generate a safer alternative, for example, if a tree is detected and cannot be crossed, fly over it instead.

[0082] Regulatory compliance check: The system will check based on the local geofence (which constitutes the flight range) and altitude restriction database.

[0083] The scheme conversion module is used to convert the approved shooting scheme into flight control commands and transmit them to the UAV flight controller for execution.

[0084] In this embodiment, the semantic reasoning module can parse vague expressions such as "I want to film my daughter playing soccer; keep it safe but make it dynamic" into structured semantic constraints; the knowledge graph retrieval module, based on photographic knowledge, maps atmosphere and subject to specific shooting rules, such as associating "dynamic" atmosphere with high-speed flight and rapid zoom parameters; the flight plan generation module generates shooting plans containing multi-stage behaviors; the feasibility check module ensures the safety and feasibility of the plan through 3D environmental models and kinematic analysis, and provides feedback to optimize the plan when path conflicts or parameter exceedances are detected. Through the above technical solutions, the system can automatically handle complex user requests, connect behaviors such as circling, following, and hovering without manual waypoint setting, and predict key events such as the moment of a goal to adjust the camera position in advance, thereby effectively solving the technical problems pointed out in the background art and significantly improving the intelligence level and user experience of drone shooting.

[0085] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0086] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for generating drone flight photography schemes based on context awareness, characterized in that, include: A large language model is used to parse the intent of user requests, decompose the task, extract the subject, atmosphere and scene, and transform the extracted content into semantic constraints. A search is performed in the knowledge graph to return shooting rules that match the semantic constraints. The shooting rules include shooting mode and shooting parameters. In the knowledge graph, atmosphere and subject correspond to shooting rules, respectively. Based on the retrieved shooting rules and semantic constraints, a large language model is used to generate a shooting plan for the drone. Based on preset rules, a feasibility check is performed on the shooting plan. If the shooting plan fails the overall check, the shooting plan is rejected, a reason for failure is generated, the large language model provides a prompt, and the shooting plan is regenerated until the overall shooting plan passes the check. The approved shooting plan is converted into flight control commands and transmitted to the drone flight controller for execution.

2. The method for generating UAV flight photography schemes based on context awareness as described in claim 1, characterized in that, The process employs a large language model to parse user requests, decomposes tasks, extracts the subject, atmosphere, and scene, and transforms the extracted content into semantic constraints, including: The user request is parsed using a large language model, which first extracts the subject, atmosphere, scene, idea, and implicit intent; Generate the main task based on the concept image, theme, atmosphere, and scene; Based on the main task and implicit intent, phased tasks are generated, including environmental framing tasks, atmosphere creation tasks, and multiple priority tasks. Break down the phased tasks into sub-tasks.

3. The method for generating UAV flight photography schemes based on context awareness as described in claim 1, characterized in that, The knowledge graph organizes knowledge entities based on photography, and the stored mapping and matching rules include: the correspondence between atmosphere tags and flight speed, flight altitude, motion trajectory type and camera parameters; and the correspondence between scene type and safety distance requirements, flight noise constraints, recommended shooting angle and aerial photography motion mode.

4. The method for generating UAV flight photography schemes based on context awareness as described in claim 3, characterized in that, The process of searching the knowledge graph and returning shooting rules that match the semantic constraints includes shooting modes and shooting parameters, including: For environmental framing tasks, the knowledge graph is searched based on the scene to obtain shooting rules. The shooting mode adopts the aerial photography motion mode corresponding to the scene, and the shooting parameters include the safety distance requirements, flight noise constraints and recommended shooting angles corresponding to the scene. The environmental framing task obtains the entire environmental layout and constructs a 3D environmental model. For the atmosphere creation task, the shooting rules are obtained by searching the knowledge graph based on the atmosphere. The shooting parameters include the flight speed, flight altitude and camera parameters corresponding to the atmosphere. The atmosphere creation task runs through the entire shooting process. For multiple priority tasks, identify the subject in each priority task, retrieve the motion trajectory type corresponding to the atmosphere from the knowledge graph based on the atmosphere, and set the object of motion trajectory tracking as the subject.

5. The method for generating UAV flight photography schemes based on context awareness as described in claim 1, characterized in that, The preset rules include kinematic consistency, environmental constraints, and regulatory compliance. Kinematic consistency includes the drone's maximum thrust, battery life, wind resistance, and maximum tilt angle. Environmental constraints include the path not intersecting with obstacles and the minimum distances between the drone and people and moving objects. Regulatory compliance includes permitted flight range and altitude limits.

6. The method for generating UAV flight photography schemes based on context awareness as described in claim 5, characterized in that, The feasibility check of the shooting plan based on preset rules includes: Based on the 3D environmental model, the range of personnel movement, and the maximum height of personnel, a safe flight zone is formed in the 3D environmental model; Based on the preset action path of the main body, the motion trajectory type and tracking relationship retrieved from the priority tasks, the corresponding drone operation trajectory is generated; Analyze whether the drone's flight path is entirely within the safe flight zone. If yes, the test is considered passed; otherwise, the test is considered failed. Further analysis of the drone's flight trajectory and planned flight time is conducted to determine whether the drone's motion parameters meet the requirements of kinematic consistency and regulatory compliance. If yes, the test is deemed passed; otherwise, the test is deemed failed. If all preset rules are deemed acceptable, the shooting plan is considered acceptable overall; otherwise, the shooting plan is considered unacceptable overall.

7. The method for generating UAV flight photography schemes based on context awareness as described in claim 1, characterized in that, The reasons for failure include: the drone's flight path intersecting with the safe flight area, flight parameters exceeding limits, and flight altitude or flight range exceeding limits.

8. The method for generating UAV flight photography schemes based on context awareness as described in claim 7, characterized in that, The step of providing prompts to the large language model and regenerating the shooting plan includes: If the failure is due to the drone's trajectory intersecting with the safe flight area, the path segment outside the safe flight area will be removed, and the severed trajectory will be reconnected within the safe flight area; if reconnection is not possible, the filming plan will be rejected.

9. A context-aware UAV flight photography scheme generation system, characterized in that, include: The semantic reasoning module is used to parse the intent of user requests using a large language model, decompose the task, extract the subject, atmosphere and scene, and transform the extracted content into semantic constraints. The knowledge graph retrieval module is used to perform retrieval in the knowledge graph and return shooting rules that match the semantic constraints. The shooting rules include shooting mode and shooting parameters. In the knowledge graph, atmosphere and subject correspond to the shooting rules respectively. The flight plan generation module is used to generate drone shooting plans based on retrieved shooting rules and semantic constraints using a large language model. The feasibility check module is used to check the feasibility of the shooting plan based on preset rules. If the shooting plan fails the overall check, the shooting plan is rejected, the reason for failure is generated, the large language model provides a prompt, and the shooting plan is regenerated until the shooting plan passes the overall check. The scheme conversion module is used to convert the approved shooting scheme into flight control commands and transmit them to the UAV flight controller for execution.