Unmanned aerial vehicle shooting system control method based on adaptive optimization
By constructing a multimodal semantic graph and graph neural network, combining Bayesian optimization and attention mechanism, and dynamically adjusting drone shooting parameters, the problem of unstable image quality during drone inspections is solved, and efficient and intelligent image acquisition and target recognition are achieved.
Patent Information
- Application Number
- CN202510840581.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Existing drones have unstable image quality and poor imaging quality during distribution network inspections, low intelligence levels, difficulty in adapting to complex and changing environments, and lack real-time evaluation and feedback adjustment mechanisms.
Construct a multimodal semantic graph, use graph neural network to extract semantic context information, combine Bayesian strategy optimization and attention mechanism to dynamically adjust shooting control parameters, introduce transfer learning to update the control strategy, and achieve adaptive optimization.
Accurately predict the spatial distribution of target equipment and the difficulty of shooting, improve the intelligence level and image quality of drone inspections, ensure the accuracy and stability of target recognition, and adapt to complex environmental changes.
Smart Images

Figure CN120610469A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) control, and in particular to a UAV photography system control method based on adaptive optimization. Background Art
[0002] With the development of intelligent power systems, distribution network equipment is vast and widely distributed. Traditional manual inspection methods are time-consuming, inefficient, and pose high safety risks. Drones, with their flexibility and unique viewing angles, are becoming an increasingly important means of distribution network inspection. However, in practice, the quality of drone-captured images is affected by factors such as ambient lighting variations, dynamic equipment position shifts, inaccurate gimbal control, and improperly set capture parameters. This makes it difficult to ensure the validity and accuracy of inspection data.
[0003] Existing technologies mostly rely on fixed parameters or simple rules to control the gimbal and camera, which makes it difficult to adapt to complex and changing inspection environments. They also lack real-time evaluation and feedback adjustment mechanisms for the quality of captured images, affecting the intelligence level and practical value of drone inspections.
[0004] Therefore, a UAV photography system control method based on adaptive optimization is proposed. Summary of the Invention
[0005] The present invention provides a method for controlling a drone photography system based on adaptive optimization, aiming to solve the problems of unstable image capture, poor imaging quality, and low intelligence level in existing drones during distribution network inspection missions.
[0006] To achieve the above object, the present invention provides the following technical solutions: The control method of the UAV shooting system based on adaptive optimization includes: Construct a multimodal semantic graph to uniformly model the target task description, equipment type, geographic location information, historical inspection data, lighting environment, meteorological information, and optical flow features, and generate a graph structure representation; Based on the semantic graph, a graph neural network is used to extract semantic context information to predict the spatial distribution probability of target devices, the shooting difficulty level, and key route nodes. Generate a candidate set of control strategies and use the Bayesian strategy optimization method to select the optimal shooting control parameter configuration suitable for the current environment from multiple prior control strategies; The UAV flies according to the key route nodes, collecting edge vision recognition data, IMU attitude information, ambient lighting information and target distribution status within the frame in real time; Dynamically fuse multi-source information based on the attention mechanism, driving the joint control module to simultaneously adjust shooting control parameters to adapt to complex imaging requirements; If the confidence level of the detected image recognition is lower than the set threshold, the retake control logic is driven based on the semantic graph to trigger local fine-tuning control adjustment; The graph state, control strategy and shooting results of the entire current task process are used as samples to input the strategy evolution module, and the control strategy model is updated using the transfer learning method to improve the adaptability of subsequent tasks.
[0007] Furthermore, the steps of generating the graph structure representation include: The target task description, distribution network equipment type, geographic location information, historical inspection data, lighting environment, meteorological information and image optical flow features are uniformly encoded; The entities corresponding to various types of information are used as graph nodes, and the relationships between different entities are used as graph edges to construct a multimodal heterogeneous graph; Assign a corresponding attribute vector to each node in the graph to generate a unified semantic graph representation.
[0008] Furthermore, the step of performing reasoning by the graph neural network includes: The constructed semantic graph is input into the graph neural network model, and the graph attention mechanism is used to aggregate the multi-dimensional attribute information of the nodes in the graph; Extract the semantic context features of each device node and combine them with the environmental status and task requirements of its adjacent nodes to achieve deep encoding of spatial structure information; During the training phase, the graph neural network is trained on multiple tasks using historical inspection trajectories and image quality as supervisory signals. In the inference stage, the spatial distribution probability of each device node, the corresponding shooting difficulty level, and the prediction results of whether it is a key route node are output.
[0009] Furthermore, the steps of screening the optimal shooting control parameter configuration include: Generate a control strategy candidate set based on the graph neural network inference results, wherein the candidate set includes multiple prior control strategies for different lighting conditions, target types, and spatial positions, and the control strategies include at least gimbal angle, shutter parameters, exposure time, and zoom level; Build a Bayesian optimization model, use the graph inference results as context input, and the captured image quality index as the evaluation function; Using the Bayesian strategy optimization method, the candidate control strategy set is iteratively sampled and scored, and the posterior expected performance of each strategy is calculated; The control strategy with the highest Bayesian expectation score is selected as the optimal shooting control parameter configuration under the current environment.
[0010] Furthermore, the step of adjusting the shooting control parameters includes: Building a multi-source information fusion network based on the attention mechanism. The multi-source information includes device semantic features, graph inference parameters, real-time image feedback, and flight status information. The weight distribution of different information sources in the current task state is calculated through attention weights to achieve dynamic weighted fusion of feature dimensions; The fused feature vector is input into the joint control module to drive the gimbal control subsystem and the camera parameter subsystem to operate in coordination. The coordinated control includes simultaneously adjusting the pitch and yaw angles of the gimbal, and the exposure parameters, zoom level, and shutter speed of the camera.
[0011] Furthermore, the fine-tuning control adjustment includes: Analyze the semantic nodes and contextual edge information associated with targets with low recognition confidence in the current image to identify potential interference factors and key imaging requirements; constructing a local perception vector based on the information as an input feature of the fine-tuning control module; A lightweight control strategy network is used to adjust the current gimbal attitude, camera parameters, and heading attitude, and a progressive parameter adjustment strategy is used to gradually approach the optimization target. After performing the adjustments, the image is recaptured and the image recognition module is reused to evaluate the recognition confidence of the updated image.
[0012] Furthermore, the steps of updating the control strategy model using the transfer learning method include: Collect flight status data, image quality indicators and corresponding control parameters under the current mission environment to build a mission domain sample set; A control strategy model pre-trained in similar environment tasks is selected as the source model, and its shallow network parameters and control feature expression structure are extracted; Freeze the network layers related to common features in the source model and only fine-tune the high-level policy decision module; Incremental learning is performed based on the current task domain sample set, and the model parameters are locally updated by minimizing the control loss function and the image quality score error.
[0013] Furthermore, the method supports the control center to manually take over the control authority of the drone, including: When a task change or manual intervention request is triggered, the control center suspends the current automatic inspection task through remote instructions and takes over manually; In manual takeover mode, the control interface receives real-time information on the drone's flight status, camera images, and mission progress to enable remote control operations or mission parameter adjustments.
[0014] The beneficial effects of the present invention are: 1. By constructing a semantic graph that integrates multi-source data such as task description, equipment type, geographic location, and environmental conditions, and introducing graph neural networks for contextual semantic reasoning, it is possible to accurately predict the spatial distribution probability and shooting difficulty level of target equipment, automatically identify key inspection nodes, and achieve task-oriented optimal route planning, greatly improving the intelligence level and scheduling efficiency of drones in performing complex inspection tasks.
[0015] 2. Based on the fusion of Bayesian strategy optimization and attention mechanism, the gimbal posture and camera imaging parameters are dynamically adjusted to adapt to interference factors such as lighting, posture, and target changes. A confidence-driven fine-tuning re-shooting mechanism and a transfer learning update strategy control model are introduced to achieve continuous adaptation to dynamic task environments and performance evolution, ensuring stable image quality and accurate target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a flow chart of the UAV photography system control method based on adaptive optimization provided by the present invention. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0018] Example 1 The control method of UAV shooting system based on adaptive optimization, such as Figure 1 Shown, including: The control method of the UAV shooting system based on adaptive optimization includes: S100: Construct a multimodal semantic graph to uniformly model the target task description, equipment type, geographic location information, historical inspection data, lighting environment, meteorological information, and optical flow features, and generate a graph structure representation; Furthermore, the steps of generating the graph structure representation include: The target task description, distribution network equipment type, geographic location information, historical inspection data, lighting environment, meteorological information and image optical flow features are uniformly encoded; The entities corresponding to various types of information are used as graph nodes, and the relationships between different entities are used as graph edges to construct a multimodal heterogeneous graph; Assign a corresponding attribute vector to each node in the graph to generate a unified semantic graph representation.
[0019] Specifically, information related to the distribution network equipment inspection task is first obtained from multiple data sources, including: target task description (such as task type, urgency level, execution time, etc.), distribution network equipment type and its geographic location information, historical inspection data (such as fault records, maintenance frequency, past image samples), environmental information (including current and predicted light intensity, weather conditions, wind speed and direction, etc.), image optical flow features (used to reflect the dynamic changes of the target and background, and assist in image stability analysis).
[0020] The above information is input into a unified data preprocessing module and encoded using the following methods: text information (task description, equipment type) is embedded and encoded using a word vector model (such as Word2Vec and BERT); geographic spatial information is normalized using latitude and longitude coordinates; historical inspection data is encoded into statistical feature vectors after feature extraction (such as fault frequency statistics and time series aggregation); ambient lighting and meteorological information are converted into structured numerical vectors; image optical flow features are extracted from the preview image sequence collected by the drone, and the dense optical flow tensor is obtained using an optical flow estimation algorithm (such as PWC-Net), and then the dimension is reduced to an optical flow feature vector.
[0021] Next, the system constructs a graph based on the semantic and business relationships between various data types. The construction rules are as follows: Each type of information (task, equipment, location, environment, etc.) is considered a graph node type. If two types of information have a direct dependency or influence in task semantics or business logic (such as equipment and location, equipment and historical failures), a graph edge is established between their corresponding nodes. Graph edges can be labeled with the association type (such as "located," "historical failure association," "environmental impact," etc.). For each node, its corresponding encoding vector is further used as a node attribute to form a preliminary representation of the graph-structured data. Ultimately, the system constructs a multimodal, heterogeneous semantic graph.
[0022] The multimodal semantic graph construction method integrates heterogeneous task-related information (such as equipment type, location, historical data, environmental conditions, and visual features) into a unified graph structure, fully exploring its inherent semantic connections and contextual logic. This graph not only enables the system to gain a global understanding of complex inspection scenarios but also provides rich, high-dimensional input features for subsequent graph neural networks.
[0023] S200: Based on the semantic graph, a graph neural network is used to extract semantic context information to predict the spatial distribution probability of the target device, the shooting difficulty level, and key route nodes; Furthermore, the step of performing reasoning by the graph neural network includes: The constructed semantic graph is input into the graph neural network model, and the graph attention mechanism is used to aggregate the multi-dimensional attribute information of the nodes in the graph; Extract the semantic context features of each device node and combine them with the environmental status and task requirements of its adjacent nodes to achieve deep encoding of spatial structure information; During the training phase, the graph neural network is trained on multiple tasks using historical inspection trajectories and image quality as supervisory signals. In the inference stage, the spatial distribution probability of each device node, the corresponding shooting difficulty level, and the prediction results of whether it is a key route node are output.
[0024] Specifically, all nodes in the semantic graph (such as devices, tasks, environmental status, etc.) and their corresponding attribute vectors are taken as input, and the graph attention mechanism in the graph neural network is used to perform weighted aggregation on the information of adjacent nodes to ensure that the system can dynamically adjust the feature fusion strength according to the importance of different adjacent nodes, thereby effectively extracting key contextual information.
[0025] Through graph convolutional propagation, each device node continuously integrates multi-source information such as task description, historical status, and surrounding environmental conditions (such as lighting, weather, and geographic location) in multiple rounds of information updates to obtain a semantically rich device node representation vector.
[0026] During the model training phase, a supervised multi-task learning strategy is employed, using the frequency of equipment appearances in historical drone inspection routes and image quality scores as label data. The network is trained to output the spatial distribution probability of a device (i.e., the likelihood that a device will be prioritized for inspection in future missions). Furthermore, the network learns the difficulty level of each device (e.g., whether it is located in an obstructed area or in complex lighting conditions), as well as whether it is a critical route node (i.e., a node that affects the overall efficiency of the mission).
[0027] During the inference phase, after inputting the graph constructed in real time, the graph neural network can output three types of prediction results for all target device nodes: spatial distribution probability, shooting difficulty level, and whether it is a key route node.
[0028] By constructing a semantic graph and introducing a graph neural network to extract contextual information, this approach achieves high-precision predictions of the target device's spatial distribution, shooting difficulty, and key route nodes, demonstrating significant intelligent and adaptive advantages. Compared to traditional rule-based or single-feature approaches, this method integrates multi-source heterogeneous data (such as environmental, historical mission, and image information) for deep semantic modeling, significantly improving the scientific nature of path planning and control decisions.
[0029] S300: Generate a candidate set of control strategies, and use a Bayesian strategy optimization method to select the optimal shooting control parameter configuration suitable for the current environment from multiple prior control strategies; Furthermore, the steps of screening the optimal shooting control parameter configuration include: Generate a control strategy candidate set based on the graph neural network inference results, wherein the candidate set includes multiple prior control strategies for different lighting conditions, target types, and spatial positions, and the control strategies include at least gimbal angle, shutter parameters, exposure time, and zoom level; Build a Bayesian optimization model, use the graph inference results as context input, and the captured image quality index as the evaluation function; Using the Bayesian strategy optimization method, the candidate control strategy set is iteratively sampled and scored, and the posterior expected performance of each strategy is calculated; The control strategy with the highest Bayesian expectation score is selected as the optimal shooting control parameter configuration under the current environment.
[0030] Specifically, based on the results of graph neural network inference, such as the spatial distribution probability of devices, the shooting difficulty level, and key flight path nodes, combined with contextual information such as ambient lighting, weather conditions, and device type, a set of prior control strategies is retrieved or generated to form a candidate control strategy set. Each strategy in this candidate set includes shooting control parameters such as gimbal angle (pitch and yaw), shutter parameters, exposure time, and zoom level, and is applicable to different target and environment combinations.
[0031] A Bayesian optimization model is constructed, using the output of graph inference as contextual conditions. Image quality metrics (such as clarity, object integrity, and exposure suitability) are defined as optimization objectives. Based on this, a Bayesian policy optimization process performs multiple rounds of sampling, scoring, and updating of candidate policy sets. In each round, the model estimates the posterior expected performance of each control policy using a Gaussian process or tree-structured Bayesian regression algorithm and selects the next sampling policy based on an acquisition function (such as expected improvement (EI) or upper confidence bound (UCB).
[0032] Finally, the control strategy with the highest expected image quality performance is selected as the optimal shooting control parameter configuration under the current environment and sent to the control module for execution to drive the drone to complete the high-quality image acquisition task.
[0033] A Bayesian strategy optimization approach can rapidly select the optimal shooting parameter configuration from multiple prior control strategies under diverse and complex environments and mission requirements, achieving coordinated optimization of gimbal angle, shutter speed, exposure time, and zoom level. This adaptive and efficient approach can improve drone imaging quality and mission completion rates in complex scenarios with dynamic lighting, heterogeneous targets, and spatially distributed objects.
[0034] S400: The UAV flies according to the key route nodes and collects edge vision recognition data, IMU attitude information, ambient lighting information and target distribution status within the frame in real time; S500: Dynamically integrates multi-source information based on the attention mechanism, driving the joint control module to simultaneously adjust shooting control parameters to meet complex imaging requirements; Furthermore, the step of adjusting the shooting control parameters includes: Building a multi-source information fusion network based on the attention mechanism. The multi-source information includes device semantic features, graph inference parameters, real-time image feedback, and flight status information. The weight distribution of different information sources in the current task state is calculated through attention weights to achieve dynamic weighted fusion of feature dimensions; The fused feature vector is input into the joint control module to drive the gimbal control subsystem and the camera parameter subsystem to operate in coordination. The coordinated control includes simultaneously adjusting the pitch and yaw angles of the gimbal, and the exposure parameters, zoom level, and shutter speed of the camera.
[0035] Specifically, we collect multi-source input information, including the semantic feature vector S of the distribution network equipment, the spatial distribution and difficulty feature vector G of the graph neural network output, the real-time image feedback feature I, and the UAV flight status information F. These features are spliced together to form a joint feature vector , then, Enter the multi-head self-attention mechanism module to calculate the attention weights between each information source. The specific calculation steps are as follows: Perform linear transformation on the input features to obtain query (Q), key (K), and value (V) matrices: ; in, 、 and is the learnable parameter matrix. Then, the attention weight matrix is calculated: ; in, is the dimension of the key vector, used for scaling. Then, by Calculate weighted features and use a multi-head mechanism to The fused feature vector is obtained by splicing and transforming it through a linear layer. Finally, the joint control module receives the fused feature vector and outputs the pitch and yaw angle adjustment of the gimbal, the exposure parameter adjustment of the camera, the zoom level, and the shutter speed adjustment.
[0036] By introducing an attention mechanism to dynamically fuse multi-source information, it intelligently assigns control weights based on multi-dimensional information such as real-time image quality, device semantic features, graph inference results, and flight status, driving coordinated adjustment of gimbal and camera parameters. This approach significantly improves the drone's adaptability in complex shooting environments, ensuring reasonable image composition, accurate exposure, and clear and complete subject matter.
[0037] S600: If the confidence level of the detected image recognition is lower than a set threshold, the retake control logic is driven based on the semantic graph to trigger a local fine-tuning control adjustment; Furthermore, the fine-tuning control adjustment includes: Analyze the semantic nodes and contextual edge information associated with targets with low recognition confidence in the current image to identify potential interference factors and key imaging requirements; constructing a local perception vector based on the information as an input feature of the fine-tuning control module; A lightweight control strategy network is used to adjust the current gimbal attitude, camera parameters, and heading attitude, and a progressive parameter adjustment strategy is used to gradually approach the optimization target. After performing the adjustments, the image is recaptured and the image recognition module is reused to evaluate the recognition confidence of the updated image.
[0038] Specifically, after image acquisition is complete, the system feeds the captured images into an edge-based object recognition module, which then outputs the recognition results and corresponding confidence levels. If the recognition confidence level of any key target falls below a preset threshold (e.g., 0.6), the system immediately initiates a retake control process.
[0039] The system focuses on the semantic nodes corresponding to targets with low recognition confidence, traverses their adjacent nodes and edge attributes (such as task type, lighting conditions, background interference type, etc.) in the pre-built multimodal semantic graph, and analyzes potential interference factors and imaging bottlenecks that may lead to imaging failure (such as backlighting, occlusion or insufficient target size).
[0040] The node features corresponding to the failed target in the recognition result, together with its context nodes (such as time, location information, light intensity, adjacent target type, etc.) are encoded into a local semantic vector , combined with the current imaging state (image brightness, blur, shooting angle) to form the input vector , for use by the fine-tuning control module.
[0041] Introducing lightweight control strategy network ,by For input and output adjustment actions , including: minor adjustments to the gimbal pitch and yaw angles; moderate changes to the camera exposure time, shutter speed, and ISO; and minor corrections to the flight direction (heading angle) or altitude when necessary. The system uses predefined step sizes. Execute a progressive parameter adjustment strategy: ; in, Indicates the current action. Indicates the updated action. Gradually approach the optimized imaging state to avoid drastic changes that may cause image degradation or unstable posture.
[0042] After performing parameter adjustments, re-collect the image and input the image into the recognition module again to obtain a new confidence evaluation value .like If the confidence level exceeds the set threshold, the image is recorded and the reshoot process is exited; otherwise, a maximum of N fine-tuning iterations can be performed until the set confidence level is reached or the maximum number of retries is reached.
[0043] By introducing a local fine-tuning control mechanism driven by semantic graphs, when image recognition confidence is insufficient, it can intelligently analyze the contextual factors of imaging failure, locate potential interference sources, and use a lightweight control strategy network to perform progressive and fine-tuning of the gimbal posture, camera parameters, and flight posture, thereby effectively improving the robustness and intelligent response capabilities of image acquisition, reducing the number of re-flights, ensuring target recognition accuracy, and enhancing the system's imaging stability and autonomous compensation capabilities in complex environments.
[0044] S700: The graph state, control strategy, and shooting results of the entire current task process are input into the strategy evolution module as samples, and the control strategy model is updated using the transfer learning method to improve the adaptability of subsequent tasks.
[0045] Furthermore, the steps of updating the control strategy model using the transfer learning method include: Collect flight status data, image quality indicators and corresponding control parameters under the current mission environment to build a mission domain sample set; A control strategy model pre-trained in similar environment tasks is selected as the source model, and its shallow network parameters and control feature expression structure are extracted; Freeze the network layers related to common features in the source model and only fine-tune the high-level policy decision module; Incremental learning is performed based on the current task domain sample set, and the model parameters are locally updated by minimizing the control loss function and the image quality score error.
[0046] Specifically, after the current inspection task is completed, the system automatically organizes and collects the IMU attitude information during the UAV flight, the graph reasoning status (such as the spatial distribution prediction of the device nodes, the difficulty level, the control parameter selection, etc.), the actual shooting control parameters (including the gimbal attitude, exposure time, shutter speed, zoom level, etc.), and the quality score of the corresponding captured image (such as clarity, completeness, exposure score) and recognition confidence and other indicators to form a sample data set under the current task domain. ,in represents the input feature vector, Indicates image quality and control effect labels.
[0047] Select pre-trained models with good performance from control strategy models that have been trained in similar environments (such as lighting, weather, or device types) in historical tasks As a source model, this model is usually a multi-input and multi-output neural network structure, and its underlying network has the ability to extract common features of the environment and equipment.
[0048] Source Model The parameters of the first several layers (such as the image feature extraction layer and the graph state encoding layer) are frozen, that is, they do not participate in the subsequent back-propagation training; the strategy decision layer (such as the control action selection module) is retained as a trainable module to build the target model This structure allows transfer learning to focus more on improving policy adaptation and avoids the risk of overfitting caused by full parameter retraining.
[0049] Using the current task sample set Model Perform incremental fine-tuning training. Minimize the following composite loss function during training: ; in, represents the control action prediction error (such as the mean square error between the control parameters and the optimal parameters), represents the error between the predicted image quality score and the true score, is an adjustable weight coefficient used to balance control performance and imaging quality. After fine-tuning is completed, the updated control strategy model is obtained. .
[0050] By introducing a transfer learning mechanism, the task map status, control strategy, and shooting results are automatically integrated after each mission is completed, task domain samples are constructed, and the control strategy model is incrementally updated. This approach retains the general capabilities of the source model while performing local fine-tuning for specific environments and imaging requirements. This improves the model's adaptability and generalization performance under complex and changing conditions, enables continuous evolution and intelligent optimization of the control strategy, and effectively enhances the self-learning ability and long-term operational efficiency of the drone photography system.
[0051] Furthermore, the method supports the control center to manually take over the control authority of the drone, including: When a task change or manual intervention request is triggered, the control center can suspend the current automatic inspection task through remote instructions and manually take over; In manual takeover mode, the control interface receives real-time information on the drone's flight status, camera images, and mission progress to enable remote control operations or mission parameter adjustments.
[0052] This method supports the control center to manually take over the drone, and can intervene in time when the mission changes suddenly or in special circumstances, improving the flexibility and safety of drone inspections; by transmitting flight status and image information in real time, it ensures the operator's comprehensive perception and precise control of the drone, effectively preventing accidents.
[0053] Example 2 This embodiment uses the daily inspection of urban distribution network equipment as a specific application scenario to describe the specific operation process of the drone shooting system control method based on adaptive optimization.
[0054] First, the system obtains a list of distribution equipment involved in the inspection, along with its objectives and priorities. It also collects multi-source data, including the area's geographic information, historical inspection records, lighting conditions, weather conditions, and optical flow feature images. This heterogeneous information is then unified into a model to construct a multimodal semantic graph. This graph structure represents each entity node (such as equipment type, geographic coordinates, historical status, and so on) and their relationships.
[0055] Subsequently, a graph neural network is used to extract features from the semantic graph, and the context node information is aggregated through the graph attention mechanism to predict the spatial distribution probability of each target device, the current shooting difficulty level, and its criticality in the inspection task, and identify key route nodes as priority inspection objects.
[0056] Next, the system constructs a candidate set of control strategies based on the results of graph reasoning, including a priori parameter combinations under multiple lighting, angles, and device types (such as gimbal angle, shutter speed, exposure time, zoom level, etc.), and introduces a Bayesian strategy optimization method to screen out the optimal shooting control parameter configuration under the current environment through sampling and scoring iterations.
[0057] The drone flies autonomously according to key route nodes and collects edge vision recognition results, heading attitude (IMU) information, light intensity and distribution status of targets in image frames in real time.
[0058] During the flight, the attention mechanism is used to fuse the above-mentioned multi-source information (including graph semantic features, current visual feedback, flight status, etc.), and the joint control module is used to coordinately adjust the shooting control parameters such as the gimbal angle and camera parameters to dynamically adapt to lighting changes, angle deviations and target feature differences to ensure imaging quality.
[0059] If the recognition module detects that the image confidence is lower than the set threshold, the system will trace back the relevant semantic graph nodes and context information to determine possible interference factors, and execute a progressive angle or parameter adjustment strategy through the local fine-tuning control module to re-shoot the local area to compensate for the incomplete target recognition problem.
[0060] Finally, the map states, flight control behaviors, shooting parameters, and imaging results recorded throughout the inspection mission are used as samples and fed into the strategy evolution module. Based on transfer learning, the general capabilities of the source control model are retained. By fine-tuning and updating the strategy decision module, the control model is continuously optimized and its adaptability is enhanced, thus providing more precise shooting control strategies for subsequent inspection missions.
[0061] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0062] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0063] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0064] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0065] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0066] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A control method for a UAV photography system based on adaptive optimization, characterized in that: include: Construct a multimodal semantic graph to uniformly model the target task description, equipment type, geographic location information, historical inspection data, lighting environment, meteorological information, and optical flow features, and generate a graph structure representation; Based on the semantic graph, a graph neural network is used to extract semantic context information to predict the spatial distribution probability of target devices, the shooting difficulty level, and key route nodes. Generate a candidate set of control strategies and use the Bayesian strategy optimization method to select the optimal shooting control parameter configuration suitable for the current environment from multiple prior control strategies; The UAV flies according to the key route nodes, collecting edge vision recognition data, IMU attitude information, ambient lighting information and target distribution status within the frame in real time; Dynamically fuse multi-source information based on the attention mechanism, driving the joint control module to simultaneously adjust shooting control parameters to adapt to complex imaging requirements; If the confidence level of the detected image recognition is lower than the set threshold, the retake control logic is driven based on the semantic graph to trigger local fine-tuning control adjustment; The graph state, control strategy and shooting results of the entire current task process are used as samples to input the strategy evolution module, and the control strategy model is updated using the transfer learning method to improve the adaptability of subsequent tasks.
2. The method for controlling a drone photography system based on adaptive optimization according to claim 1, wherein: The steps to generate a graph structure representation include: The target task description, distribution network equipment type, geographic location information, historical inspection data, lighting environment, meteorological information and image optical flow features are uniformly encoded; The entities corresponding to various types of information are used as graph nodes, and the relationships between different entities are used as graph edges to construct a multimodal heterogeneous graph. Assign a corresponding attribute vector to each node in the graph to generate a unified semantic graph representation.
3. The method for controlling a drone photography system based on adaptive optimization according to claim 1, wherein: The steps of performing reasoning in the graph neural network include: The constructed semantic graph is input into the graph neural network model, and the graph attention mechanism is used to aggregate the multi-dimensional attribute information of the nodes in the graph; Extract the semantic context features of each device node and combine them with the environmental status and task requirements of its adjacent nodes to achieve deep encoding of spatial structure information; During the training phase, the graph neural network is trained on multiple tasks using historical inspection trajectories and image quality as supervisory signals. In the inference stage, the spatial distribution probability of each device node, the corresponding shooting difficulty level, and the prediction results of whether it is a key route node are output.
4. The method for controlling a drone photography system based on adaptive optimization according to claim 1, wherein: The steps for screening the optimal shooting control parameter configuration include: Generate a control strategy candidate set based on the graph neural network inference results, wherein the candidate set includes multiple prior control strategies for different lighting conditions, target types, and spatial positions, and the control strategies include at least gimbal angle, shutter parameters, exposure time, and zoom level; Build a Bayesian optimization model, use the graph inference results as context input, and the captured image quality index as the evaluation function; Using the Bayesian strategy optimization method, the candidate control strategy set is iteratively sampled and scored, and the posterior expected performance of each strategy is calculated; The control strategy with the highest Bayesian expectation score is selected as the optimal shooting control parameter configuration under the current environment.
5. The method for controlling a drone photography system based on adaptive optimization according to claim 4, wherein: The steps to adjust the shooting control parameters include: Building a multi-source information fusion network based on the attention mechanism. The multi-source information includes device semantic features, graph inference parameters, real-time image feedback, and flight status information. The weight distribution of different information sources in the current task state is calculated through attention weights to achieve dynamic weighted fusion of feature dimensions; The fused feature vector is input into the joint control module to drive the gimbal control subsystem and the camera parameter subsystem to operate in coordination. The coordinated control includes simultaneously adjusting the pitch and yaw angles of the gimbal, and the exposure parameters, zoom level, and shutter speed of the camera.
6. The method for controlling a drone photography system based on adaptive optimization according to claim 1, wherein: The fine-tuning control adjustment includes: Analyze the semantic nodes and contextual edge information associated with targets with low recognition confidence in the current image to identify potential interference factors and key imaging requirements; constructing a local perception vector based on the information as an input feature of the fine-tuning control module; A lightweight control strategy network is used to adjust the current gimbal attitude, camera parameters, and heading attitude, and a progressive parameter adjustment strategy is used to gradually approach the optimization target. After performing the adjustments, the image is recaptured and the image recognition module is reused to evaluate the recognition confidence of the updated image.
7. The method for controlling a drone photography system based on adaptive optimization according to claim 1, wherein: The steps to update the control strategy model using transfer learning methods include: Collect flight status data, image quality indicators and corresponding control parameters under the current mission environment to build a mission domain sample set; A control strategy model pre-trained in similar environment tasks is selected as the source model, and its shallow network parameters and control feature expression structure are extracted; Freeze the network layers related to common features in the source model and only fine-tune the high-level policy decision module; Incremental learning is performed based on the current task domain sample set, and the model parameters are locally updated by minimizing the control loss function and the image quality score error.
8. The method for controlling a drone photography system based on adaptive optimization according to claim 1, wherein: The method supports the control center to manually take over the control authority of the drone, including: When a task change or manual intervention request is triggered, the control center suspends the current automatic inspection task through remote instructions and takes over manually; In manual takeover mode, the control interface receives real-time information on the drone's flight status, camera images, and mission progress to enable remote control operations or mission parameter adjustments.
Citation Information
Patent Citations
Unmanned aerial vehicle visual navigation method and device based on deep learning
CN119131643A
Urban greening system based on big data and optimization method thereof
CN119558450A
Unmanned aerial vehicle aerial image processing method, system and device and storage medium
CN119625576A
Unmanned aerial vehicle electric power inspection multi-modal data fusion system and unmanned aerial vehicle
CN120125947A
Method and device for automatically guiding an autonomous aircraft
US20240118706A1
Cited By
Large-scene multi-modal analysis system and method based on unmanned aerial vehicle
CN120894720A
Infrared target and alignment control method based on unmanned aerial vehicle
CN121115865A
Hunting camera imaging quality optimization method based on multimode data fusion
CN121304463A
Hunting camera imaging quality optimization method based on multi-mode data fusion
CN121304463B
Flight attitude compensation method and system for dynamic environment prediction
CN121454950A