Vehicle-mounted unmanned aerial vehicle follow-up system visual angle comfort degree adaptive control method and system

By establishing a geometric relationship model and a reinforcement learning control model to collaboratively optimize the drone's flight position and gimbal angle, the problem of improper viewpoint adjustment in the vehicle-mounted drone tracking system was solved, realizing adaptive viewpoint control in dynamic driving environments and improving visual comfort and stability.

CN122386654APending Publication Date: 2026-07-14DONGFENG MOTOR GRP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGFENG MOTOR GRP
Filing Date
2026-04-01
Publication Date
2026-07-14

Smart Images

  • Figure CN122386654A_ABST
    Figure CN122386654A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of unmanned aerial vehicle control, and provides a view angle comfort degree self-adaptive control method and system of a vehicle-mounted unmanned aerial vehicle follow-up shooting system. The method comprises the following steps: acquiring system state data; wherein the system state data comprises at least one of vehicle state information, unmanned aerial vehicle state information and visual feature information of the vehicle in a camera picture of the unmanned aerial vehicle; based on the system state data, a quantitative value of a generation cost representing the current follow-up shooting view angle comfort degree is calculated according to a pre-established geometric relation model; the system state data and the quantitative value of the generation cost are input into a pre-trained reinforcement learning control model, and a control instruction containing an unmanned aerial vehicle flight position adjustment amount and a gimbal angle adjustment amount is output by the reinforcement learning control model; the flight position of the unmanned aerial vehicle is adjusted and the shooting angle of the gimbal is adjusted according to the control instruction; and the steps are repeatedly executed. The method can quantitatively evaluate the comfort degree of the follow-up shooting view angle, realize multi-degree-of-freedom collaborative optimization control and autonomously maintain the optimal follow-up shooting view angle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) control technology, specifically to a method and system for adaptive control of viewpoint comfort in a vehicle-mounted UAV tracking system. Background Technology

[0002] With the rapid development of intelligent mobility and autonomous driving technologies, vehicle-mounted drone systems are gradually becoming important auxiliary tools for vehicles, widely used in scenarios such as wilderness exploration, road reconnaissance, driving recording, and creative photography. In these applications, drones need to continuously track vehicles in follow-up mode and transmit video footage in real time to help drivers obtain environmental information about the vehicle's surroundings and the road ahead. Currently, related technical solutions mainly focus on the implementation of basic flight control functions for drones, such as automatic takeoff, landing, follow-up flight, and obstacle avoidance.

[0003] However, the visual quality of the video footage acquired during vehicle-mounted drone tracking missions is also a crucial factor affecting user experience. In practical use, the shooting angle of the video directly determines the driver's ability to judge the relative position of the road ahead and the vehicle. If the shooting angle is too low, the image will feel oppressive, and the visible range of the road ahead will be severely compressed; if the shooting angle is too high, although the image is open, the sense of depth of the road ahead will be significantly reduced, which is not conducive to distance judgment. In addition, the proportion of the vehicle in the image also directly affects the driver's perception: if the vehicle is too small, it is difficult to discern its relative position to the road; if the vehicle is too large, it will encroach on the space for presenting road information in the image.

[0004] Current tracking and control solutions have significant limitations in addressing these issues. Firstly, most solutions simply place the vehicle in the center of the frame without considering or adjusting factors affecting visual quality, such as the angle of view and the size of the vehicle within the frame. This results in significant fluctuations in the visual quality of the image under different driving speeds and road conditions, making it difficult to maintain a consistently comfortable viewing experience. Secondly, existing adjustment methods are typically limited to adjusting the drone's flight position or the gimbal's shooting angle individually, lacking coordination between the two and failing to achieve ideal overall visual effects in dynamic driving scenarios. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method and system for adaptive control of viewing angle comfort in a vehicle-mounted drone following system. This method can quantitatively evaluate the comfort of the following viewing angle and achieve multi-degree-of-freedom collaborative optimization control of the drone's flight position and gimbal angle through reinforcement learning, thereby autonomously maintaining the best following viewing angle in a dynamic driving environment.

[0006] A first aspect of this invention provides a method for adaptive control of viewing comfort in a vehicle-mounted drone tracking system, comprising: Acquire system status data; wherein the system status data includes at least one of vehicle status information, UAV status information, and visual feature information of the vehicle in the UAV camera image; Based on the system status data, a quantitative value representing the comfort of the current following angle is calculated according to a pre-established geometric relationship model. The system state data and the quantified cost are input into a pre-trained reinforcement learning control model, and the reinforcement learning control model outputs control commands that include the UAV flight position adjustment amount and the gimbal angle adjustment amount. According to the control commands, the drone is coordinated to adjust its flight position and the gimbal is driven to adjust the shooting angle. The steps of calculating the quantified cost value representing the comfort of the current following viewpoint are repeated until the steps of coordinating the drone to adjust its flight position and driving the gimbal to adjust the shooting angle are executed to form a closed-loop control.

[0007] A second aspect of this invention provides a field-of-view comfort adaptive control system for a vehicle-mounted drone tracking system, comprising: A state perception unit is used to acquire system state data; wherein, the system state data includes at least one of vehicle state information, UAV state information, and visual feature information of the vehicle in the UAV camera image; The comfort assessment unit is used to calculate the quantitative value representing the comfort of the current following angle based on the system state data and a pre-established geometric relationship model. The decision unit is used to input the system state data and the quantified value into a pre-trained reinforcement learning control model, and the reinforcement learning control model outputs control commands including the UAV flight position adjustment amount and the gimbal angle adjustment amount. The collaborative execution unit is used to collaboratively drive the UAV to adjust its flight position and drive the gimbal to adjust its shooting angle according to the control commands. The comfort assessment unit, the decision-making unit, and the collaborative execution unit are configured to operate in a loop to form a closed-loop control.

[0008] The first aspect of this invention establishes a quantifiable evaluation mechanism for visual comfort based on a geometric relationship model, transforming subjective visual comfort into a calculable quantifiable value. By leveraging a pre-trained reinforcement learning control model, it directly outputs control commands containing flight position adjustment and gimbal angle adjustment in a continuous action space, achieving multi-degree-of-freedom collaborative optimization control between the UAV flight platform and the gimbal shooting angle. Through continuous iterative optimization via closed-loop control, it can autonomously maintain the optimal following angle in dynamic driving environments, improving the visual comfort and stability of the following footage.

[0009] It is understandable that the beneficial effects of the second aspect mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating an adaptive control method for viewing comfort in a vehicle-mounted drone tracking system according to an embodiment of this application. Figure 2 This is a schematic diagram of a UAV-vehicle tracking geometric relationship model provided in one embodiment of this application; Figure 3 This is a schematic diagram of the viewpoint comfort adaptive control system of a vehicle-mounted drone tracking system provided in one embodiment of this application; Figure 4 This is a schematic diagram of the viewpoint comfort adaptive control system of a vehicle-mounted drone tracking system provided in another embodiment of this application; Figure 5 This is a schematic diagram illustrating the working principle of the adaptive control system for the viewing comfort of a vehicle-mounted drone tracking system provided in one embodiment of this application. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0014] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0015] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0016] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0017] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0018] The adaptive viewpoint comfort control method provided in this invention can be applied to various vehicles equipped with vehicle-mounted drone systems, enabling real-time optimization and adjustment of the shooting perspective during drone tracking. The vehicle can be a passenger car, SUV, commercial vehicle, or other ground-based vehicle; the drone can be a multi-rotor drone, fixed-wing drone, or other aircraft with hovering and trajectory tracking capabilities. This invention does not impose any restrictions on the specific types of vehicles and drones. This method is particularly suitable for scenarios requiring long-duration, high-quality visual perception, such as road exploration in wilderness adventures, tracking and shooting in dashcams, and creative aerial photography.

[0019] like Figure 1 As shown, in one embodiment, the viewpoint comfort adaptive control method of the vehicle-mounted drone tracking system provided by this invention includes the following steps S101 to S105: Step S101: Obtain system status data; wherein the system status data includes at least one of vehicle status information, UAV status information, and visual feature information of the vehicle in the UAV camera image.

[0020] In applications, system status data serves as the information foundation for driving subsequent comfort assessments and intelligent decision-making. Vehicle status information refers to data reflecting the vehicle's current motion state, including but not limited to longitudinal speed, lateral speed, and direction of travel. This information can be obtained through the vehicle's own GPS / INS integrated navigation system. UAV status information refers to data reflecting the UAV's current flight attitude and position, including but not limited to the UAV's flight altitude H relative to the vehicle, horizontal following distance D, gimbal pitch angle φ, and relative position offset between the vehicle and the UAV. This information can be obtained through the UAV's onboard multi-sensor fusion positioning module, which integrates GPS, RTK (Real-Time Kinematics), IMU (Inertial Measurement Unit), visual odometry, and other positioning methods to provide high-precision vehicle-to-UAV relative pose. Visual feature information refers to feature data obtained through image recognition and analysis of real-time images captured by the UAV's onboard camera, including but not limited to the top-down angle α and the vehicle's area proportion in the image. This includes information such as the vehicle's image center coordinates. The acquisition of all this information is achieved through real-time transmission via a high-speed, low-latency communication link between the vehicle and the drone.

[0021] Step S102: Based on the system state data, calculate the quantitative value representing the comfort of the current following angle according to the pre-established geometric relationship model.

[0022] In applications, the geometric relationship model is a pre-established mathematical model based on the spatial geometric relationship between the drone and the vehicle, such as... Figure 2 As shown, this model describes the mapping relationship between flight altitude H, horizontal following distance D, camera field of view, and the vehicle's imaging effect in the frame. The quantified cost value is a scalar value used to quantitatively measure the degree to which the current following view deviates from the ideal viewpoint; the lower the value, the closer the current viewpoint is to the ideal state, i.e., the more comfortable it is. The calculation of this quantified cost value comprehensively considers multiple sub-objectives such as image composition, observation angle, and energy consumption control, and integrates multi-dimensional evaluation indicators into a unified cost function through weighted summation. The specific geometric relationship model and quantified cost value calculation formula will be elaborated in detail in subsequent embodiments.

[0023] Step S103: Input the system state data and the quantized cost into the pre-trained reinforcement learning control model, and have the reinforcement learning control model output control commands including the UAV flight position adjustment amount and the gimbal angle adjustment amount.

[0024] In application, reinforcement learning control models are intelligent control models pre-trained using reinforcement learning algorithms. Their core capability lies in directly outputting the optimal control actions based on the current system state and comfort assessment results. Specifically, flight position adjustment refers to the change in flight altitude ΔH and the change in horizontal following distance ΔD that the UAV needs to perform in the next control cycle; gimbal angle adjustment refers to the change in pitch angle that the gimbal needs to perform in the next control cycle. Through extensive training in a simulation environment, this model has learned strategies for adjusting the drone's position and gimbal angle to achieve optimal viewing comfort in various driving scenarios. Deployed on an in-vehicle control terminal or the drone's onboard computer, the model can complete inference calculations at speeds that meet real-time requirements.

[0025] Step S104: According to the control command, coordinate the drone to adjust its flight position and drive the gimbal to adjust the shooting angle.

[0026] In application, coordinated operation is a key feature of this invention, meaning that the adjustment of the UAV's flight position and the adjustment of the gimbal's shooting angle are executed synchronously within the same control cycle, rather than being adjusted independently. Specifically, the UAV flight control system adjusts the UAV's flight altitude and horizontal position according to the flight altitude adjustment amount ΔH and the horizontal following distance adjustment amount ΔD in the control commands, while the gimbal control system adjusts the gimbal pitch angle according to the control commands. Adjust the camera's shooting angle. Through the coordinated control of the flight platform and the gimbal, multi-degree-of-freedom joint optimization can be achieved, overcoming the limitation that it is difficult to take into account multiple perspective evaluation indicators when only adjusting a single degree of freedom.

[0027] Step S105: Repeat the steps of calculating the quantitative cost value representing the comfort of the current following angle to the steps of coordinating the drone to adjust its flight position and driving the gimbal to adjust the shooting angle, so as to form a closed-loop control.

[0028] In application, closed-loop control refers to the iterative execution of steps S102 to S104. Within each control cycle, the system reacquires the latest state data, recalculates the current quantization cost, re-outputs control commands through the reinforcement learning control model, and re-executes the cooperative driving action. This cyclical process runs continuously at a preset control frequency, for example, 4Hz, meaning a complete "evaluation-decision-execution" cycle is performed every 0.25 seconds. Through this continuously iterative closed-loop mechanism, the system can track dynamic changes in the driving environment (such as vehicle acceleration / deceleration, turning, and incline / exit), continuously adjusting the tracking angle to maintain a comfortable viewing experience.

[0029] The adaptive control method for viewing comfort provided in this invention establishes a geometric relationship model to quantify viewing comfort into a calculable cost. Combined with a reinforcement learning control model, it directly outputs coordinated control commands for flight position and gimbal angle in a continuous action space. Through closed-loop control and continuous optimization, it can autonomously maintain the best following angle in a dynamic driving environment, effectively improving the visual comfort and stability of the vehicle-mounted drone's following footage.

[0030] In one embodiment, calculating the quantitative cost representing the comfort of the current following camera angle based on a pre-established geometric relationship model includes: The downward angle α is calculated based on the flight altitude H and the horizontal following distance D. Based on the flight altitude H, the horizontal following distance D, the camera field of view, and the actual size of the vehicle, calculate the area ratio of the vehicle in the image. ; Based on the top angle α and the area ratio The deviations from their respective target values, combined with energy consumption control, are weighted and summed to obtain the quantified cost value.

[0031] In applications, such as Figure 2 As shown, the downward angle α is the angle between the line connecting the drone and the vehicle and the horizontal plane. This angle determines the "sense of height" in the image. If the downward angle is too small, the image is close to eye level, the road ahead is obscured by the vehicle body, and the ground texture is severely distorted; if the downward angle is too large, the image is close to looking down, and the sense of depth of the road ahead is significantly reduced. The target value for a comfortable downward angle is usually set between 30° and 50°. Area ratio This is the ratio of the pixel area occupied by the vehicle in the image to the total pixel area of ​​the image. This metric determines the visual size of the vehicle in the image. When the area ratio is too small, it is difficult for the driver to judge the relative position of the vehicle to the road ahead; when the area ratio is too large, the vehicle occupies too much of the image, compressing the visible area of ​​the road ahead. A comfortable area ratio target range is typically 15% to 25%. Control energy consumption refers to the sum of the squares of the drone's flight position adjustments, used to constrain the amplitude of control actions, encouraging smooth adjustments and avoiding drastic positional changes by the drone. Weighted summation refers to multiplying each of the above sub-objectives by its corresponding weight coefficient and then adding them together to obtain a comprehensive quantitative cost.

[0032] By using the above methods, subjective visual comfort is transformed into a quantifiable mathematical indicator, providing a clear optimization target for subsequent reinforcement learning control.

[0033] In one embodiment, the top-down view The calculation formula is: Formula (1) The area ratio Vehicle imaging height ratio Vehicle imaging width ratio The product is obtained, where: Formula (2) Formula (3) in, and These are the vehicle's actual height and width, respectively. The slant distance between the drone and the vehicle and , and These are the camera's vertical and horizontal field of view, respectively.

[0034] In application, formula (1) calculates the angle between the line connecting the UAV and the vehicle and the horizontal plane using the arctangent function, based on the vertical distance H and the horizontal distance D between the UAV and the vehicle. This angle is the downward angle α. When the flight altitude H increases or the horizontal distance D decreases, the downward angle α increases, and the image tends to have a top-down effect.

[0035] Formulas (2) and (3) are derived approximately from the pinhole camera imaging model and are used to estimate the proportion of the vehicle's image size in the image. The vehicle is simplified as a figure with a height... and width The rectangular plane. The imaging height ratio calculated by formula (2) This reflects the proportion of the vehicle's height in the image. The cosα term in the numerator represents the projected shortening effect of the vehicle's height due to the downward angle. The denominator... Indicates the current slope distance Lower camera vertical field of view The corresponding actual height range of the image. The imaging width ratio calculated by formula (3) This reflects the proportion of the vehicle's width in the frame, with the additional cosα term in the denominator representing the change in the horizontal imaging range due to the downward angle. The above formula holds true under the conditions that the vehicle is located in the center of the camera's field of view and perspective distortion is negligible (i.e., the vehicle occupies no more than 30% of the frame area). In actual vehicle-mounted tracking scenarios, typical vehicle dimensions are 1.5 to 1.8 meters in height and 1.8 to 2.0 meters in width, with a tracking distance L ranging from 15 to 50 meters, all of which satisfy the above approximate conditions.

[0036] The above formula system allows for the precise calculation of quantitative indicators characterizing the composition effect of the image from the relative pose parameters of the drone and the vehicle, providing a mathematical basis for the numerical evaluation of viewing comfort.

[0037] In one embodiment, the formula for calculating the quantized cost value by weighted summation is as follows: Formula (4) in, For the quantified cost value, , , These are the weighting coefficients for each sub-objective. For the target area percentage, From the perspective of the target, For flight altitude adjustment, This refers to the horizontal following distance adjustment amount; the quantified cost... The lower the value, the more comfortable the current viewing angle.

[0038] In application, formula (4) is a joint optimization cost function under multi-objective constraints. This formula contains three sub-objective terms: To mitigate the deviation of the image from the sub-target, the deviation between the current vehicle's image area ratio and the target value is measured, and a squared form is used, with a larger deviation resulting in a heavier penalty; (α-α_target) 2 To measure the deviation of the viewpoint from the target value, we need to consider the deviation between the current overhead viewpoint and the target value. To control the energy consumption sub-objective, the magnitude of flight position adjustment maneuvers is measured to encourage smooth adjustments and avoid abrupt changes in the UAV's position. Each of the three sub-objectives is multiplied by its corresponding weighting coefficient. , , The values ​​are then added together to obtain the comprehensive quantitative value S. The typical range of values ​​for the weighting coefficients is as follows: =1.0 (the composition ratio is the primary optimization target). =0.8 (perspective is secondary target). =0.3 (Energy consumption is a secondary indicator with a low weight). In actual deployment, the value can be fine-tuned within ±30% of the above value according to the specific scenario.

[0039] By designing the aforementioned cost function, the multi-dimensional evaluation of viewpoint comfort is unified into a single scalar cost value, enabling the subsequent reinforcement learning control model to be trained and make decisions with minimizing this cost value as the optimization objective.

[0040] In one embodiment, the reward function of the reinforcement learning control model The calculation formula is: Formula (5) in, and The penalty coefficient is... The gimbal pitch angle adjustment amount; the quantified cost value in the reward function. Taking a negative value makes maximizing the reward function equivalent to minimizing the quantized cost value.

[0041] In applications, the reward function is a core design element of reinforcement learning algorithms, serving to provide evaluation signals for each step of the control model's decision-making. In formula (5), due to the cost... The cost function should be as small as possible, while the training objective of reinforcement learning is to maximize the cumulative reward. Therefore, the reward function should have a smaller value for the cumulative reward. Take the negative value (i.e. -) This minimizes the cost. This is equivalent to maximizing the reward R. Compared to the control energy consumption term already included in formula (4), formula (5) additionally introduces the adjustment amount for the gimbal pitch angle. The squared penalty term This is used to constrain the range of gimbal angle adjustments, preventing severe gimbal jitter. Typical values ​​for the penalty coefficient are: =0.1 (position adjustment penalty) =0.05 (Angle adjustment penalty, with a low weight, because gimbal adjustment consumes far less energy than flight position adjustment).

[0042] Through the design of the above reward function, the reinforcement learning control model not only pursues the optimal visual comfort during the training process, but also takes into account the smoothness of the control actions, avoiding the negative impact of too frequent or drastic adjustments on flight stability and image smoothness.

[0043] In one embodiment, the reinforcement learning control model is trained using a deep deterministic policy gradient algorithm, and the reinforcement learning control model includes: A policy network is used to output action vectors based on the input state vectors. The action vectors include flight altitude adjustment, horizontal following distance adjustment, and gimbal pitch angle adjustment. A value network is used to evaluate the Q-value of the current state-action pair based on the concatenated input of the state vector and the action vector.

[0044] In applications, the Deep Deterministic Policy Gradient (DDPG) algorithm is a deep reinforcement learning algorithm suitable for continuous action spaces. The policy network (also called the actor network) learns a deterministic policy, meaning that given the current state, it directly outputs a definite action, rather than outputting the probability distribution of the action. The input to the policy network is a 12-dimensional state vector, including vehicle speed (2D, longitudinal and lateral speeds), drone state (3D, flight altitude H, horizontal following distance D, gimbal pitch angle φ), vehicle-drone relative position (3D, relative x / y / z offset), and visual features (4D, top-down angle α, area percentage). Vehicle image center coordinates u c and v c All state variables are normalized to the range [0,1]. The output of the policy network is a 3-dimensional action vector, including flight altitude adjustment ΔH (range [-2.0m, +2.0m] / step), horizontal distance adjustment ΔD (range [-3.0m, +3.0m] / step), and gimbal pitch angle adjustment Δφ (range [-5°, +5°] / step). The structure of the policy network is: input layer (12-dimensional) → hidden layer 1 (256 neurons, ReLU activation) → hidden layer 2 (128 neurons, ReLU activation) → output layer (3-dimensional, scaled to the action space range after tanh activation).

[0045] The role of the value network (also known as the Critic network) is to evaluate the long-term expected reward, or Q-value, of performing a specific action in a given state. The network's input is a concatenation of a 12-dimensional state vector and a 3-dimensional action vector (15 dimensions in total), passed through hidden layer 1 (256 neurons, ReLU activation) → hidden layer 2 (128 neurons, ReLU activation) → output layer (1-dimensional Q-value). During training, the policy network updates its parameters based on the Q-value gradient of the value network, ensuring that the actions output by the policy network achieve higher Q-values; the value network, in turn, updates its parameters based on the actual reward signal received, making the Q-value estimation more accurate. The two networks cooperate and alternately optimize, ultimately converging to the optimal policy.

[0046] By employing the DDPG algorithm, the control model can directly output precise control quantities in a continuous motion space, avoiding the accuracy loss caused by discretizing the motion space. This makes it suitable for tracking shooting scenarios that require fine adjustment of the drone's position and gimbal angle.

[0047] In one embodiment, the training process of the deep deterministic policy gradient algorithm includes: The initial strategy is obtained by offline training in a joint simulation environment based on vehicle dynamics model and UAV flight dynamics model; In actual deployment, an online fine-tuning strategy is adopted to reduce the exploration noise to a preset proportion at the end of offline training, and to simultaneously reduce the learning rate to a preset proportion at the time of offline training.

[0048] In the application, offline training is the model training process conducted in a simulation environment. This co-simulation environment is built based on vehicle dynamics and UAV flight dynamics models, capable of simulating various driving scenarios (straight-line constant speed, curves, acceleration / deceleration, uphill / downhill), and also simulating real-world factors such as wind disturbances and GPS signal fluctuations. Each training round corresponds to a complete tracking task, lasting approximately 60 seconds, with a control frequency of 4Hz (one decision every 0.25 seconds). Key hyperparameters in the training process include: a policy network learning rate of 1×10⁻⁶. -4 The value network learning rate is 1×10⁻⁶. -3 Discount factor γ = 0.99, soft update coefficient τ = 0.005, experience replay buffer capacity 10. 6 The training yielded 256 experience points. Noise was explored using the Ornstein-Uhlenbeck process with an initial standard deviation σ = 0.2 and a decay coefficient of 0.15, linearly decreasing to 0.02 during training. The total training steps were approximately 500,000.

[0049] Online fine-tuning refers to the continuous optimization process of the control model after actual deployment. After obtaining the initial policy through offline training, the controller retains the DDPG experience replay mechanism during actual deployment, but reduces the exploration noise to 10% of that at the end of offline training (i.e., σ). online =0.02), to continuously explore new environmental features with small increments. The learning rate is synchronously reduced to 10% of the offline rate. During fine-tuning, parameter updates are only triggered when the reward for the current round is significantly lower than the average reward during offline training (below 70%), avoiding over-tuning when performance is already good.

[0050] By combining offline training with online fine-tuning, the control model can not only perform well in common scenarios, but also adapt to new scenarios, thus improving the problem of insufficient generalization performance of pure offline models when facing unseen scenarios.

[0051] In one embodiment, before coordinating the adjustment of the flight position of the drone and the adjustment of the shooting angle of the gimbal according to the control command, the method further includes: The control commands are constrained and validated. When the horizontal following distance is satisfied after the control command is executed Less than the preset minimum safe distance Slant distance between the drone and the vehicle Greater than the preset maximum communication distance Flight altitude Height greater than the preset maximum regulatory limit Vehicle imaging height ratio or vehicle image width ratio Exceeding the preset lower limit of proportion and the upper limit of proportion When any of the defined ranges are met, the control command is modified so that the modified control command satisfies the corresponding constraints.

[0052] In applications, constraint verification is a safety assurance step inserted after the reinforcement learning control model outputs control commands and before the actual execution of control actions. Among these, minimum safe distance... This is designed to prevent collisions between drones and vehicles due to excessive proximity; a typical value is 10 meters. Maximum communication distance. Used to ensure the quality of image transmission and control signals between drones and vehicles, typically 100 meters; maximum regulatory-restricted altitude. This refers to the maximum flight altitude limit for drones as determined by airspace management regulations, typically 120 meters; the lower limit is... and the upper limit of proportion To ensure the vehicle is visible in the image without being too large, typical values ​​are 0.1 and 0.5, respectively. When the control command output by the reinforcement learning control model causes any of the above constraints to be violated after execution, the system will prune and correct the corresponding component of the control command, limiting it to the nearest feasible value that satisfies the constraints. This ensures that the corrected control command maintains the optimization direction of the original command as much as possible while meeting all safety and regulatory requirements.

[0053] Through the above constraint verification mechanism, the system is guaranteed to meet the basic requirements of safe distance, communication reliability, regulatory compliance and screen availability while optimizing viewing comfort, thereby improving the security and robustness of the system in actual deployment environments.

[0054] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0055] This invention also provides a perspective comfort adaptive control system for a vehicle-mounted drone tracking system, used to execute the steps in the above-described perspective comfort adaptive control method embodiments.

[0056] In one embodiment, such as Figure 3 As shown, the viewing comfort adaptive control system 100 provided in this embodiment of the invention includes: The state perception unit 101 is used to acquire system state data; wherein, the system state data includes at least one of vehicle state information, UAV state information, and visual feature information of the vehicle in the UAV camera image. The comfort assessment unit 102 is used to calculate the quantitative value representing the comfort of the current following angle based on the system state data and a pre-established geometric relationship model. Decision unit 103 is used to input the system state data and the quantified cost into a pre-trained reinforcement learning control model, and the reinforcement learning control model outputs control commands including UAV flight position adjustment amount and gimbal angle adjustment amount. The collaborative execution unit 104 is used to collaboratively drive the UAV to adjust its flight position and drive the gimbal to adjust its shooting angle according to the control command. The comfort assessment unit, the decision-making unit, and the collaborative execution unit are configured to operate in a loop to form a closed-loop control.

[0057] In application, the aforementioned units can be software program units deployed on the vehicle control terminal or the UAV's onboard computer, or implemented through different logic circuits integrated in a processor, or through the collaborative implementation of multiple distributed processors. The state perception unit collects real-time data from both the vehicle and the UAV via a communication link and aggregates it into system state data in a unified format. The comfort assessment unit executes the calculation logic of the geometric relationship model and outputs quantified cost values. The decision-making unit runs a trained reinforcement learning control model, completes inference calculations, and outputs control commands. The collaborative execution unit distributes control commands to the UAV flight control system and gimbal control system, driving the actuators to move synchronously.

[0058] In one embodiment, such as Figure 4 As shown, the viewpoint comfort adaptive control system further includes: An unmanned aerial vehicle (UAV) flight platform is used to perform flight position adjustments according to the flight position adjustment amount in the control command. A gimbal camera is mounted on the drone flight platform and can adjust its shooting angle independently of the drone flight platform; A multi-sensor fusion positioning module is used to provide relative pose data between the vehicle and the drone, the relative pose data being used to generate at least a portion of the system state data; A communication link is used to transmit the control commands, system status data, and video streams between the vehicle and the drone.

[0059] In application, the drone flight platform is a flight carrier equipped with a gimbal camera and onboard computer, possessing precise hovering and trajectory tracking capabilities, and able to adjust flight altitude and horizontal position in real time according to control commands. The gimbal camera is equipped with a high-definition camera, which achieves independent rotation relative to the flight platform through the gimbal mechanism, thereby independently adjusting the shooting angle while the drone adjusts its flight position. The multi-sensor fusion positioning module integrates multiple positioning sensors such as GPS, RTK, IMU, and visual odometry, providing centimeter-level accuracy vehicle-drone relative pose data through multi-source data fusion algorithms. The communication link adopts a high-speed, low-latency wireless communication method for bidirectional transmission of control commands, status data, and video streams between the vehicle and the drone, ensuring the real-time performance of the system's closed-loop control.

[0060] like Figure 5 As shown, after the system is started, it enters the follow-shooting mode. It acquires the drone's attitude, gimbal attitude, vehicle attitude, and visual features through state perception, and calculates a quantitative value representing the comfort of the current follow-shooting angle based on a pre-established geometric relationship model. This is the comfort assessment, which may specifically include calculating the pitch angle and area ratio, based on the aforementioned top-down angle. and the area ratio The deviations from their respective target values, combined with control energy consumption, are weighted and summed to obtain a comfort score. Then, the system state data and comfort score are input into a pre-trained reinforcement learning control model for intelligent decision-making, outputting control commands containing UAV flight position adjustments and gimbal angle adjustments to the UAV and gimbal.

[0061] This application also provides an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in the above-described method embodiments.

[0062] In applications, the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0063] In applications, memory can be an internal storage unit of an electronic device in some embodiments, such as a hard drive or RAM. In other embodiments, memory can be an external storage device of the electronic device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal and external storage units of the electronic device. Memory is used to store operating systems, applications, bootloaders, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.

[0064] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0065] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0066] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the various method embodiments above.

[0067] This application provides a computer program product, including a computer program, which, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.

[0068] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0069] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0070] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0071] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0072] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0073] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for adaptive control of viewing comfort in a vehicle-mounted drone tracking system, characterized in that, include: Acquire system status data; wherein the system status data includes at least one of vehicle status information, UAV status information, and visual feature information of the vehicle in the UAV camera image; Based on the system status data, a quantitative value representing the comfort of the current following angle is calculated according to a pre-established geometric relationship model. The system state data and the quantified cost are input into a pre-trained reinforcement learning control model, and the reinforcement learning control model outputs control commands that include the UAV flight position adjustment amount and the gimbal angle adjustment amount. According to the control commands, the drone is coordinated to adjust its flight position and the gimbal is driven to adjust the shooting angle. The steps of calculating the quantified cost value representing the comfort of the current following viewpoint are repeated until the steps of coordinating the drone to adjust its flight position and driving the gimbal to adjust the shooting angle are executed to form a closed-loop control.

2. The adaptive control method for viewing comfort as described in claim 1, characterized in that, The calculation of the quantitative cost value representing the comfort of the current following camera angle based on a pre-established geometric relationship model includes: According to flight altitude and horizontal following distance Calculate the top angle ; According to the flight altitude The horizontal following distance Based on the camera's field of view and the vehicle's actual dimensions, calculate the area occupied by the vehicle in the image. ; According to the aforementioned top-down perspective and the area ratio The deviations from their respective target values, combined with energy consumption control, are weighted and summed to obtain the quantified cost value.

3. The adaptive control method for viewing comfort as described in claim 2, characterized in that, The top-down angle The calculation formula is: ; The area ratio Vehicle imaging height ratio Vehicle imaging width ratio The product is obtained, where: ; ; in, and These are the vehicle's actual height and width, respectively. The slant distance between the drone and the vehicle and , and These are the camera's vertical and horizontal field of view, respectively.

4. The adaptive control method for viewing comfort as described in claim 2, characterized in that, The formula for calculating the quantization cost value by weighted summation is as follows: ; in, For the quantified cost value, , , These are the weighting coefficients for each sub-objective. For the target area percentage, From the perspective of the target, For flight altitude adjustment, This refers to the horizontal following distance adjustment amount; the quantified cost... The lower the value, the more comfortable the current viewing angle.

5. The adaptive control method for viewing comfort as described in claim 4, characterized in that, The reward function of the reinforcement learning control model The calculation formula is: ; in, and The penalty coefficient is... This is the adjustment amount for the gimbal's pitch angle.

6. The adaptive control method for viewing comfort as described in claim 5, characterized in that, The reinforcement learning control model is trained using a deep deterministic policy gradient algorithm, and the reinforcement learning control model includes: A policy network is used to output action vectors based on the input state vectors. The action vectors include flight altitude adjustment, horizontal following distance adjustment, and gimbal pitch angle adjustment. A value network is used to evaluate the Q-value of the current state-action pair based on the concatenated input of the state vector and the action vector.

7. The adaptive control method for viewing comfort as described in claim 6, characterized in that, The training process of the deep deterministic policy gradient algorithm includes: The initial strategy is obtained by offline training in a joint simulation environment based on vehicle dynamics model and UAV flight dynamics model; In actual deployment, an online fine-tuning strategy is adopted to reduce the exploration noise to a preset proportion at the end of offline training, and to simultaneously reduce the learning rate to a preset proportion at the time of offline training.

8. The adaptive control method for viewing comfort as described in any one of claims 1 to 7, characterized in that, Before coordinating the adjustment of the drone's flight position and the adjustment of the gimbal's shooting angle according to the control commands, the process also includes: The control commands are constrained and validated. When the horizontal following distance is satisfied after the control command is executed Less than the preset minimum safe distance Slant distance between the drone and the vehicle Greater than the preset maximum communication distance Flight altitude Height greater than the preset maximum regulatory limit Vehicle imaging height ratio or vehicle image width ratio Exceeding the preset lower limit of proportion and the upper limit of proportion When any of the defined ranges are met, the control command is modified so that the modified control command satisfies the corresponding constraints.

9. A field-of-view comfort adaptive control system for a vehicle-mounted drone tracking system, characterized in that, include: A state perception unit is used to acquire system state data; wherein, the system state data includes at least one of vehicle state information, UAV state information, and visual feature information of the vehicle in the UAV camera image; The comfort assessment unit is used to calculate the quantitative value representing the comfort of the current following angle based on the system state data and a pre-established geometric relationship model. The decision unit is used to input the system state data and the quantified value into a pre-trained reinforcement learning control model, and the reinforcement learning control model outputs control commands including the UAV flight position adjustment amount and the gimbal angle adjustment amount. The collaborative execution unit is used to collaboratively drive the UAV to adjust its flight position and drive the gimbal to adjust its shooting angle according to the control commands. The comfort assessment unit, the decision-making unit, and the collaborative execution unit are configured to operate in a loop to form a closed-loop control.

10. The viewpoint comfort adaptive control system as described in claim 9, characterized in that, Also includes: An unmanned aerial vehicle (UAV) flight platform is used to perform flight position adjustments according to the flight position adjustment amount in the control command. A gimbal camera is mounted on the drone flight platform and can adjust its shooting angle independently of the drone flight platform; A multi-sensor fusion positioning module is used to provide relative pose data between the vehicle and the drone, the relative pose data being used to generate at least a portion of the system state data; A communication link is used to transmit the control commands, system status data, and video streams between the vehicle and the drone.