An unmanned aerial vehicle battery intelligent charging strategy method
Patent Information
- Application Number
- CN202611182937.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]然而,该技术方案仍存在以下不足:其一,场景识别模块输出离散的场景类型标识,通过查找表方式映射至充电策略,当无人机进入未定义的场景类型时无法生成适配策略,且离散标签丢失了环境因素连续变化的信息;其二,深度确定性策略梯度算法的策略收敛依赖于大量在线交互试错,探索成本高,无法满足充电任务的即时性要求;其三,缺乏充电策略执行前的事前评估能力,只能事后评价,无法在策略执行前预判其效果
第一,通过因果图注意力网络将环境监测数据映射为连续向量空间的因果表征,保留了各环境因素对充电效果的差异化因果影响信息,有助于缓解离散场景标签与连续策略参数之间的表征鸿沟问题,在一定程度上增强了对未知场景的泛化推理能力。
Smart Images

Figure CN122808542A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent charging technology for unmanned aerial vehicles (UAVs), and in particular to an intelligent charging strategy and method for UAV batteries. Background Technology
[0002] With the widespread application of drones in fields such as power grid inspection and logistics delivery, drone battery charging management has become a crucial aspect of ensuring their continuous and stable operation. The environmental and power supply conditions vary significantly across different application scenarios, necessitating differentiated adaptation requirements for battery charging strategies.
[0003] In the prior art, Chinese invention patent application CN119527602A proposes a smart charging control method for drones based on multi-scenario adaptation. Its core solution is as follows: environmental monitoring data is collected through multiple types of sensors, and scene recognition and classification are performed using a deep learning model to generate scene feature vectors and scene type identification information; based on the scene type identification, a basic charging strategy is retrieved from a preset charging strategy library, and the strategy is optimized using a deep deterministic strategy gradient algorithm to generate charging control parameters; during the charging process, battery status parameters are monitored in real time, the battery health index is calculated and fed back to the optimization model to achieve closed-loop control.
[0004] However, this technical solution still has the following shortcomings: First, the scene recognition module outputs discrete scene type identifiers, which are mapped to the charging strategy through a lookup table. When the drone enters an undefined scene type, it cannot generate an appropriate strategy, and the discrete labels lose information about the continuous changes in environmental factors. Second, the strategy convergence of the deep deterministic strategy gradient algorithm depends on a large number of online interactive trial and error, which has high exploration costs and cannot meet the real-time requirements of the charging task. Third, it lacks the ability to evaluate the charging strategy before execution and can only evaluate it afterward, making it impossible to predict its effect before the strategy is executed. Summary of the Invention
[0005] In view of the above-mentioned prior art, the present invention provides a smart charging strategy method for drone batteries, which mainly solves the technical problems existing in the background art.
[0006] To achieve the above objectives, the technical solution of this invention is implemented as follows: This invention provides a smart charging strategy method for drone batteries, the method comprising the following steps: S1: Obtain environmental monitoring data, battery status data, and scene data of historical charging scenarios; construct a structural causal graph containing environmental variable nodes, battery status nodes, and charging effect nodes; and use a causal graph attention network to extract the causal effect strength of each environmental variable node on the charging effect node, and generate a causal representation vector based on the causal effect strength. S2: Construct a policy network. Based on the causal representation vector and the scene data of historical charging scenarios, use a model-independent meta-learning algorithm to meta-train the policy network to obtain meta-policy parameters. Update the meta-policy parameters with gradients based on the running data of the new scenario to obtain specific charging policy parameters. Generate multiple candidate charging policy parameters based on the specific charging policy parameters. S3: Construct a charging process proxy model, input the parameters of each candidate charging strategy into the charging process proxy model for forward deduction, obtain the expected charging effect corresponding to each candidate charging strategy parameter, and select the candidate strategy that satisfies the safety constraints and has the best effect as the charging strategy to be executed based on the expected charging effect. S4: Execute the charging strategy to be executed, monitor the battery status parameters in real time during the charging process and compare them with the preset safety threshold to obtain the actual charging effect, calculate the prediction deviation between the actual charging effect and the expected charging effect, and update the charging process proxy model according to the prediction deviation. S5: Correct the causal effect strength based on the prediction bias, and update the policy network with the corrected causal effect strength.
[0007] As a preferred embodiment of the present invention, in step S1, the specific process of extracting the causal effect strength of each environmental variable node on the charging effect node using a causal graph attention network, and generating a causal representation vector based on the causal effect strength includes: The environmental monitoring data or battery status data corresponding to each node in the aforementioned causal graph are used as the feature vectors of each node. The feature vectors of each node are input into the causal graph attention network. The causal graph attention network uses the directed edges in the causal graph as constraints for attention calculation and uses the attention weights of each environmental variable node to the charging effect node as the causal effect strength of each environmental variable node to each charging effect node. The causal effect strength of each environmental variable node on each charging effect node is arranged in a preset order and then vector-concatenated with the current measurement value of each environmental variable to obtain the causal representation vector.
[0008] As a preferred embodiment of the present invention, in step S2, the specific process of performing meta-training on the policy network using a model-independent meta-learning algorithm based on the causal representation vector and scene data of historical charging scenarios to obtain meta-policy parameters includes: The policy network is a multi-layer fully connected neural network consisting of an input layer, multiple hidden layers and an output layer. The input layer is used to receive the causal representation vector. Each hidden layer sequentially performs a nonlinear transformation on the output of the previous layer. The output layer maps the output of the last hidden layer to a charging current setting value and a charging voltage setting value. Acquire scene data from multiple historical charging scenarios, and treat each historical charging scenario as a separate meta-training task. Based on the current meta-training task, starting from the current meta-policy parameters, the first loss function is calculated using the training set, and gradient descent is performed on the gradient of the current meta-policy parameters according to the first loss function to obtain the corresponding task-specific parameters. After traversing all meta-training tasks, the sum of the second loss functions of each task-specific parameter on the corresponding task's validation set is calculated. The current meta-policy parameters are updated according to the gradient of the sum of the second loss functions, and the meta-policy parameters at the end of the update are used as the meta-policy parameters after training is completed.
[0009] As a preferred embodiment of the present invention, in step S2, the meta-strategy parameters are updated using gradients based on the operating data of the new scenario to obtain specific charging strategy parameters. The specific process of generating multiple candidate charging strategy parameters based on the specific charging strategy parameters includes: Obtain operational data under the new scenario, wherein the operational data includes multiple sets of mapping samples consisting of the causal representation vector and the actual charging strategy parameters adopted; Starting from the meta-policy parameters, the causal representation vector in each of the mapping samples is used as the input of the policy network, and the charging policy parameters in each of the mapping samples are used as the supervision target to calculate the third loss function of the policy network on the running data. The meta-policy parameters are updated iteratively according to the third loss function, and the updated parameters are used as the dedicated charging policy parameters. Using the charging current setpoint and the charging voltage setpoint as the mean of the first Gaussian distribution and the second Gaussian distribution, respectively, multiple candidate charging current values are sampled from the first Gaussian distribution, and multiple candidate charging voltage values are sampled from the second Gaussian distribution. The candidate charging current values and the candidate charging voltage values are combined to generate multiple candidate charging strategy parameters.
[0010] As a preferred embodiment of the present invention, the specific process of constructing the charging process proxy model in step S3 includes: Acquire scenario data of historical charging scenarios, which includes multiple sets of training samples consisting of initial environmental conditions, initial battery state, charging strategy parameters adopted, and corresponding actual charging effects. Using the initial environmental conditions, the initial state of the battery, and the charging strategy parameters as inputs, and the actual charging effect as the supervision target, the charging process proxy model is obtained.
[0011] As a preferred embodiment of the present invention, step S3, specifically the process of inputting each of the candidate charging strategy parameters into the charging process proxy model for counterfactual inference to obtain the expected charging effect corresponding to each candidate charging strategy parameter, includes: Obtain the initial environmental conditions and battery initial state in the current scenario, and combine the initial environmental conditions and battery initial state with the parameters of each candidate charging strategy to generate the input vector corresponding to each candidate strategy. Each input vector is input into the charging process proxy model for forward inference, and the expected charging effect corresponding to each candidate strategy is output. The expected charging effect includes the expected charging efficiency, the expected battery temperature rise rate, and the expected capacity decay rate.
[0012] As a preferred embodiment of the present invention, the specific process of selecting the candidate strategy that satisfies the safety constraints and has the best effect as the charging strategy to be executed in step S3 includes: A comprehensive score for each candidate strategy is calculated based on the expected charging effect corresponding to each candidate strategy. The comprehensive score is positively correlated with the expected charging efficiency and negatively correlated with the expected battery temperature rise rate and the expected capacity decay rate. With the expected battery temperature rise rate not exceeding a first threshold and the expected capacity decay rate not exceeding a second threshold as safety constraints, candidate strategies that meet the safety constraints are selected from all candidate strategies. The candidate strategy with the highest comprehensive score that meets the security constraints is selected as the charging strategy to be executed.
[0013] As a preferred embodiment of the present invention, step S4, which involves executing the charging strategy to be executed and monitoring battery state parameters during the charging process and comparing them with a preset safety threshold to obtain the actual charging effect, includes the following specific steps: The charging strategy to be executed is performed. During the charging process, the battery terminal voltage, charging current and battery temperature are collected at a preset sampling frequency as the battery state parameters. The battery state parameters are compared with the corresponding preset safety thresholds in real time. When any of the battery state parameters exceeds the corresponding preset safety threshold for a duration exceeding a first preset time, the charging current is reduced; When the duration exceeds the second preset duration, the charging process is terminated, where the second preset duration is greater than the first preset duration. After charging is completed, the actual charging efficiency, actual average temperature rise rate, and actual equivalent capacity decay rate of this charge are obtained and used as the actual charging effect.
[0014] As a preferred embodiment of the present invention, the specific process of calculating the prediction deviation between the actual charging effect and the expected charging effect in step S4, and updating the charging process proxy model according to the prediction deviation, includes: Obtain the expected charging effect corresponding to the charging strategy to be executed, calculate the difference between the actual charging effect and the expected charging effect in each corresponding dimension, and use the difference in each dimension as the prediction deviation. The initial environmental conditions, initial battery state, the charging strategy to be executed, and the actual charging effect are combined into new training samples, and the charging process proxy model is incrementally updated based on the new training samples.
[0015] As a preferred embodiment of the present invention, step S5 specifically includes: The prediction bias is obtained, and a correction loss function is constructed using the prediction bias as a supervision signal. The gradient of the correction loss function with respect to the causal effect strength parameter in the causal graph attention network is calculated. The causal effect strength parameter is updated once according to the gradient along the gradient descent direction, and the updated causal effect strength parameter is used as the corrected causal effect strength. The corrected causal effect strength is used as the output of the causal graph attention network in the next inference to generate the updated causal representation vector. The updated causal representation vector is input into the policy network, which performs forward propagation with the updated causal representation vector as input and outputs the updated charging policy parameters.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: First, by mapping environmental monitoring data into causal representations in a continuous vector space through a causal graph attention network, the differentiated causal impact information of various environmental factors on charging effect is preserved. This helps to alleviate the representation gap between discrete scene labels and continuous policy parameters and, to some extent, enhances the ability to generalize reasoning about unknown scenarios.
[0017] Second, by using the model-independent meta-learning algorithm, policy adaptation in new scenarios can be completed with only a small number of samples and a finite number of gradient updates. This helps to reduce the reliance of policy optimization on a large number of online interactive trial and error, thereby reducing exploration costs to a certain extent and improving the policy response speed in new scenarios.
[0018] Third, by using the charging process proxy model and fact-based deduction mechanism, the expected effects of different candidate strategies can be deduced and compared before the charging strategy is executed. This helps to provide a predictive reference for strategy selection and reduces the potential risks caused by direct execution of inappropriate strategies to a certain extent.
[0019] Fourth, by feeding the prediction bias back to the causal graph attention network to correct the causal effect strength, and then passing the corrected causal effect strength to the policy network, the actual observation information of the charging effect can be transmitted to the perception and representation stage, thus realizing the rapid scene adaptation and end-to-end closed-loop optimization of the drone battery charging strategy. Attached Figure Description
[0020] Figure 1 A flowchart illustrating the steps of a smart charging strategy for drone batteries. Detailed Implementation
[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. In the following description, the expression "some embodiments" refers to a subset of all possible embodiments; however, it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0022] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without one or more of these details. In other instances, certain technical features well-known in the art have not been described in order to avoid obscuring the invention.
[0023] It should be understood that the present invention can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the invention to those skilled in the art. Furthermore, the terminology used herein is intended only to describe particular embodiments and is not intended to limit the invention. When used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “compose” and / or “comprising,” when used in this specification, identify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups. When used herein, the term “and / or” includes any and all combinations of the associated listed items.
[0024] It should also be noted that when an element is referred to as being "fixed to" another element, it can be directly attached to the other element or there may be an intervening element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or there may be an intervening element. The terms "vertical," "horizontal," "inner," "outer," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only possible implementation.
[0025] To fully understand this invention, a detailed structure will be presented in the following description to illustrate the technical solution proposed by this invention. Optional embodiments of the invention are described in detail below; however, in addition to these detailed descriptions, the invention may have other embodiments.
[0026] Please refer to the attached document. Figure 1 This application provides a smart charging strategy method for drone batteries, the method comprising the following steps: S1: Obtain environmental monitoring data, battery status data, and scene data of historical charging scenarios, construct a structural causal graph containing environmental variable nodes, battery status nodes, and charging effect nodes, and use a causal graph attention network to extract the causal effect strength of each environmental variable node on the charging effect node, and generate a causal representation vector based on the causal effect strength.
[0027] In this embodiment, the environmental variable nodes include temperature, humidity, air pressure, light irradiance, and power supply voltage; the battery status nodes include state of charge, battery temperature, and health status; and the charging effect nodes include charging efficiency, battery temperature rise rate, and capacity decay rate.
[0028] Furthermore, the environmental variable node, the battery state node, and the charging effect node are used as the node set of the structural causal graph; based on the physical influence relationship of the environmental variable node on the battery state node and the causal relationship of the environmental variable node and the battery state node on the charging effect node, the directed edges between the nodes in the node set are determined, and the node set and the directed edges constitute the structural causal graph.
[0029] Furthermore, the causal graph attention network includes two graph attention layers, each containing four parallel attention heads, and each attention head has an embedding dimension of 64. The first layer takes the feature vectors of each node as input, and outputs the intermediate embedding representation of each node after attention aggregation; the second layer takes the output of the first layer as input and outputs the attention weight matrix of each environmental variable node to the charging effect node.
[0030] Specifically, the feature vectors of each node are constructed as follows: the feature vector of an environmental variable node is the measurement value of the environmental sensor corresponding to that environmental variable; that is, the feature vector of a temperature node is the current measurement value of the temperature sensor. The feature vector of the humidity node is the current measurement value of the humidity sensor. Real-time measurement value of air pressure sensor Real-time measurement value of light sensor Real-time measured value of voltage transformer And so on. The feature vector of the battery state node is the battery state parameter output by the battery management system, that is, the feature vector of the state of charge node is the state of charge (SOC), and the feature vector of the battery temperature node is the battery temperature. The feature vector of the health status node is the health status (SOH). The feature vector of the charging effect node is the corresponding actual charging effect index value in the historical charging data.
[0031] Specifically, using the directed edges in the aforementioned structural causal graph as constraints: only when the node To the node When there is a directed edge between nodes Only the node Attention calculations contribute to ensuring that the network does not learn spurious correlations that violate the laws of physics.
[0032] For the existence of directed edges Node pairs, attention heads compute nodes For nodes Original attention coefficient :
[0033] in, and They are nodes and eigenvectors, For attention head The linear transformation weight matrix, For attention head Attention parameter vector, This represents a vector concatenation operation. This is the activation function.
[0034] Then, the original attention coefficients are normalized using Softmax to obtain the node. For all its neighboring nodes Normalized attention weights:
[0035] in, Indicate attention head Next node For his specific neighbors The attention weights obtained after Softmax normalization Represents nodes There exists a directed edge pointing to the node. The set consisting of all neighboring nodes; For summation index variables, representing Any neighboring node in the array, Indicate attention head Next node For any of its neighboring nodes The original attention coefficient.
[0036] Furthermore, the environmental monitoring data and battery status data at the current moment are input into the trained causal graph attention network to extract the attention weight matrix from the second layer output. The attention weight matrix The dimension is 5×3.
[0037] attention weight matrix Expanding the vector row-wise into a 15-dimensional vector and concatenating it with the measured values of the current 5 environmental variables, we obtain a causal representation vector with a dimension of 20. :
[0038] in, This indicates that the attention weight matrix The operation of expanding by row into a 15-dimensional column vector.
[0039] S2: Construct a policy network. Based on the causal representation vector and the scene data of historical charging scenarios, use a model-independent meta-learning algorithm to meta-train the policy network to obtain meta-policy parameters. Update the meta-policy parameters with gradients based on the running data of the new scenario to obtain specific charging policy parameters. Generate multiple candidate charging policy parameters based on the specific charging policy parameters.
[0040] In this embodiment, the policy network is a multi-layer fully connected neural network consisting of an input layer, multiple hidden layers, and an output layer. Specifically, the input layer has a dimension of 20 (corresponding to the dimension of the causal representation vector), the first hidden layer contains 128 neurons, the second hidden layer contains 128 neurons, and the third hidden layer contains 64 neurons, all using the ReLU activation function. The output layer contains 2 neurons, corresponding to the charging current setting value and the charging voltage setting value, respectively. The activation function of the output layer uses the tanh function, linearly mapped to an effective range. The effective range of the charging current setting value is 0.5A to 10A, and the effective range of the charging voltage setting value is 12V to 30V.
[0041] Furthermore, scenario data from multiple historical charging scenarios are acquired, and each historical charging scenario is treated as a separate meta-training task. Each meta-training task includes a training set and a validation set, both of which consist of mapping samples formed by the causal representation vector under that historical charging scenario and the parameters of the actual charging strategy adopted.
[0042] In this embodiment, 1000 meta-training tasks are constructed from a historical database, covering various scene types such as urban areas, suburbs, industrial parks, and wilderness, as well as environmental conditions at different seasons and times. The training set for each task contains 50 samples, and the validation set contains 20 samples.
[0043] Furthermore, a model-independent meta-learning algorithm is used to meta-train the policy network. The goal of the meta-training is to find a set of meta-policy parameters that, starting from these meta-policy parameters, enable the policy network to achieve better performance in handling new tasks after one or more gradient updates.
[0044] For the current meta-training task, from the current meta-policy parameters Let's start by using the training set for this task. Calculate the first loss function :
[0045] in, Indicates the current meta-policy parameters The policy network, The causal representation vector in the training samples. These are the parameters of the charging strategy actually used in the training samples. This is the mean squared error loss function.
[0046] Based on the gradient of the current meta-policy parameters using the first loss function, perform a gradient descent update to obtain the task-specific charging policy parameters. :
[0047] in, , For the inner loop learning rate, .
[0048] After traversing all meta-training tasks, calculate the charging strategy parameters specific to each task. In the corresponding task verification set The sum of the second loss functions on:
[0049] in, Indicates the charging strategy parameters specific to the task. Policy network with parameters For the meta-loss function, The total number of training tasks. .
[0050] Update the current meta-policy parameters based on the gradient of the sum of the second loss functions:
[0051] in, The learning rate for the outer loop. .
[0052] Repeat the above process of traversing all meta-training tasks until a preset termination condition is reached. In this embodiment, the preset termination condition is 500 meta-iterations or the validation set loss does not decrease for 10 consecutive rounds. The meta-policy parameters at the point where the termination condition is reached are used as the meta-policy parameters after training is complete.
[0053] Furthermore, when the drone enters a new scenario, it acquires operational data for that new scenario. This operational data includes multiple sets of mapping samples consisting of causal representation vectors and the parameters of the actual charging strategy used.
[0054] Specifically, from the meta-policy parameters after training Starting from this point, the causal representation vectors in each mapped sample are used as input to the policy network, and the charging policy parameters in each mapped sample are used as supervision targets. The third loss function of the policy network on the running data is then calculated. In this embodiment, the mathematical expression of the third loss function is consistent with the principle of the first loss function.
[0055] The gradient of the meta-policy parameters is iteratively updated using gradient descent based on the gradient of the third loss function. In this embodiment, starting from the meta-policy parameters, the learning rate is used to update the gradient. Perform two gradient descent iterations, and use the updated parameters as the parameters for the dedicated charging strategy. .
[0056] Furthermore, the charging current setting value in the dedicated charging strategy parameters. The mean of the first Gaussian distribution is the charging voltage setting value in the dedicated charging strategy parameters. The standard deviation of the first Gaussian distribution is preset to be the mean of the second Gaussian distribution. The standard deviation of the second Gaussian distribution is preset to be . Wherein, the first Gaussian distribution refers to the continuous probability distribution used to generate candidate charging current values, and the second Gaussian distribution refers to the continuous probability distribution used to generate candidate charging voltage values.
[0057] Multiple candidate charging current values are sampled from a first Gaussian distribution, and multiple candidate charging voltage values are sampled from a second Gaussian distribution. The candidate charging current values and candidate charging voltage values are combined to generate multiple candidate charging strategy parameters. In this embodiment, the generation... One candidate charging strategy parameter.
[0058] S3: Construct a charging process proxy model, input the parameters of each candidate charging strategy into the charging process proxy model for forward deduction, obtain the expected charging effect corresponding to each candidate charging strategy parameter, and select the candidate strategy that satisfies the safety constraints and has the best effect as the charging strategy to be executed based on the expected charging effect.
[0059] In this embodiment, scenario data from historical charging scenarios are acquired. Initial environmental conditions, initial battery state, and charging strategy parameters are used as inputs, and actual charging performance is used as the supervision target. A supervised learning algorithm is used to train an initial model, and the trained initial model is used as a proxy model for the charging process.
[0060] Furthermore, the initial environmental conditions and initial battery state under the current charging scenario are obtained. These initial environmental conditions and initial battery state are then combined with the parameters of each candidate charging strategy to generate the input vector corresponding to each candidate strategy. :
[0061] in, This indicates the initial ambient temperature in the current charging scenario. This indicates the initial relative humidity of the current charging environment. This indicates the initial atmospheric pressure under the current charging scenario. This indicates the initial irradiance under the current charging scenario. This indicates the initial supply voltage in the current charging scenario. This indicates the initial state of charge of the battery in the current charging scenario. This indicates the initial surface temperature of the battery under the current charging scenario. This indicates the initial health state of the battery under the current charging scenario. This indicates the charging current setting value in the candidate charging strategy. This indicates the charging voltage setting value in the candidate charging strategy.
[0062] Furthermore, given that the proxy model for the charging process is a pre-trained random forest regressor, the random forest regressor consists of multiple decision trees, with 100 decision trees, each decision tree being an independent regression model. The input vectors are... The charging process proxy model is fed into the input vector for forward inference. Specifically, starting from the root node of the decision tree, at each internal node, the value of the corresponding feature in the input vector is compared with the threshold based on the splitting feature index and splitting threshold corresponding to that node. The comparison result determines whether to enter the left or right child node, until a leaf node is reached. This leaf node stores the mean of the output values of all samples falling into that node in the training set. This mean is used as the input vector of the decision tree. A single prediction result; Traverse all 100 decision trees, and output the predicted charging efficiency value independently for each decision tree. Predicted average temperature rise rate and predicted capacity decay rate The arithmetic mean of the predictions from all decision trees is used to obtain the expected charging effect for each candidate strategy:
[0063]
[0064]
[0065] in, Indicates the expected charging efficiency. Indicates the expected average temperature rise rate. This represents the expected equivalent capacity decay rate. For the first A decision tree.
[0066] Furthermore, a comprehensive score for each candidate strategy is calculated based on the expected charging effect of each strategy:
[0067] in, These are the weighting coefficients for each dimension. These correspond to the relative importance of expected charging efficiency, expected average temperature rise rate, and expected capacity decay rate in the overall evaluation, respectively.
[0068] Furthermore, the safety constraints are that the expected battery temperature rise rate does not exceed a first threshold and the expected capacity decay rate does not exceed a second threshold. In this embodiment, the first threshold... The second threshold From all candidate strategies, select those that meet the two safety constraints mentioned above, and then select the one with the highest comprehensive score from among the candidate strategies that meet the safety constraints as the charging strategy to be executed.
[0069] S4: Execute the charging strategy to be executed, monitor the battery status parameters in real time during the charging process and compare them with the preset safety threshold to obtain the actual charging effect, calculate the prediction deviation between the actual charging effect and the expected charging effect, and update the charging process proxy model according to the prediction deviation.
[0070] In this embodiment, during the charging process, the battery's terminal voltage, charging current, and battery temperature are collected at a preset sampling frequency as battery state parameters; these battery state parameters are then compared in real time with corresponding preset safety thresholds. The preset safety thresholds are as follows: upper limit of terminal voltage. Upper limit of charging current Battery temperature limit Battery temperature limit .
[0071] If any one of the terminal voltage, charging current, or battery temperature exceeds the corresponding preset safety threshold for a duration exceeding a first preset duration, the charging current is reduced to a safe value. In this embodiment, the first preset duration is 100 milliseconds, and the charging current is reduced by controlling the charger to reduce the output current to 70% of the current value through a pulse width modulation signal.
[0072] The charging process is terminated when the duration exceeds a second preset duration, which is longer than a first preset duration. In this embodiment, the second preset duration is 500 milliseconds, and the charging is terminated by disconnecting the main contactor of the charging circuit.
[0073] Furthermore, after charging is complete, the actual charging effect of this charging is obtained, as follows: First, actual charging efficiency :
[0074] in, To record the actual amount of electricity charged into the battery using a coulomb counter, The power consumption of the charger from the power supply side is recorded by the power metering chip.
[0075] Second, the actual average temperature rise rate
[0076]
[0077] in, This refers to the battery temperature at the start of charging. Temperature at the time of termination. This represents the total charging time.
[0078] Third, actual equivalent capacity decay rate :
[0079] in, This represents the capacity decay measured before and after this charging process through capacity testing. The rated capacity of the battery. One standard charge-discharge cycle.
[0080] Furthermore, the expected charging effect corresponding to the charging strategy to be executed, including the expected charging efficiency. Expected average temperature rise rate and expected equivalent capacity decay rate .
[0081] Calculate the difference between the actual charging effect and the expected charging effect along each corresponding dimension, and use the difference along each dimension as the prediction bias. The specific formula is as follows:
[0082] The initial environmental conditions for this charging: the initial ambient temperature in the current charging scenario. The initial relative humidity of the current charging scenario Initial atmospheric pressure under current charging scenario Initial irradiance under the current charging scenario Initial power supply voltage under the current charging scenario Initial state of battery: The initial state of charge of the battery under the current charging scenario. The initial surface temperature of the battery under the current charging scenario The initial health status of the battery under the current charging scenario Charging strategy parameters to be executed: charging current setting value in the candidate charging strategy. The charging voltage setting value in the candidate charging strategy and actual charging efficiency Actual average temperature rise rate Actual equivalent capacity attenuation rate Combine them into a new training sample.
[0083] The charging process proxy model is incrementally updated based on the new training sample. The specific method of the incremental update is as follows: a new decision tree is constructed using the new training sample, the new decision tree is added to the existing random forest regressor, and the decision tree with the earliest construction time is removed, so that the total number of decision trees remains unchanged at 100.
[0084] S5: Correct the causal effect strength based on the prediction bias, and update the policy network with the corrected causal effect strength.
[0085] In this embodiment, the prediction bias calculated during acquisition is normalized according to the historical standard deviation of each dimension: ,in, For the first The standard deviation of the charging performance index in historical data.
[0086] Constructing a modified loss function using the normalized prediction bias as the monitoring signal. :
[0087] in, The learnable correction scaling parameter is initialized to [value]. =1.0, which is adaptively adjusted with the number of iterations.
[0088] Calculate the corrected loss function The gradient of the causal effect strength parameter in the causal graph attention network.
[0089] The causal effect strength parameter is updated by correcting it along the gradient descent direction based on the gradient, and the specific formula is as follows:
[0090] in, In this embodiment, to correct the step size, The updated causal effect strength parameter is used as the corrected causal effect strength.
[0091] The corrected causal effect strength parameter As the output of the causal graph attention network in the next iteration, it is used to generate the updated causal representation vector. :
[0092] The updated causal representation vector is input into the policy network. The policy network performs forward propagation with the updated causal representation vector as input and outputs the updated charging policy parameters. The updated charging policy parameters are used for execution in the next charging scenario.
[0093] Through the above mechanism, the actual effect data of each charge is used to correct the causal effect strength and update the strategy network parameters, realizing rapid scenario adaptation and end-to-end closed-loop optimization of the drone battery charging strategy.
[0094] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A smart charging strategy method for drone batteries, characterized in that, The method includes the following steps: S1: Obtain environmental monitoring data, battery status data, and scene data of historical charging scenarios; construct a structural causal graph containing environmental variable nodes, battery status nodes, and charging effect nodes; and use a causal graph attention network to extract the causal effect strength of each environmental variable node on the charging effect node, and generate a causal representation vector based on the causal effect strength. S2: Construct a policy network. Based on the causal representation vector and the scene data of historical charging scenarios, use a model-independent meta-learning algorithm to meta-train the policy network to obtain meta-policy parameters. Update the meta-policy parameters with gradients based on the running data of the new scenario to obtain specific charging policy parameters. Generate multiple candidate charging policy parameters based on the specific charging policy parameters. S3: Construct a charging process proxy model, input the parameters of each candidate charging strategy into the charging process proxy model for forward deduction, obtain the expected charging effect corresponding to each candidate charging strategy parameter, and select the candidate strategy that satisfies the safety constraints and has the best effect as the charging strategy to be executed based on the expected charging effect. S4: Execute the charging strategy to be executed, monitor the battery status parameters in real time during the charging process and compare them with the preset safety threshold to obtain the actual charging effect, calculate the prediction deviation between the actual charging effect and the expected charging effect, and update the charging process proxy model according to the prediction deviation. S5: Correct the causal effect strength based on the prediction bias, and update the policy network with the corrected causal effect strength.
2. The intelligent charging strategy method for UAV batteries according to claim 1, characterized in that, In step S1, the specific process of extracting the causal effect strength of each environmental variable node on the charging effect node using a causal graph attention network, and generating a causal representation vector based on the causal effect strength, includes: The environmental monitoring data or battery status data corresponding to each node in the aforementioned causal graph are used as the feature vectors of each node. The feature vectors of each node are input into the causal graph attention network. The causal graph attention network uses the directed edges in the causal graph as constraints for attention calculation and uses the attention weights of each environmental variable node to the charging effect node as the causal effect strength of each environmental variable node to each charging effect node. The causal effect strength of each environmental variable node on each charging effect node is arranged in a preset order and then vector-concatenated with the current measurement value of each environmental variable to obtain the causal representation vector.
3. The intelligent charging strategy method for UAV batteries according to claim 2, characterized in that, In step S2, the specific process of performing meta-training on the policy network using a model-independent meta-learning algorithm to obtain meta-policy parameters based on the causal representation vector and scene data of historical charging scenarios includes: The policy network is a multi-layer fully connected neural network consisting of an input layer, multiple hidden layers and an output layer. The input layer is used to receive the causal representation vector. Each hidden layer sequentially performs a nonlinear transformation on the output of the previous layer. The output layer maps the output of the last hidden layer to a charging current setting value and a charging voltage setting value. Acquire scene data from multiple historical charging scenarios, and treat each historical charging scenario as a separate meta-training task. Based on the current meta-training task, starting from the current meta-policy parameters, the first loss function is calculated using the training set, and gradient descent is performed on the gradient of the current meta-policy parameters according to the first loss function to obtain the corresponding task-specific parameters. After traversing all meta-training tasks, the sum of the second loss functions of each task-specific parameter on the corresponding task's validation set is calculated. The current meta-policy parameters are updated according to the gradient of the sum of the second loss functions, and the meta-policy parameters at the end of the update are used as the meta-policy parameters after training is completed.
4. The intelligent charging strategy method for UAV batteries according to claim 3, characterized in that, In step S2, the meta-policy parameters are updated using gradients based on the running data of the new scenario to obtain specific charging policy parameters. The specific process of generating multiple candidate charging policy parameters based on the specific charging policy parameters includes: Obtain operational data under the new scenario, wherein the operational data includes multiple sets of mapping samples consisting of the causal representation vector and the actual charging strategy parameters adopted; Starting from the meta-policy parameters, the causal representation vector in each of the mapping samples is used as the input of the policy network, and the charging policy parameters in each of the mapping samples are used as the supervision target to calculate the third loss function of the policy network on the running data. The meta-policy parameters are updated iteratively according to the third loss function, and the updated parameters are used as the dedicated charging policy parameters. Using the charging current setpoint and the charging voltage setpoint as the mean of the first Gaussian distribution and the second Gaussian distribution, respectively, multiple candidate charging current values are sampled from the first Gaussian distribution, and multiple candidate charging voltage values are sampled from the second Gaussian distribution. The candidate charging current values and the candidate charging voltage values are combined to generate multiple candidate charging strategy parameters.
5. The intelligent charging strategy method for UAV batteries according to claim 4, characterized in that, In step S3, the specific process of constructing the charging process proxy model includes: Acquire scenario data of historical charging scenarios, which includes multiple sets of training samples consisting of initial environmental conditions, initial battery state, charging strategy parameters adopted, and corresponding actual charging effects. Using the initial environmental conditions, the initial state of the battery, and the charging strategy parameters as inputs, and the actual charging effect as the supervision target, the charging process proxy model is obtained.
6. The intelligent charging strategy method for drone batteries according to claim 5, characterized in that, In step S3, the specific process of inputting each of the candidate charging strategy parameters into the charging process proxy model for counterfactual inference to obtain the expected charging effect corresponding to each candidate charging strategy parameter includes: Obtain the initial environmental conditions and battery initial state in the current scenario, and combine the initial environmental conditions and battery initial state with the parameters of each candidate charging strategy to generate the input vector corresponding to each candidate strategy. Each input vector is input into the charging process proxy model for forward inference, and the expected charging effect corresponding to each candidate strategy is output. The expected charging effect includes the expected charging efficiency, the expected battery temperature rise rate, and the expected capacity decay rate.
7. The intelligent charging strategy method for UAV batteries according to claim 6, characterized in that, In step S3, the specific process of selecting the candidate strategy that satisfies the safety constraints and has the best effect as the charging strategy to be executed based on the expected charging effect includes: A comprehensive score for each candidate strategy is calculated based on the expected charging effect corresponding to each candidate strategy. The comprehensive score is positively correlated with the expected charging efficiency and negatively correlated with the expected battery temperature rise rate and the expected capacity decay rate. With the expected battery temperature rise rate not exceeding a first threshold and the expected capacity decay rate not exceeding a second threshold as safety constraints, candidate strategies that meet the safety constraints are selected from all candidate strategies. The candidate strategy with the highest comprehensive score that meets the security constraints is selected as the charging strategy to be executed.
8. The intelligent charging strategy method for UAV batteries according to claim 7, characterized in that, In step S4, the specific process of executing the charging strategy to be executed and monitoring the battery status parameters during the charging process and comparing them with a preset safety threshold to obtain the actual charging effect includes: The charging strategy to be executed is performed. During the charging process, the battery terminal voltage, charging current and battery temperature are collected at a preset sampling frequency as the battery state parameters. The battery state parameters are compared with the corresponding preset safety thresholds in real time. When any of the battery state parameters exceeds the corresponding preset safety threshold for a duration exceeding a first preset time, the charging current is reduced; When the duration exceeds the second preset duration, the charging process is terminated, where the second preset duration is greater than the first preset duration. After charging is completed, the actual charging efficiency, actual average temperature rise rate, and actual equivalent capacity decay rate of this charge are obtained and used as the actual charging effect.
9. The intelligent charging strategy method for UAV batteries according to claim 8, characterized in that, In step S4, the specific process of calculating the prediction deviation between the actual charging effect and the expected charging effect, and updating the charging process proxy model based on the prediction deviation, includes: Obtain the expected charging effect corresponding to the charging strategy to be executed, calculate the difference between the actual charging effect and the expected charging effect in each corresponding dimension, and use the difference in each dimension as the prediction deviation. The initial environmental conditions, initial battery state, the charging strategy to be executed, and the actual charging effect are combined into new training samples, and the charging process proxy model is incrementally updated based on the new training samples.
10. The intelligent charging strategy method for UAV batteries according to claim 9, characterized in that, Step S5 specifically includes: The prediction bias is obtained, and a correction loss function is constructed using the prediction bias as a supervision signal. The gradient of the correction loss function with respect to the causal effect strength parameter in the causal graph attention network is calculated. The causal effect strength parameter is updated once according to the gradient along the gradient descent direction, and the updated causal effect strength parameter is used as the corrected causal effect strength. The corrected causal effect strength is used as the output of the causal graph attention network in the next inference to generate the updated causal representation vector. The updated causal representation vector is input into the policy network, which performs forward propagation with the updated causal representation vector as input and outputs the updated charging policy parameters.
Citation Information
Patent Citations
Unmanned aerial vehicle intelligent charging control method and system based on multi-scene adaptation
CN119527602A