Intelligent charging control method and system for unmanned aerial vehicle based on multi-scenario adaptation
Through multi-type sensors to collect environmental data and deep learning algorithms for scene recognition, combined with reinforcement learning algorithms to optimize charging strategies, the problem that existing UAV charging control methods cannot be adaptable is solved, and an efficient and safe charging process is achieved.
Patent Information
- Application Number
- CN202510082222.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing UAV charging control methods cannot be adaptively adjusted according to the characteristics of the actual scenario, resulting in low charging efficiency or safety hazards, and lack of comprehensive considerations for environmental factors and battery health status.
By setting up multiple types of sensors to collect environmental monitoring data, using deep learning algorithms to identify scenes and extract features, and generating scene feature vectors. Based on the reinforcement learning algorithm, the charging optimization model is built, the charging strategy is dynamically optimized, the charging parameters are adjusted adaptively, and the battery status is monitored in real time through the battery management system to establish a closed-loop control mechanism for the charging process.
It significantly improves the intelligence level and adaptability of the UAV charging system, improves the charging efficiency and battery life, and ensures the safety and reliability of the charging process.
Smart Images

Figure CN119527602B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and in particular to an intelligent charging control method and system for unmanned aerial vehicles based on multi-scenario adaptation. Background Art
[0002] With the widespread application of drones in various industries, drone charging management has become a key link to ensure their continuous and stable operation. When drones perform tasks in different scenarios, the environmental conditions and power supply requirements they face vary significantly. In addition, there are also large differences in power supply conditions in different scenarios, including supply voltage fluctuations, power limitations and other factors, which will affect the charging process. Therefore, developing intelligent charging control methods for different application scenarios is of great significance to improving the endurance and service life of drones.
[0003] There are some major deficiencies in the existing UAV charging control methods. The fixed charging strategy cannot be adaptively adjusted according to the characteristics of the actual scenario, resulting in low charging efficiency or safety hazards. There is a lack of comprehensive consideration of environmental factors, making it difficult to cope with complex and changeable application scenarios. The impact of battery health status on the charging process is ignored, and dynamic optimization cannot be achieved, affecting the battery life. The charging control process lacks an effective closed-loop feedback mechanism and cannot respond to environmental changes and battery status changes in a timely manner, posing a greater safety risk.
[0004] In summary, there is an urgent need for an intelligent charging control method for drones based on multi-scenario adaptation, which can collect environmental data in real time, perform scene recognition and feature extraction, establish a mapping relationship between scene features and charging strategies, dynamically optimize the charging strategy, and achieve adaptive adjustment of charging parameters, and establish a closed-loop control mechanism for the charging process to ensure the safety and reliability of the charging process. The present invention can solve the problems in the prior art. Summary of the invention
[0005] The embodiments of the present invention provide a method and system for intelligent charging control of a drone based on multi-scenario adaptation, which can solve the problems in the prior art.
[0006] According to a first aspect of an embodiment of the present invention, a method for intelligent charging control of a drone based on multi-scenario adaptation is provided, comprising: collecting environmental monitoring data through multi-type sensors arranged on the drone, inputting the environmental monitoring data into a scene recognition model, wherein the scene recognition model extracts features of the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector; inputting the scene feature vector into a scene classification unit to perform scene type recognition to obtain scene type identification information; calculating environmental parameter values and power supply power values based on the scene feature vector to generate a scene parameter data set; retrieving a corresponding basic charging strategy from a preset charging strategy database according to the scene type identification information, inputting the scene parameter data set into a charging optimization model, wherein the charging optimization model constructs a strategy evaluation function based on a reinforcement learning algorithm, and performs an evaluation on the strategy according to the calculation result of the strategy evaluation function. The basic charging strategy is optimized and adjusted to generate charging control parameters; the charging power upper limit is determined based on the power supply value, and a charging power adjustment curve is generated in combination with the environmental parameter value; the charging control parameters are dynamically configured according to the charging power adjustment curve, and a charging control instruction is output; the charging control instruction is executed, and the voltage data, current data and temperature data of the battery are collected through the battery management system, and the voltage data, current data and temperature data are respectively compared with the corresponding preset safety thresholds in real time. When any data exceeds the corresponding preset safety threshold, the charging protection control is triggered; the battery health index is calculated based on the voltage data and current data, and the battery health index is input into the battery performance evaluation unit to generate battery performance status data, and the data is fed back to the charging optimization model, and the optimization target of the strategy evaluation function is updated to complete the closed-loop control of the charging process.
[0007] In an optional embodiment, environmental monitoring data is collected by multiple types of sensors installed on an unmanned aerial vehicle, and the environmental monitoring data is input into a scene recognition model. The scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector, including: obtaining environmental monitoring data including scene point cloud data, scene spectral data and scene environment data through a millimeter wave radar array, a spectral imager and an environmental multi-parameter sensor network installed on the unmanned aerial vehicle; generating denoised data through denoising processing based on the environmental monitoring data, and generating a feature tensor through spatiotemporal mapping of the denoised data, and obtaining a multimodal feature tensor through calculation and decomposition; constructing a dynamic neural architecture search framework, and the dynamic neural architecture search framework includes a super network module and a sub-network module; the super network module includes a network structure encoding unit, an enhanced A learning agent unit and a structure evaluation unit; the feature dimension information of the multimodal feature tensor is encoded into a network structure search space through a network structure encoding unit; the reinforcement learning agent unit performs a strategy search in the network structure search space based on a deep reinforcement learning algorithm to determine the network structure; the structure evaluation unit performs a performance evaluation on the network structure and outputs a network structure optimization strategy; the sub-network module constructs a deformable convolution unit based on the network structure optimization strategy, including a spatial sampling sub-unit and a feature weighting sub-unit; the multimodal feature tensor is adaptively spatially sampled through the spatial sampling sub-unit to obtain feature point position information, a multi-head attention network is constructed based on the feature point position information, and feature weight distribution is calculated; the feature weight distribution is tensor multiplied with the multimodal feature tensor to generate a scene feature vector.
[0008] In an optional embodiment, based on the environmental monitoring data, denoising data is generated through denoising processing, the denoised data is mapped through time and space to generate a feature tensor, and after calculation and decomposition, a multimodal feature tensor is obtained, including: performing spatial filtering processing on the scene point cloud data to obtain first denoised data, performing frequency domain filtering processing on the scene spectral data to obtain second denoised data, and performing wavelet transform processing on the scene environment data to obtain third denoised data; constructing a time and space mapping matrix, the time and space mapping matrix includes a time feature mapping submatrix and a space feature mapping submatrix; inputting the first denoised data into the time feature mapping submatrix for time series decomposition to obtain a first feature tensor, and the second The denoising data is input into the spatial feature mapping submatrix for spatial decomposition to obtain a second feature tensor, and the third denoising data is respectively input into the time feature mapping submatrix and the spatial feature mapping submatrix for spatiotemporal decomposition to obtain a third feature tensor; a feature fusion unit including a core tensor generation module and a projection matrix calculation module is constructed, and a tensor product operation is performed on the first feature tensor, the second feature tensor and the third feature tensor through the core tensor generation module to obtain a fused core tensor, and the projection matrix calculation module calculates a modal projection matrix based on the fused core tensor, and a Tucker decomposition operation is performed on the fused core tensor and the modal projection matrix to generate a multimodal feature tensor.
[0009] In an optional embodiment, the scene feature vector is input into a scene classification unit for scene type identification to obtain scene type identification information; environmental parameter values and power supply values are calculated based on the scene feature vector to generate a scene parameter data set, including: inputting the scene feature vector into a scene classification unit, the scene classification unit constructing a node feature matrix based on the scene feature vector, constructing an edge relationship matrix based on the spatial position relationship between nodes, combining the node feature matrix and the edge relationship matrix to generate a scene element relationship graph; inputting the scene element relationship graph into a graph convolutional network for feature aggregation operation to obtain a node vector, and performing a global pooling operation on the node vector to obtain a scene structure Features; perform spatial dimension feature aggregation on the scene structure features to obtain a spatial feature vector, perform channel dimension feature aggregation on the spatial feature vector to obtain scene type identification information; perform probability encoding on the scene feature vector to obtain a priori distribution parameters, construct a Gaussian process regression model based on the prior distribution parameters, input the scene feature vector into the Gaussian process regression model to calculate the environmental parameter value; input the scene feature vector into a load prediction model, and correct the prediction result of the load prediction model based on the environmental parameter value to obtain the power supply value; combine the environmental parameter value and the power supply power value according to a preset data structure format to generate a scene parameter data set.
[0010] In an optional embodiment, according to the scenario type identification information, the corresponding basic charging strategy is retrieved from a preset charging strategy database, the scenario parameter data set is input into a charging optimization model, the charging optimization model constructs a strategy evaluation function based on a reinforcement learning algorithm, and the basic charging strategy is optimized and adjusted according to the calculation result of the strategy evaluation function, and the generation of charging control parameters includes: retrieving the basic charging strategy from a preset charging strategy database according to the scenario type identification information, using a multi-head attention mechanism to perform time series feature analysis on the basic charging strategy to obtain feature weights, inputting the feature weights into a residual connection network to extract long-term and short-term dependency features, and generating a strategy feature vector including a charging time window parameter, a power allocation ratio parameter, and a priority setting parameter based on the long-term and short-term dependency features; inputting the scenario parameter data set and the strategy feature vector into the charging optimization model, and constructing a reward function, the reward function including a charging efficiency benefit value calculated based on an adaptive weight coefficient, a time response benefit value calculated based on an exponential decay function, and a power balance benefit value calculated based on a power fluctuation penalty term; based on the charging efficiency benefit value, the time response benefit value, and the power The strategy evaluation function is constructed based on the rate balance benefit value, and the strategy evaluation function includes an action value network and a strategy network; the action value network takes the strategy feature vector as input, adopts a double-delay deep deterministic policy gradient algorithm to calculate the state action value, and constructs a value optimization target based on the time difference error between the state action value and the target state action value; the strategy network inputs the state action value into a deterministic policy gradient model, calculates the policy gradient value based on the deterministic policy gradient model, and updates the reward function parameters according to the policy gradient value to obtain a strategy evaluation result; the strategy evaluation result is compared with a preset priority threshold, and when the strategy evaluation result is greater than the priority threshold, the corresponding strategy adjustment sample is marked as a high-value strategy sample, a priority queue is constructed based on the high-value strategy sample, and the priority queue is stored in a priority experience replay memory; the high-value strategy sample is extracted from the priority experience replay memory in the order of the priority queue, and the charging time window parameter, power allocation ratio parameter and priority setting parameter of the basic charging strategy are soft-updated and optimized by using an exponential sliding average method to generate charging control parameters.
[0011] In an optional embodiment, a charging power upper limit is determined based on the power supply power value, a charging power adjustment curve is generated in combination with the environmental parameter value, the charging control parameter is dynamically configured according to the charging power adjustment curve, and the charging control instruction is output, including: determining a charging power upper limit based on the power supply power value, inputting the environmental parameter value into a gated cyclic unit network to extract time-varying features, and inputting the time-varying features into a temperature-sensitive adaptive layer and a humidity-sensitive adaptive layer respectively to obtain a temperature adjustment coefficient and a humidity adjustment coefficient; constructing a power adjustment function based on the temperature adjustment coefficient and the humidity adjustment coefficient, and converting the charging power upper limit into the power adjustment function. The method comprises the following steps: adaptively weighting the actual power and the charging power adjustment curve to obtain a charging power adjustment curve; using a dynamic programming algorithm to calculate a deviation between the actual power and the charging power adjustment curve, and dynamically configuring the charging control parameters based on the deviation, including: when the deviation exceeds a preset threshold, dynamically adjusting the power by using a sliding average filter; when the deviation is within a preset threshold, maintaining the charging control parameters unchanged by using an adaptive dead zone control; performing a safety check on the configured charging control parameters based on fuzzy control rules, fine-tuning the checked charging control parameters according to the statistical characteristics of historical charging data, and combining the finely adjusted charging control parameters to generate a charging control instruction.
[0012] In an optional embodiment, the charging control instruction is executed, and the voltage data, current data and temperature data of the battery are collected through the battery management system, and the voltage data, current data and temperature data are respectively compared with the corresponding preset safety thresholds in real time. When any data exceeds the corresponding preset safety threshold, the charging protection control is triggered; the battery health index is calculated based on the voltage data and current data, and the battery health index is input into the battery performance evaluation unit to generate battery performance status data, and the battery performance status data is fed back to the charging optimization model, and the optimization target of the strategy evaluation function is updated to complete the closed-loop control of the charging process. The battery management system is used to collect the single cell voltage data and the total voltage data according to the preset sampling frequency, Input the single cell voltage data and the total voltage data into a digital filter to eliminate high-frequency noise and obtain filtered voltage data; use a Hall current sensor to collect current data, and perform temperature drift correction on the current data based on a preset temperature compensation coefficient to obtain corrected current data; use a thermistor to collect temperature data, and input the temperature data into a piecewise linear interpolation model to obtain a temperature characteristic value; calculate a dynamic voltage threshold based on the temperature characteristic value, the corrected current data and a preset reference voltage; calculate the battery state of charge value based on the corrected current data, and determine the segmented current threshold based on the battery state of charge value; determine the temperature protection upper limit value and the temperature protection lower limit value based on the temperature characteristic value; The filtered voltage data is compared with the dynamic voltage threshold in real time, the corrected current data is compared with the segmented current threshold in real time, the temperature characteristic value is compared with the temperature protection upper limit value and the temperature protection lower limit value in real time, and a charging protection signal is output when any comparison result exceeds the corresponding threshold; the battery capacity value is calculated by the coulomb counting method based on the corrected current data, and the battery capacity value is corrected according to the temperature characteristic value to obtain the corrected capacity value; the voltage jump amount is calculated based on the filtered voltage data, the current step amount is calculated based on the corrected current data, and the battery internal resistance value is calculated according to the voltage jump amount and the current step amount; the corrected capacity value is compared with The ratio of the preset initial capacity and the ratio of the battery internal resistance to the preset initial internal resistance are respectively weightedly calculated with the corresponding weight coefficients to obtain a battery health index; the health state impact value is calculated based on the battery health index, and the health attenuation impact value is calculated based on the change rate of the battery health index; the health state impact value and the health attenuation impact value are weightedly processed to obtain an optimization target correction value, and the optimization target correction value is weightedly fused with the preset charging optimization target to obtain an updated optimization target; a dynamic programming method is adopted, with the dynamic voltage threshold, the segmented current threshold, the temperature protection upper limit value and the temperature protection lower limit value as constraints, and the optimal charging strategy is solved based on the updated optimization target.
[0013] According to a second aspect of an embodiment of the present invention, there is provided a drone intelligent charging control system based on multi-scenario adaptation, comprising: a first unit, for collecting environmental monitoring data through multi-type sensors arranged on a drone, inputting the environmental monitoring data into a scene recognition model, wherein the scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector; inputting the scene feature vector into a scene classification unit for scene type recognition to obtain scene type identification information; calculating environmental parameter values and power supply power values based on the scene feature vector to generate a scene parameter data set; a second unit, for retrieving a corresponding basic charging strategy from a preset charging strategy database according to the scene type identification information, inputting the scene parameter data set into a charging optimization model, wherein the charging optimization model constructs a strategy evaluation function based on a reinforcement learning algorithm, and according to the calculation of the strategy evaluation function The basic charging strategy is optimized and adjusted according to the result to generate charging control parameters; the charging power upper limit is determined based on the power supply value, and a charging power adjustment curve is generated in combination with the environmental parameter value; the charging control parameters are dynamically configured according to the charging power adjustment curve, and a charging control instruction is output; a third unit is used to execute the charging control instruction, collect the voltage data, current data and temperature data of the battery through the battery management system, and compare the voltage data, current data and temperature data with the corresponding preset safety thresholds in real time. When any data exceeds the corresponding preset safety threshold, the charging protection control is triggered; the battery health index is calculated based on the voltage data and current data, and the battery health index is input into the battery performance evaluation unit to generate battery performance status data, and the data is fed back to the charging optimization model, the optimization target of the strategy evaluation function is updated, and the closed-loop control of the charging process is completed.
[0014] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0015] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0016] In the embodiment of the present invention, environmental data is collected by multi-type sensors, and scene recognition and feature extraction are performed using deep learning algorithms. It is possible to accurately identify different scene types and formulate charging strategies in a targeted manner, significantly improving the intelligence level and adaptability of the UAV charging system; a charging optimization model based on reinforcement learning is adopted, and the charging strategy can be dynamically adjusted according to real-time scene parameters and battery performance status, so as to achieve precise control and real-time optimization of charging power, and effectively improve charging efficiency and battery life; the voltage, current and temperature are monitored in real time and compared with safety thresholds through the battery management system, and a charging protection mechanism is established. At the same time, a battery health index evaluation is introduced, which can effectively ensure charging safety, and realize the self-adaptation and continuous optimization of the charging process through a closed-loop control mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a method for intelligent charging control of a drone based on multi-scenario adaptation according to an embodiment of the present invention;
[0018] Figure 2 The figure is a schematic diagram of the structure of an intelligent charging control system for unmanned aerial vehicles based on multi-scenario adaptation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0021] Figure 1 FIG. 1 is a flow chart of a method for intelligent charging control of a drone based on multi-scenario adaptation according to an embodiment of the present invention. Figure 1As shown, the method includes: S101. Collecting environmental monitoring data through multiple types of sensors arranged on the unmanned aerial vehicle, inputting the environmental monitoring data into a scene recognition model, and the scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector; inputting the scene feature vector into a scene classification unit to identify the scene type and obtain scene type identification information; calculating environmental parameter values and power supply values based on the scene feature vector to generate a scene parameter data set; in this embodiment, by extracting features from the environmental monitoring data through a deep learning algorithm, different scene types can be more accurately identified and classified; combined with the scene feature vector, the environmental parameter values and power supply values can be automatically calculated to improve the degree of automation of data processing; the generated scene parameter data set can be used for further analysis or decision support, which is conducive to the continuous optimization and intelligent management of environmental monitoring.
[0022] In an optional embodiment, environmental monitoring data is collected by multiple types of sensors installed on an unmanned aerial vehicle, and the environmental monitoring data is input into a scene recognition model. The scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector, including: obtaining environmental monitoring data including scene point cloud data, scene spectral data and scene environment data through a millimeter wave radar array, a spectral imager and an environmental multi-parameter sensor network installed on the unmanned aerial vehicle; generating denoised data through denoising processing based on the environmental monitoring data, and generating a feature tensor through spatiotemporal mapping of the denoised data, and obtaining a multimodal feature tensor through calculation and decomposition; constructing a dynamic neural architecture search framework, and the dynamic neural architecture search framework includes a super network module and a sub-network module; the super network module includes a network structure encoding unit, an enhanced A learning agent unit and a structure evaluation unit; the feature dimension information of the multimodal feature tensor is encoded into a network structure search space through a network structure encoding unit; the reinforcement learning agent unit performs a strategy search in the network structure search space based on a deep reinforcement learning algorithm to determine the network structure; the structure evaluation unit performs a performance evaluation on the network structure and outputs a network structure optimization strategy; the sub-network module constructs a deformable convolution unit based on the network structure optimization strategy, including a spatial sampling sub-unit and a feature weighting sub-unit; the multimodal feature tensor is adaptively spatially sampled through the spatial sampling sub-unit to obtain feature point position information, a multi-head attention network is constructed based on the feature point position information, and feature weight distribution is calculated; the feature weight distribution is tensor multiplied with the multimodal feature tensor to generate a scene feature vector.
[0023] In an optional implementation, first, a millimeter-wave radar array, a spectral imager, and an environmental multi-parameter sensor network are configured on the drone. The millimeter-wave radar array consists of 16 transmitting units and 32 receiving units, with an operating frequency of 77GHz, a scanning range of 120 degrees horizontally and 60 degrees vertically, and a sampling frequency of 10Hz; the spectral imager has a band range of 400-1000nm, a spectral resolution of 2nm, and an imaging resolution of 1920×1080 pixels; the environmental multi-parameter sensor network includes a temperature sensor, a humidity sensor, an air pressure sensor, and a light sensor, and the sampling frequency is 1Hz.
[0024] The collected raw data was preprocessed for denoising. For millimeter-wave radar data, the adaptive Gaussian filtering method was used, with a filter window size of 5×5 and a smoothing coefficient of 0.8; the spectral data was denoised using wavelet transform, with the db4 wavelet basis and 3 decomposition layers; the environmental data was median filtered with a filter window size of 3. The denoised data was converted into a feature tensor through spatiotemporal mapping, with the dimension of point cloud data being 256×256×32, the dimension of spectral data being 256×256×64, and the dimension of environmental data being 256×256×8.
[0025] A dynamic neural architecture search framework is constructed. The network structure encoding unit in the hypernetwork module encodes the dimensional information of the feature tensor into a search space. The search space contains the number of convolutional layers in the range of [3, 8], the convolution kernel size can be selected from [3, 5, 7], and the number of channels can be selected from [32, 64, 128]. The reinforcement learning agent unit uses a deep deterministic policy gradient algorithm to search the network structure. The exploration cycle is 200 rounds, and 50 candidate structures are evaluated in each round. The structure evaluation unit is based on a comprehensive score of accuracy and computational complexity, with a score weight ratio of 7:3.
[0026] The subnetwork module adopts a deformable convolution structure. The number of sampling points of the spatial sampling subunit is 9, and the sampling range is 3×3 area. The feature weighting subunit constructs an 8-head attention network with an attention dimension of 64 and a dropout rate of 0.1. Through iterative optimization, the search is stopped when the score improvement is less than 0.1% for 5 consecutive rounds. The final output scene feature vector dimension is 1×2048.
[0027] In a specific example, data is collected in an urban environment. The point cloud data shows that there are fixed obstacles such as buildings and trees within a range of 100 meters. The spectral data reflects the characteristic spectra of vegetation and building materials. The environmental data shows a temperature of 25°C, a relative humidity of 65%, an air pressure of 101.3kPa, and a light intensity of 15,000 lux. After feature extraction and fusion, the generated scene feature vector can accurately characterize the environmental characteristics for subsequent scene recognition and decision control.
[0028] The entire processing flow realizes closed-loop control from multi-source heterogeneous data collection to feature extraction and fusion. Each processing link is interconnected to ensure the real-time and accuracy of feature extraction. This method can adapt to the scene perception needs in different complex environments and provide a reliable environmental perception foundation for UAV intelligent control.
[0029] In this embodiment, through the point cloud, spectrum and environmental data collected by various sensors, more comprehensive and accurate environmental monitoring data can be generated, thereby improving the accuracy of scene analysis; feature tensors are generated through noise reduction processing and spatiotemporal mapping, and the optimal deep learning network structure is constructed using dynamic neural architecture search, thereby improving the accuracy and robustness of feature extraction; through the deep reinforcement learning agent and structural evaluation unit in the super network module, the network structure can be dynamically optimized to improve the adaptability and efficiency in different scenarios and data modes; in the sub-network module, through adaptive spatial sampling and multi-head attention mechanism, feature weights are accurately calculated and feature vectors are generated to ensure that the extracted features have high distinctiveness.
[0030] In an optional embodiment, based on the environmental monitoring data, denoising data is generated through denoising processing, the denoised data is mapped through time and space to generate a feature tensor, and after calculation and decomposition, a multimodal feature tensor is obtained, including: performing spatial filtering processing on the scene point cloud data to obtain first denoised data, performing frequency domain filtering processing on the scene spectral data to obtain second denoised data, and performing wavelet transform processing on the scene environment data to obtain third denoised data; constructing a time and space mapping matrix, the time and space mapping matrix includes a time feature mapping submatrix and a space feature mapping submatrix; inputting the first denoised data into the time feature mapping submatrix for time series decomposition to obtain a first feature tensor, and the second The denoising data is input into the spatial feature mapping submatrix for spatial decomposition to obtain a second feature tensor, and the third denoising data is respectively input into the time feature mapping submatrix and the spatial feature mapping submatrix for spatiotemporal decomposition to obtain a third feature tensor; a feature fusion unit including a core tensor generation module and a projection matrix calculation module is constructed, and a tensor product operation is performed on the first feature tensor, the second feature tensor and the third feature tensor through the core tensor generation module to obtain a fused core tensor, and the projection matrix calculation module calculates a modal projection matrix based on the fused core tensor, and a Tucker decomposition operation is performed on the fused core tensor and the modal projection matrix to generate a multimodal feature tensor.
[0031] In an optional implementation, an adaptive Gaussian filter is used for spatial filtering of scene point cloud data, the filter window size is set to 7×7 pixels, the filter radius is 2.5 meters, and the Gaussian kernel parameters are adaptively adjusted according to the point cloud density. Taking an actual scene as an example, the density of the original point cloud data is 80 points / square meter, and the noise standard deviation is 0.15 meters. After spatial filtering, the point cloud noise standard deviation is reduced to 0.03 meters, while maintaining the clarity of edge features, and obtaining the first noise reduction data.
[0032] When filtering the scene spectral data in the frequency domain, first convert the data to the frequency domain through fast Fourier transform, set the cutoff frequency to 1 / 4 of the original sampling frequency, and use a Butterworth low-pass filter for processing. The original resolution of the spectral data is 1nm, and the data volume is 1024 bands. After filtering, 250 valid bands are retained, and the smoothness of the spectral curve is improved by 90%, obtaining the second noise reduction data.
[0033] The wavelet transform processing of the scene environment data uses the db4 wavelet basis function and performs a 4-layer decomposition. The original environment data contains parameters such as temperature, humidity, and air pressure, and the sampling frequency is 10Hz. The high-frequency noise is removed by threshold processing, and the signal-to-noise ratio is improved by 15dB after reconstruction, and the third noise reduction data is obtained.
[0034] In the process of constructing the spatiotemporal mapping matrix, the temporal feature mapping submatrix adopts a sliding time window mechanism with a window length of 2 seconds, a step size of 0.2 seconds, and a feature dimension of 10. The spatial feature mapping submatrix adopts a multi-scale spatial pyramid structure, which contains 3 scale layers, with a downsampling rate of 2 for each layer and a feature dimension of 16.
[0035] The first denoised data is decomposed in time series to generate a time series feature sequence with a dimension of 256×10, and then the first feature tensor is obtained through time feature mapping with a dimension of 256×10×16. The second denoised data is decomposed in space to generate a multi-scale feature map, and the second feature tensor is obtained through spatial feature mapping with a dimension of 256×16×16. The third denoised data is decomposed in time and space at the same time to obtain the third feature tensor with a dimension of 256×10×16.
[0036] In the feature fusion unit, the core tensor generation module uses a weighted tensor product operation, and the weight coefficient is determined by adaptive learning. Taking the actual data as an example, the weight of the first feature tensor is 0.4, the weight of the second feature tensor is 0.35, and the weight of the third feature tensor is 0.25. The fusion obtains a core tensor with a dimension of 256×16×16.
[0037] The projection matrix calculation module constructs projection matrices of three modes based on the feature distribution of the core tensor, with dimensions of 256 × 64, 16 × 32, and 16 × 32. Through Tucker decomposition operation, the core tensor is projected into a unified feature space, and finally a multimodal feature tensor with a dimension of 64 × 32 × 32 is generated.
[0038] In this embodiment, spatial, frequency and wavelet transform are performed on the point cloud, spectrum and environmental data respectively to reduce noise interference and make the data cleaner and more reliable; time and space decomposition is performed through the space-time mapping matrix, so that the space-time characteristics of each data type can be extracted, and the accuracy of feature separation and expression can be improved; the feature fusion unit uses tensor product operation and Tucker decomposition to fuse each modal feature into a multimodal feature tensor, which retains the complementarity of multi-source information and improves the quality and distinctiveness of the overall feature expression; through the calculation of the modal projection matrix and the decomposition of the core tensor, the multimodal feature data can be effectively compressed, the computational complexity can be reduced, and the system processing efficiency can be improved.
[0039] In an optional embodiment, the scene feature vector is input into a scene classification unit for scene type identification to obtain scene type identification information; environmental parameter values and power supply values are calculated based on the scene feature vector to generate a scene parameter data set, including: inputting the scene feature vector into a scene classification unit, the scene classification unit constructing a node feature matrix based on the scene feature vector, constructing an edge relationship matrix based on the spatial position relationship between nodes, combining the node feature matrix and the edge relationship matrix to generate a scene element relationship graph; inputting the scene element relationship graph into a graph convolutional network for feature aggregation operation to obtain a node vector, and performing a global pooling operation on the node vector to obtain a scene structure Features; perform spatial dimension feature aggregation on the scene structure features to obtain a spatial feature vector, perform channel dimension feature aggregation on the spatial feature vector to obtain scene type identification information; perform probability encoding on the scene feature vector to obtain a priori distribution parameters, construct a Gaussian process regression model based on the prior distribution parameters, input the scene feature vector into the Gaussian process regression model to calculate the environmental parameter value; input the scene feature vector into a load prediction model, and correct the prediction result of the load prediction model based on the environmental parameter value to obtain the power supply value; combine the environmental parameter value and the power supply power value according to a preset data structure format to generate a scene parameter data set.
[0040] In an optional implementation, a scene classification unit is first constructed, and the scene feature vector with an input dimension of 2048 is reshaped into a 32×64 node feature matrix. Based on the spatial distribution of elements in the scene, the distance and azimuth between nodes are calculated to generate a 32×32 edge relationship matrix. When the distance between two nodes is less than a preset threshold of 5 meters, a connection relationship is established in the edge relationship matrix. The node feature matrix and the edge relationship matrix together constitute a scene element relationship graph, which is used to describe the scene structure characteristics.
[0041] Taking the urban environment scene as an example, the node feature matrix contains the feature information of scene elements such as buildings, roads, and greenery, and the edge relationship matrix reflects the spatial topological relationship between these elements. The scene element relationship graph is input into a three-layer graph convolutional network, and the number of output channels of each layer is 128, 64, and 32 respectively. The first layer captures local structural features, the second layer integrates medium-scale information, and the third layer extracts global semantic features. After graph convolution, feature vectors of 32 nodes are obtained, and each vector dimension is 32.
[0042] By combining global maximum pooling and average pooling, the node vectors are aggregated to generate 256-dimensional scene structure features. During the pooling process, maximum pooling retains significant features, average pooling retains the overall feature distribution, and the pooling weight ratio is 6:4.
[0043] The scene structure features are aggregated in the spatial dimension using a self-attention mechanism with 8 attention heads and 32 hidden dimensions to obtain a 64-dimensional spatial feature vector. The spatial feature vector is then aggregated in the channel dimension and the softmax function is used to map the features to the predefined scene category space, and finally the scene type identification information is output. The scene types include urban areas, suburbs, industrial parks, natural environments, and other categories.
[0044] The 2048-dimensional scene feature vector is probabilistically encoded, and the mean vector and variance vector are generated as prior distribution parameters through the fully connected layer. The dimensions of both vectors are 64. A Gaussian process regression model is constructed based on these parameters. The kernel function selects the RBF kernel, the length scale is 0.5, and the signal variance is 1.0. The scene feature vector is input into the model to predict the environmental parameter values including temperature, humidity, air pressure, and light.
[0045] The load forecasting model uses a three-layer long short-term memory network structure with 128 hidden units. It inputs the scene feature vector and predicts the power supply change in the next 30 minutes. The forecast results are corrected based on the environmental parameter values, and the correction coefficient is adaptively adjusted according to the degree of influence of the environmental parameters. For example, when the temperature exceeds 35°C, the predicted power value is increased by 15% to consider the impact of increased heat dissipation demand.
[0046] Finally, the environmental parameter values and power supply values are organized into a scene parameter data set according to the preset format. The data set adopts a key-value pair structure and contains four main fields: timestamp, scene ID, environmental parameter array, and power array. The environmental parameter array contains 8 parameter values, and the power array contains the power values of 30 predicted time points.
[0047] In this embodiment, an element relationship graph is constructed through scene feature vectors, and feature aggregation is performed using a graph convolutional network to extract the structural features of the scene, thereby improving the accuracy of scene classification. Feature aggregation in spatial and channel dimensions further extracts global and local information of the scene, enhances the expressiveness of scene features, and ensures accurate acquisition of scene type identification information. Environmental parameters are automatically inferred from scene features through probability coding and Gaussian process regression models to achieve flexible adaptation to changing environments. The load prediction model is corrected using environmental parameters to make the power supply prediction more in line with actual needs and improve the reliability of the prediction. Environmental parameters and power values are organized into standardized data sets to facilitate subsequent analysis or application and support intelligent decision-making and management.
[0048] S102. According to the scene type identification information, the corresponding basic charging strategy is retrieved from the preset charging strategy database, and the scene parameter data set is input into the charging optimization model. The charging optimization model constructs a strategy evaluation function based on the reinforcement learning algorithm, optimizes and adjusts the basic charging strategy according to the calculation result of the strategy evaluation function, and generates charging control parameters; the charging power upper limit is determined based on the power supply power value, and the charging power adjustment curve is generated in combination with the environmental parameter value. The charging control parameters are dynamically configured according to the charging power adjustment curve, and the charging control instructions are output; in this embodiment, according to the scene type identification information, the basic charging strategy is extracted from the charging strategy database, and the strategy is optimized by the reinforcement learning algorithm to ensure that the charging scheme adapts to different environments and needs and improves the charging efficiency; the charging power upper limit is determined based on the power supply power value, and the charging power adjustment curve is generated in combination with the environmental parameters, which can flexibly respond to environmental changes and ensure the stability and efficiency of the charging process; by dynamically configuring the charging control parameters, the power distribution in the charging process is optimized, energy waste is reduced and charging efficiency is improved; the charging control instructions finally generated can automatically adjust the charging system, reduce manual intervention, and improve the automation and intelligence level of the system.
[0049] In an optional embodiment, according to the scenario type identification information, the corresponding basic charging strategy is retrieved from a preset charging strategy database, the scenario parameter data set is input into a charging optimization model, the charging optimization model constructs a strategy evaluation function based on a reinforcement learning algorithm, and the basic charging strategy is optimized and adjusted according to the calculation result of the strategy evaluation function, and the generation of charging control parameters includes: retrieving the basic charging strategy from a preset charging strategy database according to the scenario type identification information, using a multi-head attention mechanism to perform time series feature analysis on the basic charging strategy to obtain feature weights, inputting the feature weights into a residual connection network to extract long-term and short-term dependency features, and generating a strategy feature vector including a charging time window parameter, a power allocation ratio parameter, and a priority setting parameter based on the long-term and short-term dependency features; inputting the scenario parameter data set and the strategy feature vector into the charging optimization model, and constructing a reward function, the reward function including a charging efficiency benefit value calculated based on an adaptive weight coefficient, a time response benefit value calculated based on an exponential decay function, and a power balance benefit value calculated based on a power fluctuation penalty term; based on the charging efficiency benefit value, the time response benefit value, and the power The strategy evaluation function is constructed based on the rate balance benefit value, and the strategy evaluation function includes an action value network and a strategy network; the action value network takes the strategy feature vector as input, adopts a double-delay deep deterministic policy gradient algorithm to calculate the state action value, and constructs a value optimization target based on the time difference error between the state action value and the target state action value; the strategy network inputs the state action value into a deterministic policy gradient model, calculates the policy gradient value based on the deterministic policy gradient model, and updates the reward function parameters according to the policy gradient value to obtain a strategy evaluation result; the strategy evaluation result is compared with a preset priority threshold, and when the strategy evaluation result is greater than the priority threshold, the corresponding strategy adjustment sample is marked as a high-value strategy sample, a priority queue is constructed based on the high-value strategy sample, and the priority queue is stored in a priority experience replay memory; the high-value strategy sample is extracted from the priority experience replay memory in the order of the priority queue, and the charging time window parameter, power allocation ratio parameter and priority setting parameter of the basic charging strategy are soft-updated and optimized by using an exponential sliding average method to generate charging control parameters.
[0050] First, the basic charging strategy is retrieved from the charging strategy database according to the scene type identification information. The database presets basic strategy templates for different scenarios such as urban areas, suburbs, and industrial parks. An 8-head attention mechanism is used to analyze the time series characteristics of the basic charging strategy. The dimension of the attention head is 64. The query matrix, key matrix, and value matrix are generated through linear mapping, and the feature weights are calculated.
[0051] The feature weights are input into a residual connection network with four residual blocks, each of which contains two convolutional layers and a skip connection. The network extracts the long-term and short-term dependency features in the strategy sequence and outputs a feature vector with a dimension of 256. Based on the feature vector, a strategy feature vector is generated, which contains three sets of data: charging time window parameters, power allocation ratio parameters, and priority setting parameters. The dimension of each set of parameters is 32.
[0052] The charging optimization model constructs a reward function that includes three benefit components. The charging efficiency benefit value is calculated using an adaptive weight coefficient, which is dynamically adjusted according to the current load rate. The weight is highest when the load rate is in the range of 60%-80%. The time response benefit value is calculated using an exponential decay function with a decay rate of 0.05. The shorter the response time, the higher the benefit value. The power balance benefit value is calculated using the power fluctuation penalty term, and a negative incentive is generated when the power fluctuation exceeds the set threshold.
[0053] The policy evaluation function consists of an action value network and a policy network. The action value network uses a four-layer fully connected network structure, with 256, 128, 64 and 32 hidden units, respectively. The policy feature vector is input to calculate the state action value. A double-delayed deep deterministic policy gradient algorithm is used, with a target network update cycle of 200 steps and a soft update coefficient of 0.005. The value optimization target is constructed based on the temporal difference error.
[0054] The policy network inputs the state action value into the deterministic policy gradient model. The model contains three fully connected layers, and the output layer uses the tanh activation function to limit the action range. The policy gradient value guides the reward function parameter update, with an update step of 0.001, and the target policy network parameters are updated every 1000 steps.
[0055] Taking the actual scenario as an example, the basic charging strategy of a charging station in a certain city during the morning rush hour on weekdays is: charging time window 7:00-9:00, power allocation ratio 0.6, priority setting parameter 0.8. After strategy evaluation, the charging efficiency benefit value is 0.85, the time response benefit value is 0.92, and the power balance benefit value is 0.78, and the comprehensive evaluation result is 0.85.
[0056] Compare the strategy evaluation result with the preset priority threshold of 0.8. If the evaluation result is greater than the threshold, the strategy adjustment sample is marked as a high-value strategy sample. Construct a priority queue with a queue length of 1000 and sort in descending order by evaluation score. Store the priority queue in a priority experience replay memory with a capacity of 10,000.
[0057] The high-value strategy samples are extracted from the memory in priority order, and the exponential sliding average method is used for parameter soft update, with a smoothing coefficient of 0.1. The optimized charging control parameters are: the charging time window is adjusted to 6:30-9:30, the power allocation ratio is increased to 0.75, and the priority setting parameter is reduced to 0.7.
[0058] In this embodiment, a multi-head attention mechanism is used to analyze the timing characteristics of the basic charging strategy, and the long-term and short-term dependency characteristics are extracted through the residual connection network, so that the generated strategy characteristics are more in line with the actual dynamic needs and the adaptability of the charging strategy is improved; through the reward function including charging efficiency, time response, and power balance, the charging efficiency and response speed can be optimized under different trade-off conditions, and the power fluctuation can be reduced to achieve a comprehensive balance of benefits; the constructed strategy evaluation function is combined with the double-delay deep deterministic policy gradient algorithm to dynamically calculate and optimize the state action value, thereby improving the accuracy and effect of the charging strategy adjustment and making the charging process more efficient; based on the strategy evaluation results, a priority queue of high-value strategy samples is established and stored in the experience replay memory, and the high-value strategy is continuously strengthened and optimized through the priority experience replay mechanism, which significantly improves the continuous optimization ability of the strategy; the exponential sliding average method is used to soft-update the charging strategy parameters, so that the adjustment of the charging control parameters is smoother and more stable, and the interference caused by the sudden change of strategy switching is reduced.
[0059] In an optional embodiment, a charging power upper limit is determined based on the power supply power value, a charging power adjustment curve is generated in combination with the environmental parameter value, the charging control parameter is dynamically configured according to the charging power adjustment curve, and the charging control instruction is output, including: determining a charging power upper limit based on the power supply power value, inputting the environmental parameter value into a gated cyclic unit network to extract time-varying features, and inputting the time-varying features into a temperature-sensitive adaptive layer and a humidity-sensitive adaptive layer respectively to obtain a temperature adjustment coefficient and a humidity adjustment coefficient; constructing a power adjustment function based on the temperature adjustment coefficient and the humidity adjustment coefficient, and converting the charging power upper limit into the power adjustment function. The method comprises the following steps: adaptively weighting the actual power and the charging power adjustment curve to obtain a charging power adjustment curve; using a dynamic programming algorithm to calculate a deviation between the actual power and the charging power adjustment curve, and dynamically configuring the charging control parameters based on the deviation, including: when the deviation exceeds a preset threshold, dynamically adjusting the power by using a sliding average filter; when the deviation is within a preset threshold, maintaining the charging control parameters unchanged by using an adaptive dead zone control; performing a safety check on the configured charging control parameters based on fuzzy control rules, fine-tuning the checked charging control parameters according to the statistical characteristics of historical charging data, and combining the finely adjusted charging control parameters to generate a charging control instruction.
[0060] First, the upper limit of charging power is determined according to the power supply value. Taking a charging station as an example, when the power supply value is 500kW, considering the 15% safety margin, the upper limit of charging power is set to 425kW. The environmental parameter values are input into the double-layer gated recurrent unit network. The number of hidden units in the network is 128. The input includes 8 environmental parameters such as temperature, humidity, and air pressure. The tanh activation function is used to extract time-varying features.
[0061] The time-varying features are input into the temperature-sensitive adaptive layer and the humidity-sensitive adaptive layer respectively. The temperature-sensitive adaptive layer contains three layers of fully connected networks, with 64, 32 and 16 hidden units respectively, and uses the ReLU activation function. When the ambient temperature varies from -10°C to 45°C, a temperature adjustment coefficient of 0.6 to 1.2 is generated. The humidity-sensitive adaptive layer has the same structure, and when the relative humidity varies from 30% to 90%, a humidity adjustment coefficient of 0.8 to 1.1 is generated.
[0062] When constructing the power regulation function, a weighted combination method is used to integrate the effects of temperature and humidity. The temperature influence weight is 0.6, and the humidity influence weight is 0.4. For example, when the ambient temperature is 35°C and the relative humidity is 75%, the temperature regulation coefficient is 0.85, the humidity regulation coefficient is 0.95, and the power regulation coefficient is 0.89 obtained by adaptive weighting. Multiply the charging power upper limit of 425kW by the power regulation function to obtain the charging power regulation curve, and the maximum value of the curve is 378.25kW.
[0063] The dynamic programming algorithm is used to calculate the deviation between the actual power and the charging power regulation curve. The algorithm uses a forward recursive method with a time step of 1 minute. The state space is the power deviation value and the decision space is the power regulation amount. When the actual power is 390kW, the deviation from the regulation curve is 11.75kW, which exceeds the preset 10kW threshold.
[0064] At this time, the dynamic power adjustment mechanism is triggered, and a 5-point sliding average filter is used for power adjustment, with filter window weights of 0.1, 0.2, 0.4, 0.2, and 0.1 respectively. After filtering, the power adjustment amount is determined to be -8kW. When the deviation value drops to 7kW, it enters the preset ±10kW threshold range, and an adaptive dead zone control strategy is adopted with a dead zone width of 5kW, keeping the charging control parameters unchanged.
[0065] The configured charging control parameters are safety checked based on fuzzy control rules. The fuzzy control rules include temperature rule set and humidity rule set, each of which contains 5 fuzzy levels. For example, when the temperature is "high" and the humidity is "medium", the power limit level is "medium to low". According to the rule reasoning results, the charging power is limited to 85% of the rated power.
[0066] When fine-tuning the calibrated charging control parameters, the charging data of the last 24 hours is counted and the standard deviation of power fluctuation is calculated. When the standard deviation is greater than the set threshold, the power limit is reduced by 5%; when the standard deviation is less than 50% of the threshold, the power limit is increased by 3%. The charging control instructions finally generated include: charging power upper limit value 370kW, power adjustment coefficient 0.89, time window parameter 6:30-9:30 and other control parameters.
[0067] In this embodiment, an adjustment coefficient is dynamically generated based on environmental parameters (such as temperature and humidity), and time-varying features are extracted through a gated recurrent unit network to ensure that the charging process can adapt to environmental changes, thereby optimizing power regulation; by constructing a power adjustment function and adaptively weighting the upper limit of the charging power, a charging power adjustment curve is generated, which can flexibly adjust the charging power according to environmental changes and power requirements, thereby improving the efficiency and stability of the charging process; a dynamic programming algorithm is used to calculate the power deviation, and it is adjusted through a sliding average filter and an adaptive dead zone control strategy to ensure that it can be automatically adjusted when the deviation is large and remain stable when the deviation is small, thereby avoiding excessive adjustment or unnecessary fluctuations; a safety check is performed based on fuzzy control rules to ensure that the charging control parameters operate within a reasonable range; and the charging control parameters are fine-tuned through the statistical characteristics of historical data to further optimize the charging strategy and improve the safety and accuracy of the charging process; through comprehensive adjustment, dynamic adjustment and fine-tuning mechanisms, efficient, safe and stable control is achieved during the charging process, thereby avoiding a decrease in charging efficiency or risks caused by environmental changes.
[0068] S103. Execute the charging control instruction, collect the voltage data, current data and temperature data of the battery through the battery management system, compare the voltage data, current data and temperature data with the corresponding preset safety thresholds in real time, and trigger the charging protection control when any data exceeds the corresponding preset safety threshold; calculate the battery health index based on the voltage data and current data, input the battery health index into the battery performance evaluation unit, generate battery performance status data, and feed it back to the charging optimization model, update the optimization target of the strategy evaluation function, and complete the closed-loop control of the charging process.
[0069] In this embodiment, the battery management system collects voltage, current and temperature data in real time and compares them with preset thresholds, so as to detect anomalies in time and trigger protection control to prevent potential dangers such as overcharging and overheating, thereby improving the safety of the charging process; the battery health index is calculated based on the voltage and current data, which helps to dynamically evaluate the health status of the battery so that the charging strategy can be adjusted as the battery condition changes; the battery performance status data is fed back to the charging optimization model, and the optimization target of the strategy evaluation function is updated to achieve closed-loop control of the charging process, so that the charging strategy continuously adapts to the battery health status and improves the charging efficiency and battery life; the performance evaluation unit continuously monitors the battery status, which can effectively prevent the battery performance from deteriorating, ensuring that the charging process is more intelligent and health maintenance is carried out simultaneously.
[0070] In an optional embodiment, the charging control instruction is executed, and the voltage data, current data and temperature data of the battery are collected through the battery management system, and the voltage data, current data and temperature data are respectively compared with the corresponding preset safety thresholds in real time. When any data exceeds the corresponding preset safety threshold, the charging protection control is triggered; the battery health index is calculated based on the voltage data and current data, and the battery health index is input into the battery performance evaluation unit to generate battery performance status data, and the battery performance status data is fed back to the charging optimization model, and the optimization target of the strategy evaluation function is updated to complete the closed-loop control of the charging process. The battery management system is used to collect the single cell voltage data and the total voltage data according to the preset sampling frequency, Input the single cell voltage data and the total voltage data into a digital filter to eliminate high-frequency noise and obtain filtered voltage data; use a Hall current sensor to collect current data, and perform temperature drift correction on the current data based on a preset temperature compensation coefficient to obtain corrected current data; use a thermistor to collect temperature data, and input the temperature data into a piecewise linear interpolation model to obtain a temperature characteristic value; calculate a dynamic voltage threshold based on the temperature characteristic value, the corrected current data and a preset reference voltage; calculate the battery state of charge value based on the corrected current data, and determine the segmented current threshold based on the battery state of charge value; determine the temperature protection upper limit value and the temperature protection lower limit value based on the temperature characteristic value; The filtered voltage data is compared with the dynamic voltage threshold in real time, the corrected current data is compared with the segmented current threshold in real time, the temperature characteristic value is compared with the temperature protection upper limit value and the temperature protection lower limit value in real time, and a charging protection signal is output when any comparison result exceeds the corresponding threshold; the battery capacity value is calculated by the coulomb counting method based on the corrected current data, and the battery capacity value is corrected according to the temperature characteristic value to obtain the corrected capacity value; the voltage jump amount is calculated based on the filtered voltage data, the current step amount is calculated based on the corrected current data, and the battery internal resistance value is calculated according to the voltage jump amount and the current step amount; the corrected capacity value is compared with The ratio of the preset initial capacity and the ratio of the battery internal resistance to the preset initial internal resistance are respectively weightedly calculated with the corresponding weight coefficients to obtain a battery health index; the health state impact value is calculated based on the battery health index, and the health attenuation impact value is calculated based on the change rate of the battery health index; the health state impact value and the health attenuation impact value are weightedly processed to obtain an optimization target correction value, and the optimization target correction value is weightedly fused with the preset charging optimization target to obtain an updated optimization target; a dynamic programming method is adopted, with the dynamic voltage threshold, the segmented current threshold, the temperature protection upper limit value and the temperature protection lower limit value as constraints, and the optimal charging strategy is solved based on the updated optimization target.
[0071] The battery management system collects data at a sampling frequency of 100Hz. Taking a certain type of power battery as an example, the nominal voltage of a single cell is 3.7V and the total voltage is 384V. A Butterworth digital filter is used for noise elimination, with a filter order of 4 and a cutoff frequency of 10Hz. After filtering, the fluctuation range of the single cell voltage data is reduced to ±0.01V, and the fluctuation range of the total voltage is reduced to ±0.5V.
[0072] The Hall current sensor has a measurement range of -500A to 500A and an accuracy of 0.1%. The zero point drifts by about 0.2% for every 10°C increase in ambient temperature. The temperature compensation coefficient is set to 0.02% / °C to correct the current data. A 10kΩ thermistor is used to measure temperature with a range of -40°C to 85°C. The piecewise linear interpolation model divides the temperature range into 7 segments, and a linear fit is used to obtain the temperature characteristic value for each segment.
[0073] The dynamic voltage threshold is calculated based on the temperature characteristic value, the corrected current data and the preset reference voltage of 4.2V. When the temperature is 25℃ and the charging current is 0.5C, the single cell dynamic voltage threshold is 4.15V. The battery state of charge value is calculated based on the corrected current data. When the state of charge is in the range of 0-20%, the current threshold is 1x rate charging current (i.e., the charging current equal to the nominal capacity of the battery, which is 100A for a 100Ah battery); the current threshold in the range of 20%-80% is 0.8x rate charging current (i.e., the charging current equal to 0.8 times the nominal capacity of the battery, which is 80A for a 100Ah battery); the current threshold in the range of 80%-100% is 0.3x rate charging current (i.e., the charging current equal to 0.3 times the nominal capacity of the battery, which is 30A for a 100Ah battery). The upper limit of the temperature protection is set to 55℃, and the lower limit is set to 0℃.
[0074] Real-time monitoring of the deviation between data and threshold. When it is detected that the cell voltage exceeds 4.15V, the charging current exceeds the corresponding interval threshold, or the temperature exceeds the range of 0-55℃, the charging protection signal is immediately output to trigger the protection mechanism. For example, when a cell voltage reaches 4.18V, the system outputs an overvoltage protection signal.
[0075] The battery capacity is calculated using the coulomb counting method. The 100Ah battery was charged and discharged at 25°C, and the calculated capacity was 98.5Ah. Correction was made based on the temperature characteristic value. When the temperature was 35°C, the correction factor was 1.05, and the corrected capacity was 103.4Ah. Through a step test of a 0.1x charging current (i.e., a charging current equal to 0.1 times the nominal capacity of the battery, which is 10A for a 100Ah battery), the voltage jump and current step were measured, and the battery internal resistance was calculated to be 2.5mΩ.
[0076] In the calculation of the battery health index, the capacity attenuation weight is 0.6 and the internal resistance increase weight is 0.4. When the ratio of the corrected capacity to the initial capacity is 0.95 and the ratio of the internal resistance to the initial internal resistance is 1.1, the calculated battery health index is 0.92. Based on this health index, the health status impact value is calculated to be 0.85, and the health attenuation impact value is calculated to be 0.9 based on the recent health index change rate of -0.5% / month.
[0077] The health status impact value and health decay impact value are weighted in the ratio of 0.7:0.3 to obtain the optimization target correction value of 0.865. The correction value is weighted and fused with the preset charging optimization target, with weights of 0.4 and 0.6 respectively, to obtain the updated optimization target value.
[0078] The dynamic programming method is used to solve the optimal charging strategy, with a time step of 1 minute, the state space is the state of charge, and the decision space is the charging power. Under the constraints of voltage, current and temperature, the segmented charging strategy is obtained: 0-20% state of charge uses a 0.8x charging current (80A), 20%-80% uses a 0.6x charging current (60A), and 80%-100% uses a 0.2x charging current (20A).
[0079] In this embodiment, high-frequency noise is eliminated by a digital filter to ensure that the battery voltage data is more accurate, avoid erroneous judgments caused by noise interference, and improve data quality; the temperature compensation coefficient is used to correct the current data, and the temperature data is processed in combination with the piecewise linear interpolation model, which effectively eliminates the influence of temperature changes on current and voltage, and improves the accuracy and stability of the data; through the real-time comparison of the dynamic voltage threshold, the segmented current threshold and the temperature protection threshold, it is possible to detect and avoid situations beyond the safety range in real time, prevent risks such as overcharging or overheating, and ensure charging safety; the battery capacity is calculated by the coulomb counting method, and the weighted calculation of the battery internal resistance and the health index is combined to accurately evaluate the battery health status. Through the changes in the corrected capacity value and the battery internal resistance, the battery performance degradation is continuously monitored; through the evaluation of the battery health index and the change rate, the health state impact value and the health decay impact value are calculated, and the charging optimization target is dynamically corrected to ensure that the charging strategy is adjusted according to the real-time health status of the battery and extend the battery life; combined with dynamic programming and multiple constraints (such as voltage, temperature and current thresholds), the optimal charging strategy is generated through the updated optimization target to ensure that the charging process is not only efficient but also safe.
[0080] Figure 2 FIG. 1 is a schematic diagram of a structure of a drone intelligent charging control system based on multi-scenario adaptation according to an embodiment of the present invention. Figure 2As shown, the system includes: a first unit, which is used to collect environmental monitoring data through multiple types of sensors set on the drone, input the environmental monitoring data into a scene recognition model, and the scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector; the scene feature vector is input into a scene classification unit to identify the scene type and obtain scene type identification information; the environmental parameter value and the power supply value are calculated based on the scene feature vector to generate a scene parameter data set; the second unit is used to retrieve the corresponding basic charging strategy from a preset charging strategy database according to the scene type identification information, input the scene parameter data set into a charging optimization model, the charging optimization model constructs a strategy evaluation function based on a reinforcement learning algorithm, and optimizes the basic charging strategy according to the calculation result of the strategy evaluation function. The charging control parameters are generated by adjusting the charging power upper limit based on the power supply power value, generating a charging power adjustment curve in combination with the environmental parameter value, dynamically configuring the charging control parameters according to the charging power adjustment curve, and outputting a charging control instruction; a third unit is used to execute the charging control instruction, collect the voltage data, current data and temperature data of the battery through the battery management system, and compare the voltage data, current data and temperature data with the corresponding preset safety thresholds in real time. When any data exceeds the corresponding preset safety threshold, the charging protection control is triggered; the battery health index is calculated based on the voltage data and current data, and the battery health index is input into the battery performance evaluation unit to generate battery performance status data, and the data is fed back to the charging optimization model, the optimization target of the strategy evaluation function is updated, and the closed-loop control of the charging process is completed.
[0081] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0082] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0083] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The intelligent charging control method for drones based on multi-scenario adaptation is characterized by: include: Collect environmental monitoring data through multiple types of sensors installed on the drone, input the environmental monitoring data into a scene recognition model, and the scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector; Inputting the scene feature vector into a scene classification unit to identify the scene type and obtain scene type identification information; calculating the environmental parameter value and the power supply value based on the scene feature vector to generate a scene parameter data set; Retrieving the corresponding basic charging strategy from a preset charging strategy database according to the scenario type identification information, inputting the scenario parameter data set into a charging optimization model, constructing a strategy evaluation function based on a reinforcement learning algorithm, optimizing and adjusting the basic charging strategy according to a calculation result of the strategy evaluation function, and generating charging control parameters; Determine a charging power upper limit based on the power supply value, generate a charging power adjustment curve in combination with the environmental parameter value, dynamically configure the charging control parameter according to the charging power adjustment curve, and output a charging control instruction; Execute the charging control instruction, collect the voltage data, current data and temperature data of the battery through the battery management system, compare the voltage data, current data and temperature data with the corresponding preset safety thresholds in real time, and trigger the charging protection control when any data exceeds the corresponding preset safety threshold; Calculating a battery health index based on voltage data and current data, inputting the battery health index into a battery performance evaluation unit, generating battery performance status data, and feeding back the data to the charging optimization model, updating the optimization target of the strategy evaluation function, and completing closed-loop control of the charging process; Retrieving a basic charging strategy from a preset charging strategy database according to the scene type identification information, performing time series feature analysis on the basic charging strategy using a multi-head attention mechanism to obtain feature weights, inputting the feature weights into a residual connection network to extract long-term and short-term dependency features, and generating a strategy feature vector including a charging time window parameter, a power allocation ratio parameter, and a priority setting parameter based on the long-term and short-term dependency features; Inputting the scenario parameter data set and the strategy feature vector into the charging optimization model, constructing a reward function, wherein the reward function includes a charging efficiency benefit value calculated based on an adaptive weight coefficient, a time response benefit value calculated based on an exponential decay function, and a power balance benefit value calculated based on a power fluctuation penalty term; A policy evaluation function is constructed based on the charging efficiency benefit value, the time response benefit value, and the power balance benefit value, wherein the policy evaluation function includes an action value network and a policy network; the action value network takes the policy feature vector as input, calculates the state action value using a double-delay deep deterministic policy gradient algorithm, and constructs a value optimization target based on a time difference error between the state action value and the target state action value; the policy network inputs the state action value into a deterministic policy gradient model, calculates a policy gradient value based on the deterministic policy gradient model, and updates the reward function according to the policy gradient value to obtain a policy evaluation result; Comparing the strategy evaluation result with a preset priority threshold, when the strategy evaluation result is greater than the priority threshold, marking the corresponding strategy adjustment sample as a high-value strategy sample, building a priority queue based on the high-value strategy sample, and storing the priority queue in a priority experience replay memory; Extracting the high-value strategy samples from the priority experience replay memory in the order of the priority queue, performing soft update optimization on the charging time window parameters, power allocation ratio parameters and priority setting parameters of the basic charging strategy using an exponential moving average method, and generating charging control parameters; Determine the upper limit of charging power based on the power supply value, input the environmental parameter value into the gated recurrent unit network to extract time-varying features, and input the time-varying features into the temperature-sensitive adaptive layer and the humidity-sensitive adaptive layer respectively to obtain a temperature adjustment coefficient and a humidity adjustment coefficient; Based on the temperature adjustment coefficient and the humidity adjustment coefficient, construct a power adjustment function, and adaptively weight the charging power upper limit and the power adjustment function to obtain a charging power adjustment curve; The method uses a dynamic programming algorithm to calculate a deviation between actual power and the charging power adjustment curve, and dynamically configures the charging control parameters based on the deviation, including: When the deviation value exceeds the preset threshold, a sliding average filter is used to dynamically adjust the power; when the deviation value is within the preset threshold, an adaptive dead zone control is used to keep the charging control parameters unchanged; The configured charging control parameters are safety checked based on fuzzy control rules, the checked charging control parameters are fine-tuned according to the statistical characteristics of historical charging data, and the fine-tuned charging control parameters are combined to generate charging control instructions.
2. The method according to claim 1, characterized in that Environmental monitoring data is collected by multiple types of sensors installed on the drone, and the environmental monitoring data is input into a scene recognition model. The scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector, including: Environmental monitoring data including scene point cloud data, scene spectral data and scene environment data are obtained through the millimeter wave radar array, spectral imager and environmental multi-parameter sensor network installed on the UAV; Based on the environmental monitoring data, noise reduction processing is performed to generate noise reduction data, and the noise reduction data is mapped in time and space to generate a feature tensor, and after calculation and decomposition, a multi-modal feature tensor is obtained; Constructing a dynamic neural architecture search framework, wherein the dynamic neural architecture search framework includes a super network module and a sub network module; the super network module includes a network structure encoding unit, a reinforcement learning agent unit and a structure evaluation unit; The feature dimension information of the multimodal feature tensor is encoded into a network structure search space through a network structure encoding unit, the reinforcement learning agent unit performs a strategy search in the network structure search space based on a deep reinforcement learning algorithm to determine the network structure, and the structure evaluation unit performs a performance evaluation on the network structure and outputs a network structure optimization strategy; The sub-network module constructs a deformable convolution unit based on the network structure optimization strategy, including a spatial sampling sub-unit and a feature weighting sub-unit; the multi-modal feature tensor is adaptively spatially sampled by the spatial sampling sub-unit to obtain feature point position information, a multi-head attention network is constructed based on the feature point position information, and feature weight distribution is calculated; A tensor multiplication operation is performed on the feature weight distribution and the multimodal feature tensor to generate a scene feature vector.
3. The method according to claim 2, characterized in that Based on the environmental monitoring data, noise reduction processing is performed to generate noise reduction data. The noise reduction data is mapped in time and space to generate feature tensors. After calculation and decomposition, the multimodal feature tensors are obtained, including: Performing spatial filtering on the scene point cloud data to obtain first denoised data, performing frequency domain filtering on the scene spectrum data to obtain second denoised data, and performing wavelet transform on the scene environment data to obtain third denoised data; Constructing a spatiotemporal mapping matrix, wherein the spatiotemporal mapping matrix includes a time feature mapping submatrix and a space feature mapping submatrix; inputting the first denoised data into the time feature mapping submatrix for time series decomposition to obtain a first feature tensor; inputting the second denoised data into the space feature mapping submatrix for space decomposition to obtain a second feature tensor; and inputting the third denoised data into the time feature mapping submatrix and the space feature mapping submatrix for spatiotemporal decomposition to obtain a third feature tensor; A feature fusion unit including a core tensor generation module and a projection matrix calculation module is constructed. The core tensor generation module performs a tensor product operation on the first feature tensor, the second feature tensor and the third feature tensor to obtain a fused core tensor. The projection matrix calculation module calculates a modal projection matrix based on the fused core tensor. The fused core tensor and the modal projection matrix are subjected to Tucker decomposition operation to generate a multimodal feature tensor.
4. The method according to claim 2, characterized in that: Inputting the scene feature vector into a scene classification unit to identify the scene type and obtain scene type identification information; calculating the environmental parameter value and the power supply value based on the scene feature vector to generate a scene parameter data set includes: The scene feature vector is input into a scene classification unit, the scene classification unit constructs a node feature matrix based on the scene feature vector, constructs an edge relationship matrix based on the spatial position relationship between nodes, and combines the node feature matrix and the edge relationship matrix to generate a scene element relationship graph; Input the scene element relationship graph into a graph convolutional network to perform feature aggregation operations to obtain node vectors, and perform global pooling operations on the node vectors to obtain scene structure features; Performing spatial dimension feature aggregation on the scene structure feature to obtain a spatial feature vector, and performing channel dimension feature aggregation on the spatial feature vector to obtain scene type identification information; Probabilistically encoding the scene feature vector to obtain a priori distribution parameters, constructing a Gaussian process regression model based on the prior distribution parameters, and inputting the scene feature vector into the Gaussian process regression model to calculate and obtain an environmental parameter value; Inputting the scene feature vector into a load prediction model, and correcting the prediction result of the load prediction model based on the environmental parameter value to obtain a power supply value; The environmental parameter value and the power supply value are combined according to a preset data structure format to generate a scene parameter data set.
5. The method according to claim 1, characterized in that Execute the charging control instruction, collect the voltage data, current data and temperature data of the battery through the battery management system, compare the voltage data, current data and temperature data with the corresponding preset safety thresholds in real time, and trigger the charging protection control when any data exceeds the corresponding preset safety threshold; Calculating a battery health index based on voltage data and current data, inputting the battery health index into a battery performance evaluation unit, generating battery performance status data, and feeding back the battery performance status data to the charging optimization model, updating the optimization target of the strategy evaluation function, and completing the closed-loop control of the charging process include: A battery management system is used to collect single cell voltage data and total voltage data according to a preset sampling frequency, and the single cell voltage data and the total voltage data are input into a digital filter to eliminate high frequency noise, thereby obtaining filtered voltage data; A Hall current sensor is used to collect current data, and the current data is subjected to temperature drift correction based on a preset temperature compensation coefficient to obtain corrected current data; a thermistor is used to collect temperature data, and the temperature data is input into a piecewise linear interpolation model to obtain a temperature characteristic value; Calculate a dynamic voltage threshold according to the temperature characteristic value, the corrected current data and a preset reference voltage; calculate a battery state of charge value according to the corrected current data, and determine a segmented current threshold based on the battery state of charge value; determine a temperature protection upper limit value and a temperature protection lower limit value according to the temperature characteristic value; Compare the filtered voltage data with the dynamic voltage threshold in real time, compare the corrected current data with the segmented current threshold in real time, compare the temperature characteristic value with the temperature protection upper limit and the temperature protection lower limit in real time, and output a charging protection signal when any comparison result exceeds the corresponding threshold; The battery capacity value is calculated based on the corrected current data by using the coulomb counting method, and the battery capacity value is corrected according to the temperature characteristic value to obtain the corrected capacity value; the voltage jump amount is calculated based on the filtered voltage data, the current step amount is calculated based on the corrected current data, and the battery internal resistance value is calculated according to the voltage jump amount and the current step amount; The ratio of the corrected capacity value to the preset initial capacity, the ratio of the battery internal resistance value to the preset initial internal resistance and the corresponding weight coefficients are respectively weighted to obtain a battery health index; a health state impact value is calculated based on the battery health index, and a health attenuation impact value is calculated based on the change rate of the battery health index; The health state impact value and the health decay impact value are weighted to obtain an optimization target correction value, and the optimization target correction value is weightedly integrated with a preset charging optimization target to obtain an updated optimization target; A dynamic programming method is adopted, with the dynamic voltage threshold, the segmented current threshold, the temperature protection upper limit value and the temperature protection lower limit value as constraints, and an optimal charging strategy is obtained based on the updated optimization target.
6. An intelligent charging control system for unmanned aerial vehicles based on multi-scenario adaptation, used to implement the method described in any one of claims 1 to 5, characterized in that: include: The first unit is used to collect environmental monitoring data through multiple types of sensors arranged on the drone, input the environmental monitoring data into a scene recognition model, and the scene recognition model extracts features from the environmental monitoring data based on a deep learning algorithm to generate a scene feature vector; Inputting the scene feature vector into a scene classification unit to identify the scene type and obtain scene type identification information; calculating the environmental parameter value and the power supply value based on the scene feature vector to generate a scene parameter data set; The second unit is used to retrieve the corresponding basic charging strategy from a preset charging strategy database according to the scenario type identification information, input the scenario parameter data set into a charging optimization model, the charging optimization model constructs a strategy evaluation function based on a reinforcement learning algorithm, optimizes and adjusts the basic charging strategy according to the calculation result of the strategy evaluation function, and generates charging control parameters; Determine a charging power upper limit based on the power supply value, generate a charging power adjustment curve in combination with the environmental parameter value, dynamically configure the charging control parameter according to the charging power adjustment curve, and output a charging control instruction; The third unit is used to execute the charging control instruction, collect voltage data, current data and temperature data of the battery through the battery management system, and compare the voltage data, current data and temperature data with the corresponding preset safety thresholds in real time. When any data exceeds the corresponding preset safety threshold, the charging protection control is triggered; The battery health index is calculated based on the voltage data and the current data, and the battery health index is input into the battery performance evaluation unit to generate battery performance status data, which is fed back to the charging optimization model to update the optimization target of the strategy evaluation function and complete the closed-loop control of the charging process.
7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Cluster electric vehicle charging behavior optimization method based on deep reinforcement learning
CN111934335A
Lithium battery intelligent charging control method based on reinforcement learning
CN117578679A