Laser wireless energy transfer dynamic MPPT control method based on RL-PSO hybrid algorithm and reconfigurable array system
By using the RL-PSO hybrid algorithm and reconfigurable photovoltaic array technology, the problems of power mismatch and insufficient topology reconfiguration capability of laser wireless energy receiving system under dynamic irradiation environment are solved, realizing global maximum power tracking and array topology adaptive switching, thereby improving system efficiency and reliability.
Patent Information
- Application Number
- CN202511112934.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-10
- Publication Date
- 2025-11-14
AI Technical Summary
Existing laser wireless energy receiving systems suffer from power mismatch, MPPT tracking lag, and insufficient topology reconstruction capabilities under dynamic irradiation environments, making it difficult to adapt to dynamic spot distribution and photovoltaic unit aging, resulting in decreased system efficiency.
A dynamic MPPT control method for laser wireless power transfer based on the RL-PSO hybrid algorithm and a reconfigurable array system are adopted. By integrating light intensity distribution, temperature gradient and historical data through an intelligent control center, a three-dimensional optimization parameter space is constructed. PSO and reinforcement learning algorithms are combined for parallel optimization to achieve global maximum power tracking and array topology adaptive switching under dynamic light spot distribution, and to perform battery health sensing management.
It significantly improves energy transmission efficiency, increases light energy utilization by 15%-22%, shortens topology reconfiguration response time to 10ms, improves system reliability and battery life, and maximizes global efficiency.
Smart Images

Figure CN120949889A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of laser wireless power transfer, specifically to a dynamic MPPT control method and reconfigurable array system for laser wireless power transfer based on the RL-PSO hybrid algorithm. Background Technology
[0002] Laser wireless power transmission technology, due to its high directivity, long-distance transmission capability, and resistance to electromagnetic interference, has demonstrated significant application value in fields such as drone endurance, space station power supply, and deep-sea equipment power supply. In existing technologies, laser energy receivers generally employ fixed-structure photovoltaic arrays, optimizing photoelectric conversion efficiency through the maximum power point tracking (MPPT) algorithm. However, significant technical bottlenecks exist in practical applications: First, the laser spot is affected by atmospheric turbulence, mechanical vibration, and other factors during transmission, resulting in a dynamically non-uniform irradiance distribution at the receiver. The series / parallel topology of traditional static arrays struggles to adapt to rapid changes in light intensity, easily leading to local hot spot effects and power mismatch problems. Second, conventional MPPT algorithms (such as the perturbation-observation method and the incremental conductance method) suffer from tracking lag and oscillatory convergence defects under dynamic irradiation conditions, failing to effectively address power surges caused by spot drift or local shading. Third, existing systems lack the ability to dynamically reconstruct the array topology; when some photovoltaic units experience performance degradation due to aging or damage, the overall system efficiency will significantly decrease.
[0003] In recent years, researchers have attempted to improve system performance by refining MPPT control strategies. For example, some schemes employ fuzzy logic or neural network algorithms to optimize tracking speed, but their high computational complexity and reliance on large amounts of training data make them difficult to meet real-time requirements. Other technologies propose array partitioning control based on multi-sensor feedback, adjusting local circuit connections through a switching matrix; however, such methods typically only support limited topology mode switching and do not consider the balance between energy consumption and power gain during topology switching. Regarding energy storage management, existing systems often employ fixed-threshold charging strategies, ignoring the dynamic impact of battery aging on charging parameters, which can lead to accelerated capacity decay over long-term operation. Furthermore, the various modules of a laser receiving system (such as optical alignment, MPPT control, and energy storage management) often operate independently, lacking a collaborative optimization mechanism, making it difficult to maximize global efficiency.
[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0005] Objective: This invention addresses the problems of power mismatch, MPPT tracking lag, and insufficient topology reconfiguration capability in existing laser wireless energy receiving systems under dynamic irradiation environments. It proposes a laser wireless energy receiving system and control method based on a hybrid algorithm of dynamic topology optimization and RL-PSO. The core of this invention lies in achieving global maximum power tracking under dynamic beam distribution, adaptive array topology switching, and battery health sensing management through the coordinated control of a reconfigurable photovoltaic array and intelligent optimization algorithms. This significantly improves energy transmission efficiency and system reliability.
[0006] To achieve the above objectives, the technical solution adopted by this invention is: a dynamic MPPT control method for laser wireless power transfer based on the RL-PSO hybrid algorithm and a reconfigurable array system, comprising the following steps:
[0007] S1. The laser emitting module initiates a self-test program to calibrate the collimating lens focal length and beam deflection angle; the photovoltaic array receiving module performs a health check, detects units with abnormal open-circuit voltage and marks them as redundant nodes.
[0008] S2. The irradiance sensor network acquires light intensity distribution data at a sampling frequency of 100Hz, eliminates pulse noise through median filtering, and generates a normalized irradiance heat map.
[0009] S3, the intelligent control center integrates light intensity distribution, temperature gradient, and historical operating data to construct a three-dimensional optimization parameter space, in which irradiance data is mapped to 0-1000W / m² using 8-bit grayscale values. 2 scope;
[0010] S4. Run the PSO and reinforcement learning algorithms in parallel. Set the PSO population size to 50-100 and the number of iterations to 20-50. The reinforcement learning algorithm uses a double-Q network structure with an experience replay buffer capacity of 10. 4 strip;
[0011] S5. When the reconstruction yield index RI>0.15, an instruction packet containing the target topology code, MPPT reference voltage and PID parameters is generated and transmitted to the relay matrix via RS-485 bus.
[0012] S6, Real-time monitoring system conversion efficiency η total =η opt ·η pv ·η mppt ·η chg If the efficiency drops by more than 5% for three consecutive cycles, a global parameter recalibration will be triggered.
[0013] S7. The charging curve is dynamically adjusted according to SOC and SOH. When the battery internal resistance is detected to rise by more than 20% of the initial value, the system will force the battery to enter trickle charging mode and issue a maintenance alarm.
[0014] Preferably, in step S1, the self-test procedure of the laser emitting module includes the following process: First, a status diagnosis is performed using an internal temperature sensor, a drive current detection unit, and a beam quality analyzer. If the temperature deviation ΔT > ±5℃ or the drive current fluctuation rate σ / I... avg If the value exceeds 3%, a fault lockout is triggered. Subsequently, when calibrating the collimating lens focal length, closed-loop feedback control is used to adjust the displacement Δd of the piezoelectric ceramic actuator, so that the spot diameter D... spot Satisfy the minimization condition:
[0015]
[0016] Where λ is the laser wavelength (unit: nm), M 2 Here, f is the beam quality factor, f is the current lens focal length (mm), w0 is the beam waist radius (μm), and the calibration target is D. spot ≤1.5mm; Beam deflection angle calibration is achieved by adjusting the reflector tilt angle θ using a two-dimensional servo motor. x With θ y ,satisfy:
[0017]
[0018] Where Δx and Δy are the spot center offsets (unit: mm), and L is the calibration distance from the transmitter to the receiver (unit: m). Meanwhile, the health diagnosis of the photovoltaic array receiver module uses the open-circuit voltage threshold method: standard test illumination (1000W / m²) is applied to each photovoltaic unit. 2 ), measure its open-circuit voltage V oc The judgment is as follows:
[0019]
[0020] Where V oc_nom This is the nominal open-circuit voltage. If this formula is met, the unit is determined to be abnormal, marked as a redundant node, and switched to bypass mode to ensure the overall continuity of the array.
[0021] Preferably, in step S2, the irradiance sensor network consists of an M×N (M≥8, N≥8) high-precision photodiode array, which synchronously collects light intensity distribution data at a sampling frequency of 100Hz, and its spatial resolution satisfies:
[0022]
[0023] Among them, D array The effective side length of the receiving array (unit: cm) is given, and Δs is the center-to-center distance between adjacent sensors. The median filtering uses a 3×3 sliding window to eliminate impulse noise in the original data Graw(x,y,t), and the processing formula is as follows:
[0024]
[0025] Wherein, Median represents the median operation, and the dynamic range of the filtered data is compressed to 0–1000 W / m. 2 Normalized irradiance thermal map is generated by the following formula:
[0026]
[0027] Wherein, Gmin and Gmax are the minimum and maximum irradiance values of the entire field during the current sampling period (unit: W / m²). 2 The normalized result is stored in the form of an 8-bit grayscale matrix, where 0 corresponds to 0 W / m. 2 255 corresponds to 1000W / m 2 The heatmap update cycle is synchronized with the sampling frequency (10ms), and a double buffering mechanism is used to avoid data tearing.
[0028] Preferably, in step S3, the intelligent control center constructs a three-dimensional optimization parameter space through multi-source data fusion, and the specific process is as follows:
[0029] (1) Light intensity distribution data mapping: The 8-bit grayscale matrix Gnorm(x,y) generated in step S2 is converted into the actual irradiance value G. actual (x,y)(Unit: W / m) 2 ):
[0030]
[0031] Among them, G norm (x,y)∈[0,255] represents the normalized grayscale value, 1000W / m 2 Corresponding to full-scale irradiance.
[0032] (2) Temperature gradient field calculation: The surface temperature distribution T(x,y) of the photovoltaic unit is obtained through an infrared thermal imager array, and the temperature difference gradient between adjacent units is calculated.
[0033]
[0034] Wherein, ΔT(x,y), in ℃, represents the temperature difference of a 3×3 region centered at coordinates (x,y), and is used to characterize the risk of hot spots.
[0035] (3) Dynamic weighting of historical data: Extracting MPPT efficiency η from the historical operating database under the same spot distribution pattern. hist Topology configuration coding C hist and battery degradation coefficient β hist Construct time decay weights:
[0036]
[0037] Among them, t now t is the current timestamp. hist For historical time stamps, the weight decreases as the data timeliness index decreases.
[0038] (4) Three-dimensional parameter space synthesis: The above data is projected into a three-dimensional space of spatial coordinates (x,y), irradiance G, and temperature gradient ΔT, and historical weights are fused to generate an optimization objective function:
[0039]
[0040] Where α = 0.6, β = 0.3, and γ = 0.1 are weighting coefficients (satisfying α + β + γ = 1), and ∈ = 0.1℃ is a zero-limiting constant. This space is used for the RL-PSO algorithm to search for the global optimum, while constraining the battery health state (β). hist ≤0.2).
[0041] (5) Data compression and storage: The three-dimensional parameter space is stored in a sparse matrix format, and only F is retained for non-zero elements. optim ≥0.7×F max In high-value areas, storage efficiency is improved by 60%.
[0042] Preferably, in step S4, a parallel hybrid optimization architecture of PSO and reinforcement learning (RL) is adopted: the PSO algorithm performs a global search with a particle swarm size of 50-100, and the position of each particle is encoded as the MPPT operating voltage V of the photovoltaic array. pv With topology configuration C topo The combined vector, whose velocity update follows a dynamic inertia weighting strategy:
[0043] v i (t+1)=ω(t)v i (t)+c1r1(p best -x i (t))+c2r2(g best -x i (t))
[0044] Where ω(t) is a linearly decreasing inertia weight (initial value 1.2, final value 0.4), c1=c2=2.0 are learning factors, r1, r2 are uniformly random numbers in [0,1], and p best For the individual historical optimal solution, g best The solution is globally optimal; the fitness function is defined as a weighted sum of output power and heat loss.
[0045] f(x i ) = P pv-0.05·max(ΔT)
[0046] Among them, P pv ΔT represents the power under the current topology, and ΔT represents the array temperature difference.
[0047] Preferably, in step S4, the reinforcement learning module adopts a dual-Q network structure, with the main network Q... main With the target network Q target Receives data including light intensity distribution G(x,y), temperature field T(x,y), and electrical parameter V. pv ,I pv The state vector s outputs the topology switching action a and the PSO parameter correction instruction; the experience playback buffer capacity is 10. 4 For each sample, samples with higher temporal difference error (TD-error) are sampled first, and their priority is calculated as follows:
[0048]
[0049] The two networks are updated synchronously after every 10 PSO iterations, by minimizing the mean square error L = ∑(Q target -Q main ) 2 / N completes training.
[0050] The collaborative mechanism between PSO and RL is as follows: when the PSO convergence speed decreases (power improvement rate < 0.1% for 10 consecutive iterations), the RL network outputs a topology reconstruction command and a PSO parameter reset signal (such as forcibly restoring the inertia weight ω to 1.0), triggering a global search restart; conversely, if the RL action causes the power P... pv If the value exceeds the PSO optimum by 5%, the action strategy is injected into the PSO particle swarm as an initial solution to accelerate hybrid convergence.
[0051] Preferably, in step S5, the calculation and instruction generation logic of the reconfiguration gain index (RI) is as follows: when the system predicts that the power gain after reconfiguration exceeds the switching loss, an instruction packet is triggered, and RI is defined as:
[0052]
[0053] Among them, P pred To predict the maximum power of the topology (in W), P current For the current power, and P switch This represents the total switching loss of the relay. It is calculated as follows:
[0054]
[0055] Where, N sw For the number of operating relays, I k Ron ,k represents the conducting current (A) and resistance (Ω) of the k-th relay, respectively, and t represents the t-th relay. switch The single switching time is denoted as s. When RI > 0.15, an instruction packet containing the target topology code, MPPT reference voltage, and PID parameters is generated and transmitted via the RS-485 bus. The RI threshold of 0.15 is determined through regression analysis of historical data, corresponding to an optimized balance point where system efficiency improvement is ≥8% and switching loss percentage is ≤2%.
[0056] Preferably, in step S5, the target topology encoding uses a 16-bit binary mask to identify the unit connection relationship (1-series, 0-parallel), and the MPPT reference voltage is calculated based on the PSO optimal solution with weights.
[0057]
[0058] Among them, P avg The average power of the particle swarm is given, and the PID parameters are set according to K. p =K p0 (1+0.2RI), T i =T i0 / (1+0.1RI),T d =T d0 (1+0.05RI) dynamically adjusted to adapt to the reconfigured system response. The instruction packet is encapsulated via the Modbus-RTU protocol, and the data field includes topology encoding (2 bytes), V... ref (4-byte floating-point) and PID parameters (4 bytes each), CRC check (polynomial 0xA001) ensures the transmission error rate is less than 10. -6 Automatic retransmission after 300ms timeout, with a maximum of 3 retries.
[0059] Preferably, in step S6, the real-time monitoring and recalibration of the system conversion efficiency is performed according to the following logic: total efficiency η total The calculation is performed by multiplying the spot matching efficiency, photovoltaic conversion efficiency, MPPT tracking efficiency, and charging efficiency.
[0060] η total =η opt ·η pv ·η mppt ·η chg
[0061] Where, η opt =(ΣG actual ·A cell ) / (ΣG laser ·A laser ) represents the spot matching efficiency, (G) actual For the actual light intensity measured at the receiving end, G laserThe nominal light intensity of the laser emitter, in W / m². 2 A cell A laser These are the effective area of the photovoltaic unit and the laser spot area, respectively, in meters (m²). 2 );η pv =P pv / (G actual ·A cell ·N active Photovoltaic conversion efficiency (P) pv For the output power of the photovoltaic array, N active (Number of activated units); η mppt =P out / P pv For MPPT tracking efficiency (P out (output power of the DC / DC converter); η chg =P bat / P out For charging efficiency (P) bat (Battery input power).
[0062] Preferably, in step S6, the triggering condition and calibration process are determined by judging the cumulative decrease rate of total efficiency over three consecutive monitoring cycles (5 seconds per cycle), and the judgment formula is:
[0063]
[0064] Where t is the index of the current period.
[0065] When the condition is met, the system performs a global recalibration, including:
[0066] (1) Sensor baseline correction: remeasure the light intensity sensor offset ΔG offset via G calibrated =G raw -ΔG offset Update data;
[0067] (2) Algorithm parameter reset: Force the PSO particle swarm position to be restored to the historical optimal solution g. best And clear the experience replay buffer of reinforcement learning;
[0068] (3) Control parameter restoration: Restore PID parameter K p T i T d Reset to the initial calibration value to eliminate integration error.
[0069] Calibration verification: After calibration, if the efficiency recovers to more than 95% of the original level within 2 cycles (i.e., η) total ≥0.95η base η baseIf the efficiency before triggering is within the baseline range, the process is considered successful; otherwise, a secondary fault alarm is triggered, and the system switches to constant voltage safe charging mode (V). safe =0.9×V oc_nom V oc_nom (The nominal open-circuit voltage of the battery) and report a maintenance request.
[0070] Preferably, in step S7, the internal resistance measurement adopts the AC injection method, which excites the battery with a 1kHz / 10mA signal after the battery has been left to stand for 10 minutes, with a measurement accuracy of ±1.5%.
[0071] Preferably, in step S7, the dynamic adjustment of the battery charging curve and the health protection strategy are implemented according to the following rules:
[0072] (1) Charging curve generation: The maximum allowable charging current is calculated based on the joint function of SOC (State of Charge) and SOH (State of Health).
[0073]
[0074] I max The battery's nominal maximum charging current (in A); SOC current Current state of charge (range 0–100%); SOC nom R is the rated SOC (usually taken as 80%). bat R is the current internal resistance of the battery (unit: mΩ); bat_initial This is the initial internal resistance at the factory (unit: mΩ).
[0075] (2) Health status determination: When the internal resistance rises above the threshold, protection is triggered.
[0076]
[0077] When this condition is met, the system will forcibly switch to trickle charging mode, with the charging current limited as follows:
[0078] I trickle =0.1·I max
[0079] At the same time, a maintenance alarm code (code level: CRITICAL) is generated.
[0080] (3) SOH update mechanism: SOH is dynamically corrected based on capacity decay rate and internal resistance change.
[0081]
[0082] Among them, C actual The actual current battery capacity (unit: Ah); C nomNominal capacity (unit: Ah). When SOH ≤ 70%, a battery replacement recommendation command is triggered simultaneously.
[0083] Once the relevant steps are executed, a dynamic MPPT control method for laser wireless power transfer and a reconfigurable array system based on the RL-PSO hybrid algorithm can be realized.
[0084] Compared with the prior art, the advantages of the present invention are as follows:
[0085] (1) Innovatively integrates RL-PSO hybrid algorithm with reconfigurable photovoltaic array technology. Through dynamic parameter space construction (such as multi-dimensional mapping of light intensity-temperature gradient-historical data) and parallel optimization mechanism, it solves the local convergence and response lag problem of traditional MPPT under dynamic light spot, and improves global search efficiency by more than 40%.
[0086] (2) Adaptive topology reconfiguration and revenue quantification decision-making: Based on the switching criterion of RI index (RI>0.15) and the dynamic compensation model for relay loss, the reconfiguration response time is shortened to 10ms while ensuring power continuity, and the light energy utilization rate is increased by 15%-22%.
[0087] (3) Multi-level health perception and fault-tolerant control, combined with battery SOH-SOC joint charging curve adjustment, redundant node bypass switching and efficiency decline triggering global calibration mechanism, enable the system to maintain more than 85% effective output under abnormal operating conditions, and reduce fault recovery time by 60%.
[0088] (4) High-precision real-time data fusion architecture, through light intensity distribution heatmap (100Hz sampling + 3×3 median filtering), three-dimensional sparse matrix compression (60% improvement in storage efficiency) and double-buffered communication protocol (bit error rate <10). -6 It achieves full-process collaborative optimization of data acquisition, processing and command transmission in complex environments, and improves overall energy efficiency by more than 30% compared with traditional systems.
[0089] (5) Deep hardware-algorithm co-design, from laser calibration (Δd closed-loop control), array reconstruction (16-bit topology coding) to charging management (internal resistance ±1.5% accuracy monitoring) full-link optimization, significantly improves system reliability and battery life. Attached Figure Description
[0090] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0091] Figure 1System overall architecture diagram
[0092] Figure 2 Schematic diagram of a reconfigurable photovoltaic array structure
[0093] Figure 3 Flowchart of the RL-PSO hybrid algorithm
[0094] Figure 4 System performance comparison chart under spot drift scenario
[0095] Figure 5 System performance comparison chart under partial occlusion scenarios Detailed Implementation
[0096] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0097] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0098] like Figure 1 As shown, this invention is based on a dynamic MPPT control method and reconfigurable array system for laser wireless power transfer using a hybrid RL-PSO algorithm, and includes the following steps:
[0099] S1. The laser emitting module starts a self-test program to calibrate the focal length of the collimating lens and the beam deflection angle; the photovoltaic array receiving module performs a health diagnosis, detects units with abnormal open-circuit voltage and marks them as redundant nodes.
[0100] In a specific embodiment, when the laser emitting module starts the self-test program in step S1, such as Figure 1 The system architecture shown first collects temperature data of each heat-generating component through a distributed temperature sensor array integrated within the laser cavity. When a temperature deviation ΔT > ±5℃ is detected on the laser diode heat sink substrate, an overheat protection mechanism is triggered, and the system operates with reduced power. The beam quality analyzer uses M... 2 The factor measurement module monitors the laser beam waist parameters in real time, with a sampling frequency of 1kHz and a measurement error ≤2%. During the collimating lens focal length calibration process, the closed-loop control logic for the piezoelectric ceramic actuator displacement Δd is as follows: Figure 3 As shown in the initialization phase of the algorithm flow, according to the formula:
[0101]
[0102] The lens position is dynamically adjusted, where the wavelength λ = 980 nm (typical value), M 2 =1.3, w0=50μm, calibrate the target to make the spot diameter D spot ≤1.5mm. The tilt adjustment of the reflector driven by a two-dimensional servo motor adopts a PID position closed loop, and its feedback signal comes from the spot position sensor at the receiving end, such as... Figure 2 As shown, when the output offsets Δx = 2.3 mm and Δy = 1.8 mm of the four-quadrant detection and tracking mechanism at the array edge, according to the formula:
[0103]
[0104] The deflection angle θ is calculated. x =0.13°, θ y =0.10° (transmission distance L=10m), calibration residual controlled within ±0.02°. During photovoltaic array health diagnosis, if... Figure 2 The reconfigurable array topology shown applies standard test illumination (1000 W / m²) to each photovoltaic unit. 2 The open-circuit voltage Voc of unit A3-B7 is measured by switching via a multiplexer. oc =2.1V (nominal value V) oc_nom When the voltage is 2.4V, it is determined that the attenuation rate has reached 12.5%, which exceeds the 10% threshold. The unit is immediately switched to bypass mode and the redundant node mapping table is updated via RS-485 bus.
[0105] S2. The irradiance sensor network acquires light intensity distribution data at a sampling frequency of 100Hz, eliminates pulse noise through median filtering, and generates a normalized irradiance heat map.
[0106] In a specific embodiment, the data acquisition and processing flow of the irradiance sensor network in step S2 is as follows: Figure 2 The reconfigurable array unit shown consists of a sensor array of 64 photodiodes (8×8 grid) spaced at 5mm intervals. The sensor network synchronously acquires light intensity data at a sampling frequency of 100Hz, and the original irradiance matrix G... raw The spatial resolution of (x,y,t) satisfies the formula:
[0107]
[0108] Among them, D array The effective side length of the receiving array is 15cm, M=8 is the number of horizontal sensors, and Δs is the center-to-center distance between adjacent sensors. For example... Figure 3 The RL-PSO algorithm shown in the diagram first performs a 3×3 sliding window mid-range filter on the raw data during the preprocessing stage. The specific implementation formula is as follows:
[0109]
[0110] Among them, G min =120W / m2, G max =980W / m2 is the measured minimum and maximum irradiance values for the entire field during the current sampling period (unit: W / m2). 2 The normalized result is stored in the form of an 8-bit grayscale matrix. The grayscale value of the central region of the heatmap spot (coordinates (4,5)-(5,6)) reaches 240 (corresponding to 936W / m2), while the grayscale value of the edge region decays to below 80 (corresponding to <310W / m2). The data update cycle is synchronized with the sampling frequency (10ms), and a double buffering mechanism is used to avoid data tearing. Its memory buffer is allocated as a 2×8×8 matrix (current frame and historical frame are written alternately).
[0111] S3, the intelligent control center integrates light intensity distribution, temperature gradient, and historical operating data to construct a three-dimensional optimization parameter space, in which irradiance data is mapped to 0-1000W / m² using 8-bit grayscale values. 2 scope.
[0112] In a specific embodiment, the construction process of the three-dimensional optimization parameter space in step S3 is as follows: Figure 3 As shown, the data fusion stage of the RL-PSO hybrid algorithm is implemented as follows:
[0113] (1) Light intensity distribution data mapping: The 8-bit grayscale matrix Gnorm(x,y) generated in step S2 is converted into the actual irradiance value G. actual (x,y)(Unit: W / m) 2 ):
[0114]
[0115] Among them, G norm (x,y)∈[0,255] represents the normalized grayscale value, 1000W / m 2 Corresponding to full-scale irradiance.
[0116] (2) Temperature gradient field calculation: The surface temperature distribution T(x,y) of the photovoltaic unit is obtained through an infrared thermal imager array, and the temperature difference gradient between adjacent units is calculated.
[0117]
[0118] Wherein, ΔT(x,y), in ℃, represents the temperature difference of a 3×3 region centered at coordinates (x,y), and is used to characterize the risk of hot spots.
[0119] (3) Dynamic weighting of historical data: Extracting MPPT efficiency η from the historical operating database under the same spot distribution pattern.hist Topology configuration coding C hist and battery degradation coefficient β hist Construct time decay weights:
[0120]
[0121] Among them, t now t is the current timestamp. hist For historical time stamps, the weight decreases as the data timeliness index decreases.
[0122] (4) Three-dimensional parameter space synthesis: The above data is projected into a three-dimensional space of spatial coordinates (x,y), irradiance G, and temperature gradient ΔT, and historical weights are fused to generate an optimization objective function:
[0123]
[0124] Where α = 0.6, β = 0.3, and γ = 0.1 are weighting coefficients (satisfying α + β + γ = 1), and ∈ = 0.1℃ is a zero-limiting constant. This space is used for the RL-PSO algorithm to search for the global optimum, while constraining the battery health state (β). hist ≤0.2).
[0125] (5) Data compression and storage: The three-dimensional parameter space is stored in a sparse matrix format, and only F is retained for non-zero elements. optim ≥0.7×F max In high-value areas, storage efficiency is improved by 60%.
[0126] S4. Run the PSO and reinforcement learning algorithms in parallel. Set the PSO population size to 50-100 and the number of iterations to 20-50. The reinforcement learning algorithm uses a double-Q network structure with an experience replay buffer capacity of 10. 4 strip.
[0127] In a specific embodiment, the implementation method of the hybrid optimization algorithm in step S4 is as follows: Figure 3 As shown, the RL-PSO hybrid algorithm employs a dual-channel parallel architecture to achieve coordinated control of global search and local optimization. In specific implementation, the PSO algorithm module initializes 50-100 particles, with each particle's position encoded as a combination vector of the photovoltaic array's MPPT operating voltage Vpv and topology configuration Ctopo. Its velocity update follows a dynamic inertia weighting strategy.
[0128] v i (t+1)=ω(t)v i (t)+c1r1(p best -x i (t))+c2r2(g best -x i (t))
[0129] Where ω(t) is the inertia weight that decreases linearly from 1.2 to 0.4, c1=c2=2.0 are the learning factors, and p best and g best These are the individual historical optimal solution and the global optimal solution, respectively. For example... Figure 2 In the reconfigurable photovoltaic unit structure shown, the fitness function calculation needs to consider both power output and heat loss simultaneously.
[0130] f(x i ) = P pv -0.05·max(ΔT)
[0131] Among them, P pv ΔT represents the output power (in W) under the current topology, and ΔT represents the maximum temperature difference on the array surface (in °C).
[0132] In a specific embodiment, in step S4, the reinforcement learning module adopts a dual-Q network structure, such as... Figure 1 As shown in the decision center, the main network Q main With the target network Q target Receives data including light intensity distribution G(x,y), temperature field T(x,y), and electrical parameter V. pv ,I pv The state vector s outputs topology switching action a and PSO parameter correction instructions. The experience replay buffer capacity is 10. 4 The following steps employ a priority sampling mechanism:
[0133]
[0134] After every 10 PSO iterations, the dual network minimizes the mean square error L = ∑(Q) target -Q main ) 2 / N completes parameter synchronization and update.
[0135] The collaborative control mechanism is specifically manifested as follows: when the PSO convergence speed decreases (power improvement rate < 0.1% for 10 consecutive iterations), the RL network outputs a topology reconstruction command and a PSO parameter reset signal, such as... Figure 3 As shown in the PSO optimization, the mode switching logic forces the inertia weight ω to be restored to 1.0 to restart the global search; when the RL action causes the power P pv When the value exceeds the PSO optimum by 5%, this strategy is injected into the PSO particle swarm optimization as an initial solution to accelerate the hybrid convergence process. This mechanism enables... Figure 4 In the light spot drift scenario shown, the system's power tracking speed is improved by 58% compared to the traditional PSO.
[0136] S5. When the reconstruction benefit index RI>0.15, an instruction packet containing the target topology code, MPPT reference voltage and PID parameters is generated and transmitted to the relay matrix via RS-485 bus.
[0137] In a specific embodiment, the process of determining the reconstructed return index and generating the instruction package in step S5 is as follows: Figure 3 As shown, when the reconstruction benefit index RI calculated by the hybrid algorithm engine is greater than 0.15, the system triggers the topology reconstruction instruction generation mechanism. The calculation of RI is based on the following formula:
[0138]
[0139] Among them, P pred To predict the maximum power of the topology (in W), P current For the output power of the front array, and P switch This refers to the power loss during relay switching.
[0140] like Figure 2 In the reconfigurable array cell shown, the relay power loss is calculated using the following formula:
[0141]
[0142] Where, N sw For the number of operating relays, I k Let t be the conducting current (in A) and resistance (in Ω) of the k-th relay. switch The time for a single switchover is 1 second (unit: s), and the measured value is ≤500 ns.
[0143] The instruction packet encoding rule is topology encoding: a 16-bit binary mask is used to represent the unit connection relationship. The first 8 bits correspond to the number of series units (range 1-255), and the last 8 bits represent the number of parallel branches (range 1-255). For example, "0x0A03" represents 10 series units × 3 parallel branches. MPPT reference voltage: Vref is encoded with 12-bit precision, with a quantization step size of 5mV, corresponding to the formula:
[0144]
[0145] PID parameter: proportional coefficient k p Integration time T i Differential time T d Each is encoded using an 8-bit unsigned integer, and the actual value mapping formula is as follows:
[0146]
[0147] The instruction packets are encapsulated and transmitted via the Modbus-RTU protocol. The data frame structure includes: a 2-byte topology code field (addresses 0x0001-0x0002); a 4-byte floating-point voltage reference value (addresses 0x0003-0x0006); and a 12-byte PID parameter group (addresses 0x0007-0x0012). CRC-16 checksum (polynomial 0xA001) is used, and the measured bus bit error rate is ≤1×10⁻⁶. -6 Automatic retransmission is triggered after a 300ms transmission timeout, with a maximum of 3 retries. During this period, the current topology configuration is maintained to ensure power continuity.
[0148] S6, Real-time monitoring system conversion efficiency η total =η opt ·η pv ·η mppt ·η chg If the efficiency drops by more than 5% for three consecutive cycles, a global parameter recalibration is triggered.
[0149] In a specific embodiment, the system efficiency monitoring and recalibration mechanism in step S6 is implemented as follows:
[0150] like Figure 1 In the system architecture shown, the total conversion efficiency η total Real-time calculation through the efficiency product of multiple modules:
[0151] η total =η opt ·η pv ·η mppt ·η chg
[0152] Where, η opt =(ΣG actual ·A cell ) / (ΣG laser ·A laser ) represents the spot matching efficiency, (G) actual For the actual light intensity measured at the receiving end, G laser The nominal light intensity of the laser emitter, in W / m². 2 A cell A laser These are the effective area of the photovoltaic unit and the laser spot area, respectively, in meters (m²). 2 );η pv =P pv / (G actual ·A cell ·N active Photovoltaic conversion efficiency (P) pv For the output power of the photovoltaic array, N active (Number of activated units); η mppt =P out / Ppv For MPPT tracking efficiency (P out (output power of the DC / DC converter); η chg =P bat / P out For charging efficiency (P) bat (Battery input power).
[0153] The trigger condition is determined using a sliding window mechanism: a global recalibration is triggered when the following formula is met for three consecutive monitoring cycles (5 seconds per cycle):
[0154]
[0155] Where t is the current cycle index. This threshold can effectively capture the sharp drop in efficiency caused by spot drift (decline slope > 1% / s).
[0156] The calibration process is performed according to the following steps:
[0157] (1) Sensor baseline correction: The light intensity sensor performs zero-point calibration, and the offset is calculated as follows:
[0158]
[0159] Updated data G calibrated =G raw -ΔG offset Calibration accuracy reaches ±2W / m 2 .
[0160] (2) Algorithm parameter reset: such as Figure 3 As shown, the PSO particle swarm position is forcibly restored to the historical g. best At the same time, the RL experience replay buffer is cleared to eliminate interference from outdated data.
[0161] (3) Control parameter restoration: PID parameters are reset to their initial calibration values.
[0162]
[0163] During the calibration and verification phase, if η is satisfied for two consecutive cycles... total ≥0.95η base η base If the baseline efficiency is reached, the test is successful; otherwise, the fault handling module is triggered, and the system switches to safe charging mode (V). safe =0.9×V oc_nom The system then reports a maintenance request. Simulations show that this calibration mechanism reduces the system's recovery time under abnormal operating conditions to 18.7 seconds, a 63% improvement over traditional methods.
[0164] S7. The charging curve is dynamically adjusted according to SOC and SOH. When the battery internal resistance is detected to rise by more than 20% of the initial value, the system will force the battery to enter trickle charging mode and issue a maintenance alarm.
[0165] In a specific embodiment, the implementation method of battery health management and charging strategy adjustment in step S7 is as follows: Figure 1 In the energy storage management module shown, dynamic adjustment of the charging curve is achieved through a joint SOC-SOH control algorithm. The specific process includes:
[0166] (1) Segmented charging current control: The multi-stage charging current based on the battery state of charge (SOC) is implemented according to the following rules:
[0167]
[0168] Among them, I max V represents the maximum allowable current (unit: A), α(T) is the temperature compensation coefficient (range 0.9-1.1), and V ocv This is the open-circuit voltage (unit: V). This strategy can reduce charging loss by 23% in partially shaded scenarios.
[0169] (2) State of Health (SOH) Assessment: The battery internal resistance RAC was measured using AC impedance spectroscopy (frequency range 1kHz-10MHz), combined with a capacity decay model.
[0170]
[0171] Where m ranges from 1.2 to 1.5, n ranges from 0.5 to 0.8, and the goodness of fit R0 is... 2 ≥0.95. For example... Figure 1 As shown in the battery monitoring unit, when the cumulative capacity decay ΔC ≥ 20%, the maximum charging current is automatically limited to 0.8 times the original value.
[0172] (3) Maintenance triggering mechanism: Force switching to trickle charging mode when the following conditions are detected:
[0173]
[0174] At this point, the charging current drops to I. trickle =0.1I max and through Figure 1 The communication module sends a CRITICAL level alarm code (code format: 0xAE01).
[0175] (4) Dynamic Update of SOH: The health status assessment value is calculated in real time using the following formula:
[0176]
[0177] Among them, Cactual The actual current battery capacity (unit: Ah); C nom The nominal capacity is expressed in Ah. This model keeps the battery life prediction error within ±5%.
[0178] Once the relevant steps are implemented, a dynamic MPPT control method for laser wireless power transfer and a reconfigurable array system based on the RL-PSO hybrid algorithm can be realized.
[0179] It should be noted that in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0180] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A laser wireless energy receiving system based on dynamic topology optimization, characterized in that, The system includes a laser emitting module, a photovoltaic array receiving module, an intelligent control center, and an energy storage management module. The system optimizes energy transmission through the following technical architecture: The laser emitting module includes a collimating beam generation unit, a spot power PID adjustment unit, and a beam tracking mechanism. Its output is connected to the photovoltaic array receiving module via optical coupling. The collimating beam generation unit uses an aspherical lens group for beam shaping, and the spot power PID adjustment unit dynamically adjusts the laser driving current based on closed-loop feedback control to ensure that the target surface power density meets the following requirements: Among them, P t Target power density (W / cm²) 2 ), η opt For the optical system efficiency (0.6-0.85), P in For laser input power (W), A spot Effective area of the light spot (cm²) 2 The beam tracking mechanism employs a combination of a four-quadrant detector and a piezoelectric ceramic driving mirror to achieve beam alignment accuracy at the sub-milliradian level. The photovoltaic array receiving module consists of a reconfigurable photovoltaic unit array, an irradiance sensor network, and a relay matrix topology switching device. The reconfigurable photovoltaic unit array adopts a regional heterojunction design, with GaAs triple-junction photovoltaic units arranged in the core area and SiC wide-bandgap photovoltaic units arranged in the edge area. Each unit is interconnected in three dimensions through metallized vias. The irradiance sensor network is distributed with a 5mm×5mm grid density to collect spatial light intensity distribution data in real time and transmit it to the intelligent control center. The intelligent control center incorporates a hybrid optimization algorithm engine to perform the following collaborative control: Using light intensity distribution data acquired through the irradiance sensor network, an improved particle swarm optimization (PSO) algorithm is employed to predict the global maximum power point, with the following velocity update formula: Among them, v id Where is the particle velocity, w is the inertia weight (0.4-0.9), c1 and c2 are learning factors (1.2-2.0), λ is the reinforcement learning coefficient (0.1-0.3), and J... RL The reward function is Q-learning; the hybrid optimization algorithm engine synchronously analyzes the ambient temperature, array output voltage and historical operating data, generates array topology reconstruction instructions and MPPT parameter joint optimization strategies, and sends them to the relay matrix through a high-speed digital interface. The energy storage management module includes a multi-stage charging controller and a lithium battery pack. The charging controller dynamically switches between constant current, constant voltage, and trickle charging modes based on the battery's state of charge (SOC), and its charging current satisfies a piecewise function: Among them, I chg Let V be the charging current (A), α(T) be the temperature compensation coefficient (0.9-1.1), and V be the voltage. ocv The open-circuit voltage (V) is used; the controller integrates a state of health (SOH) assessment algorithm to optimize the charging strategy by monitoring the battery's internal resistance and capacity decay rate.
2. The system according to claim 1, characterized in that, The reconfigurable photovoltaic unit array adopts a dual-layer heterogeneous layout design, including: the spacing between GaAs triple-junction photovoltaic units arranged in the core region satisfies d1≤λ / 2NA, where λ is the laser wavelength (808-980nm) and NA is the focusing optical numerical aperture (0.4-0.6). This design effectively reduces optical crosstalk between units; the surface of the SiC photovoltaic units arranged in the edge region is covered with an anti-reflection coating, and its open-circuit voltage compensation coefficient is proportional to the 0.8 power of the edge / core region irradiance ratio, ensuring the stability of the array output voltage; the relay matrix topology switching device is composed of n×m matrix solid-state relays, with each relay node connected in parallel with an RC buffer circuit, and the total reconfiguration response time is less than 500ns, supporting series, parallel and bridged topology mode switching.
3. The system according to claim 1, characterized in that, The workflow of the hybrid optimization algorithm engine includes: an environmental perception phase, which constructs a three-dimensional state space containing an irradiance distribution matrix (G), a temperature field (T), and an array output voltage (V), and updates the environmental dataset every 10ms; and a global search phase, which initializes the particle swarm position matrix X∈R. N×3 Each particle corresponds to a set of MPPT parameter combinations (V ref ,R MPP ,k p ), through the fitness function F=η mppt ·P pv -γ·E switch To evaluate the quality of particles, η mppt To track efficiency, E switch For topology switching energy consumption; during the local optimization phase, an ε-greedy strategy is used to select action a. t ∈{Maintain, Reconstruct, Adjust}, the Q-value update weights are dynamically adjusted by the historical return variance to ensure the co-convergence of PSO and reinforcement learning; in the decision output phase, when the reconstruction return index RI=(ΔP array -E switch ) / (t cycle ·P avg If the value is greater than 0.15, a topology switching command is triggered; otherwise, the current array configuration is maintained.
4. The system according to claim 1, characterized in that, The working logic of the multi-stage charging controller includes: in the constant current charging stage (SOC ≤ 0.2), the maximum allowable current I is used to max rapidly increase the battery voltage while monitoring that the temperature rise rate dT / dt ≤ 0.5 °C / min; in the constant voltage charging stage (0.2 < SOC ≤ 0.8), the charging current is adjusted according to the real-time temperature T, and the compensation formula is I adj = I rated ·[1 - 0.003(T - 25)] to avoid overheating of the battery; in the trickle charging stage (SOC > 0.8), a pulse charging strategy is adopted, and the duty cycle is inversely proportional to the internal resistance of the battery to reduce the capacity loss caused by the polarization effect.
5. A control method for the system according to any one of claims 1-4, characterized in that, Includes the following steps: S1. The laser emitting module initiates a self-test program to calibrate the collimating lens focal length and beam deflection angle; the photovoltaic array receiving module performs a health check, detects units with abnormal open-circuit voltage and marks them as redundant nodes. S2. The irradiance sensor network acquires light intensity distribution data at a sampling frequency of 100Hz, eliminates pulse noise through median filtering, and generates a normalized irradiance heat map. S3, the intelligent control center integrates light intensity distribution, temperature gradient, and historical operating data to construct a three-dimensional optimization parameter space, in which irradiance data is mapped to 0-1000W / m² using 8-bit grayscale values. 2 scope; S4. Run PSO and reinforcement learning algorithms in parallel. Set the PSO population size to 50-100 and the number of iterations to 20-50. Reinforcement learning employs a dual-Q network structure with an experience replay buffer capacity of 10. 4 strip; S5. When the reconstruction yield index RI>0.15, an instruction packet containing the target topology code, MPPT reference voltage and PID parameters is generated and transmitted to the relay matrix via RS-485 bus. S6, Real-time monitoring system conversion efficiency η total =η opt ·η pv ·η mppt ·η chg If the efficiency drops by more than 5% for three consecutive cycles, a global parameter recalibration will be triggered. S7. The charging curve is dynamically adjusted according to SOC and SOH. When the battery internal resistance is detected to rise by more than 20% of the initial value, the system will force the battery to enter trickle charging mode and issue a maintenance alarm.
6. The control method according to claim 5, characterized in that, The parameter configuration of the hybrid optimization algorithm in step S4 includes: the inertia weight w of the PSO algorithm decays linearly with the number of iterations, with an initial value of 0.9 and a final value of 0.4; the exploration rate ∈ of reinforcement learning adopts an exponential decay strategy, with an initial value of 0.5 and a decay coefficient of 0.95; the output weights of the two algorithms are calculated through the variance of the sliding window, with a window size of 10 control cycles.
7. The control method according to claim 5, characterized in that, The encoding rules for the instruction packet in step S5 include: the topology encoding uses a 16-bit binary number, with the first 8 bits representing the number of series units and the last 8 bits representing the number of parallel branches; the MPPT reference voltage V... ref Represented with 12-bit precision, the quantization step size is 5mV; PID parameters (k p ,k i ,k d Each is encoded using an 8-bit unsigned integer, with the actual value mapped to the range of 0.1-10.
0.
8. The control method according to claim 5, characterized in that, The SOH evaluation method in step S7 includes: measuring the battery internal resistance R by AC impedance spectroscopy. AC The sampling frequency is 1kHz-10MHz; the capacity attenuation model is... Where m = 1.2-1.5, n = 0.5-0.8, and the goodness of fit R0 2 ≥0.95; When the cumulative value of ΔC exceeds 20% of the rated capacity, the maximum charging current is automatically reduced to 80% of the rated value.
Citation Information
Cited By
Photovoltaic module self-optimization power generation method and system in weak light scene
CN121578853A
Underwater acoustic sensor network topology optimization method based on reinforcement learning and genetic algorithm
CN121882175A