Reinforcement learning-based electric vehicle range extender amplifier control system and method
By using a reinforcement learning-based range extender amplifier control system, the dynamic responses of the electrical and mechanical links are coordinated, solving the problem of misalignment between the electrical and mechanical links in range extender control. This enables efficient power distribution under complex operating conditions and battery protection under low charge conditions, thereby improving the control performance and fuel economy of the range extender.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU LANYING INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-12
AI Technical Summary
Existing range extender control schemes cannot identify the response misalignment of electrical and mechanical links under dynamic operating conditions, resulting in latent aging of the battery under low charge state and increased ripple stress on the high-voltage DC bus, leading to increased fuel consumption and pollutant emissions, and failing to meet the requirements for improved range extender control performance.
An electric vehicle range extender amplifier control system based on reinforcement learning is adopted. By acquiring the forward traction demand power and battery efficiency, the excitation drag energy density is calculated. Combined with the battery state of charge and excitation suppression degree, the power sharing ratio of the range extender and the excitation suppression strategy are generated, and the dynamic response of the electrical and mechanical links is coordinated to achieve closed-loop control.
It effectively identifies differences in energy sources during power output, coordinates the range extender's power distribution and excitation regulation, reduces power distortion caused by rotor kinetic energy release and short-term battery compensation, reduces additional charging and discharging stress on the battery, improves engine operating efficiency, and reduces pollutant emissions.
Smart Images

Figure CN122186114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric vehicles, and more specifically, to an electric vehicle range extender amplifier control system and method based on reinforcement learning. Background Art
[0002] Range-extended electric vehicles combine the smoothness of pure electric drive and the endurance flexibility of fuel energy replenishment, and have become the mainstream power solution for urban trunk logistics transport vehicles. As the core energy replenishment unit, the control accuracy of the range extender directly determines the vehicle's power performance, economy, and the service life of key components. Current research on range extender control in the industry mostly focuses on upper-layer energy management optimization and steady-state high-efficiency range control, and pays insufficient attention to the response misalignment problem between the electrical and mechanical links under dynamic working conditions.
[0003] The hybrid excitation starter / generator integrated with the range extender can quickly change the output voltage and power by adjusting the excitation current. The response speed of the excitation amplification link is much faster than the establishment speed of the engine combustion torque. In complex working conditions such as continuous climbing and frequent acceleration and deceleration of the vehicle, the battery is often in a low state of charge, and the range extender needs to frequently respond to the rapid changes in power commands. At this time, the rapid adjustment of the excitation current will first lift the electrical output prior to the establishment of mechanical-side power, forming an inherent misalignment where the electrical side responds first and the mechanical side follows later. Existing control schemes mostly take power tracking error, fuel consumption, and state of charge maintenance as the core control objectives, and cannot identify the distortion of the energy source behind the surface power compliance, and are prone to misjudging the pseudo-tracking formed by the release of rotor kinetic energy and short-term battery compensation as effective power output. The long-term existence of this phenomenon will exacerbate the hidden aging of the battery under low state of charge, amplify the ripple stress of the high-voltage DC bus, and at the same time cause the actual working condition of the engine to deviate from the high-efficiency range, increasing fuel consumption and transient pollutant emissions, becoming the key bottleneck restricting the improvement of range extender control performance. Summary of the Invention
[0004] The present invention provides an electric vehicle range extender amplifier control system and method based on reinforcement learning to solve the technical problems in the above background art.
[0005] The present invention provides an electric vehicle range extender amplifier control system based on reinforcement learning, including: The first module acquires the forward traction demand power and battery efficiency; The second module calculates the excitation drag energy borrowing density based on the excitation current change rate, generator deceleration release power, battery compensation power, and battery efficiency; The third module acquires the battery state of charge, synchronously inputs the battery state of charge, forward traction demand power, battery efficiency, and excitation drag energy borrowing density into the reinforcement learning policy network to generate the range extender power sharing ratio and excitation suppression degree, and combines the range extender power sharing ratio and forward traction demand power to obtain the target range extender power; The fourth module retrieves the base speed based on the target range extender power and superimposes the energy borrowing pre-compensation term generated by the excitation drag energy borrowing density to obtain the target engine speed. The fifth module obtains the DC bus reference voltage, generates the basic magnetic flux based on the DC bus reference voltage and the target engine speed, and performs suppression processing on the basic magnetic flux by combining the excitation suppression degree and the excitation drag borrowed energy density to obtain the target synthetic magnetic flux, and then obtains the target excitation current from the target synthetic magnetic flux. The sixth module obtains the target generator electromagnetic torque from the target range extender power and the target engine speed, obtains the target quadrature axis current from the target generator electromagnetic torque and the target composite magnetic flux, obtains the target excitation duty cycle from the target excitation current, and collects and sends the target generator electromagnetic torque, target quadrature axis current and target excitation duty cycle to the power amplifier to complete closed-loop control.
[0006] This invention provides a reinforcement learning-based control method for an electric vehicle range extender amplifier, comprising the following steps: Step S1: Obtain the forward traction power demand and battery efficiency; Step S2: Calculate the excitation drag energy density based on the excitation current change rate, generator deceleration release power, battery compensation power and battery efficiency. Step S3: Obtain the battery state of charge, and simultaneously input the battery state of charge, forward traction power demand, battery efficiency and excitation drag energy density into the reinforcement learning policy network to generate the range extender power sharing ratio and excitation inhibition degree. Combine the range extender power sharing ratio and forward traction power demand to obtain the target range extender power. Step S4: Based on the target range extender power, the base speed is obtained, and the energy borrowing pre-compensation term generated by combining the excitation drag energy borrowing density is superimposed to obtain the target engine speed. Step S5: Obtain the DC bus reference voltage, generate the basic magnetic flux based on the DC bus reference voltage and the target engine speed, and perform suppression processing on the basic magnetic flux by combining the excitation suppression degree and the excitation drag borrowed energy density to obtain the target synthetic magnetic flux, and then obtain the target excitation current from the target synthetic magnetic flux. Step S6: Obtain the target generator electromagnetic torque from the target range extender power and the target engine speed; obtain the target quadrature axis current from the target generator electromagnetic torque and the target composite magnetic flux; obtain the target excitation duty cycle from the target excitation current; collect the target generator electromagnetic torque, the target quadrature axis current and the target excitation duty cycle and send them to the power amplifier to complete closed-loop control.
[0007] The beneficial effects of this invention are as follows: Addressing the misalignment of the time scale between the range extender's excitation link and the engine's mechanical response, this invention constructs excitation drag energy density characteristics, which can identify differences in energy sources during power output, providing a decision-making basis for range extender control that aligns with actual energy supply conditions. By simultaneously completing range extender energy distribution and excitation regulation constraints through reinforcement learning, combined with engine speed pre-compensation and flux suppression logic, the dynamic response rhythm of the electrical and mechanical links can be coordinated, reducing power distortion caused by rotor kinetic energy release and short-term battery compensation. Furthermore, this invention can be deployed based on existing vehicle hardware architecture. In complex operating conditions and low-charge scenarios, it can smooth bus fluctuations during the range extender's energy supply process, reduce additional charging and discharging stress on the battery, and simultaneously make the engine's operating conditions more closely match the preset high-efficiency range, adapting to the daily operating needs of range-extended electric vehicles. Attached Figure Description
[0008] Figure 1 This is a flowchart of the reinforcement learning-based electric vehicle range extender amplifier control method of the present invention; Figure 2 This is a schematic diagram of the computational scenario for the reinforcement learning-based electric vehicle range extender amplifier control of the present invention. Detailed Implementation
[0009] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0010] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of the present invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in one or more embodiments of the present invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" indicate that the element or object preceding the term encompasses the elements or objects listed following the term and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0011] like Figures 1-2 As shown, the reinforcement learning-based electric vehicle range extender amplifier control system includes: The first module obtains the forward traction power demand and battery efficiency; The second module calculates the excitation drag energy density based on the excitation current change rate, generator deceleration release power, battery compensation power and battery efficiency. The third module obtains the battery state of charge. The battery state of charge, forward traction power demand, battery efficiency and excitation drag energy density are simultaneously input into the reinforcement learning policy network to generate the range extender power sharing ratio and excitation inhibition degree. The target range extender power is obtained by combining the range extender power sharing ratio and forward traction power demand. The fourth module retrieves the base speed based on the target range extender power and superimposes the energy borrowing pre-compensation term generated by the excitation drag energy borrowing density to obtain the target engine speed. The fifth module obtains the DC bus reference voltage, generates the basic magnetic flux based on the DC bus reference voltage and the target engine speed, and performs suppression processing on the basic magnetic flux by combining the excitation suppression degree and the excitation drag borrowed energy density to obtain the target synthetic magnetic flux, and then obtains the target excitation current from the target synthetic magnetic flux. The sixth module obtains the target generator electromagnetic torque from the target range extender power and the target engine speed, obtains the target quadrature axis current from the target generator electromagnetic torque and the target composite magnetic flux, obtains the target excitation duty cycle from the target excitation current, and collects and sends the target generator electromagnetic torque, target quadrature axis current and target excitation duty cycle to the power amplifier to complete closed-loop control.
[0012] In one embodiment of the present invention, obtaining the forward traction power demand and battery efficiency includes: Obtain vehicle quality Longitudinal acceleration Gravitational acceleration Rolling resistance coefficient air density air drag coefficient Windward area Speed and road slope angle Calculate the total traction force according to the following formula. : ; in For total traction force, For the overall vehicle quality, For longitudinal acceleration, It is the acceleration due to gravity. The rolling resistance coefficient, air density, The air drag coefficient, For windward area, For vehicle speed, The road slope angle; Obtain the combined efficiency of the drive motor and inverter Calculate the total power demand according to the following formula. : ; in For total power demand, For total traction force, For vehicle speed, To improve the overall efficiency of the drive motor and inverter; The total power demand is calculated according to the following formula. Obtain the positive traction power requirement : ; in For positive traction power demand, Total power demand; Obtain the battery open circuit voltage Battery current and battery equivalent internal resistance And calculate the battery efficiency according to the following formula. : in For battery efficiency, This is the battery open-circuit voltage. Battery current, This is the equivalent internal resistance of the battery.
[0013] It should be noted that the vehicle mass is a mass parameter used to characterize the vehicle's inertia and drag under the current calculation conditions. Longitudinal acceleration is the acceleration parameter of the vehicle along the direction of travel, which can be collected by the vehicle's inertial measurement unit. Gravitational acceleration is a gravity constant used to calculate rolling resistance and gradient resistance, with a preferred value of 9.8, meaning this constant is commonly used in ground vehicle dynamics calculations and meets the accuracy requirements for vehicle control calculations. The rolling resistance coefficient is a dimensionless coefficient characterizing the degree of rolling loss between the tire and the road surface, with a preferred value of 0.008 to 0.015, meaning the rolling resistance of electric vehicle road tires and tires commonly used in urban logistics vehicles on paved roads is typically within this range. Air density is an environmental medium density parameter used to calculate the magnitude of air resistance. The air drag coefficient is a dimensionless coefficient characterizing the aerodynamic drag characteristics of the vehicle's shape, with a preferred value of 0.28 to 0.45, meaning the shape drag of vans and passenger-platform range-extended vehicles is typically within this range. Frontal area is the effective area parameter of the vehicle's frontal projection, preferably between 2.0 and 4.0, meaning the projected area of urban logistics vehicles and medium-sized range-extended electric vehicles typically falls within this range. Vehicle speed is the real-time speed parameter of the vehicle relative to the ground, which can be collected through wheel speed sensors or drive motor speed conversion. Road slope angle is the angle of inclination of the road's longitudinal slope relative to the horizontal plane. Total traction force is the combined traction force required by the vehicle at the current moment to overcome inertial drag, rolling resistance, air resistance, and slope resistance. The combined efficiency of the drive motor and inverter is the combined efficiency parameter of the drive motor and inverter in converting electrical energy into effective driving power at the wheel side at the current operating point. Total power demand is the power required at the drivetrain input end to meet the total traction force demand at the current moment. Zero is a reference constant used to truncate negative demand into non-powered operating conditions, preferably set to 0, meaning the positive traction demand power is only used to characterize the positive power range where the range extender participates in power supply. Forward traction power demand is the forward energy supply power demand obtained by retaining the portion of the total power demand that is greater than or equal to zero. Battery open-circuit voltage is the equivalent voltage parameter reflecting the electrochemical equilibrium state of the battery under equivalent no-load conditions. Battery current is the charging / discharging current parameter of the power battery at the current moment, which can be collected by a battery current sensor. Battery equivalent internal resistance is the equivalent impedance parameter that causes a voltage drop across the current under the current state of charge and temperature conditions. The positive or negative sign of the battery current is a sign state parameter used to distinguish whether the battery is in a discharging or charging state. Battery efficiency is the equivalent efficiency parameter of the battery's available output power or received charging power under current operating conditions relative to the ideal power.
[0014] It should be noted that the internal resistance voltage drop of a battery in the discharging state reduces its output voltage, while the internal resistance voltage drop in the charging state changes the equivalent voltage relationship when the battery absorbs energy. Therefore, by using the battery open-circuit voltage, battery current, and battery equivalent internal resistance, and performing segmented calculations based on the positive and negative states of the battery current, the different power flows under charging and discharging conditions can be distinguished, thus obtaining a battery efficiency consistent with the current operating condition and avoiding the confusion between charging and discharging efficiency using the same calculation relationship. The overall efficiency of the drive motor and inverter is obtained by pre-storing the efficiency calibration relationship between the drive motor and inverter in the vehicle controller, and then looking up the overall efficiency in a table using the operating point corresponding to the current drive motor speed and output torque. When the system has online identification capabilities, the overall efficiency can also be corrected in real time based on the ratio of DC side input power to mechanical side output power, and the correction result is limited and filtered. The road slope angle is obtained by using the longitudinal attitude angle or longitudinal acceleration information output by the inertial measurement unit, combined with the vehicle speed and vehicle stationary identification results to estimate the slope. During the estimation process, the influence of inertial components caused by rapid acceleration and deceleration should be eliminated, and the road slope angle used for traction force calculation should be obtained through low-pass filtering. The battery open-circuit voltage is obtained by the battery management system based on the battery state of charge, temperature, and pre-stored relationships in the battery model. When the vehicle is in the low-current steady-state range, the terminal voltage can be used to correct the estimation result. When the vehicle is in the high-current dynamic range, the estimated value from the previous cycle should be fused with the model output value to obtain the current battery open-circuit voltage. The battery equivalent internal resistance is obtained by pre-stored an equivalent internal resistance calibration table that varies with the battery state of charge and temperature in the battery management system, and then looking up the table based on the current battery state of charge and temperature. When online identification conditions are available, the lookup result can be corrected by combining the changes in terminal voltage and current, and then the corrected value is used for battery efficiency calculation. The positive and negative states of battery current are defined as follows: when the battery outputs current to the DC bus, it is positive; when the DC bus returns current to the battery, it is negative. In the control program, the sign of the battery current should be unified first, and then the battery efficiency should be calculated in segments according to the unified sign, so as to avoid calculation errors caused by inconsistent definitions of charging and discharging directions.
[0015] In one embodiment of the present invention, the excitation drag energy density is calculated based on the excitation current change rate, generator deceleration release power, battery compensation power, and battery efficiency, including: Get the current control window Internal excitation current And the excitation current change rate was obtained. ;in For excitation current, The excitation current change rate, The start time of the current control window. The current control window length; Obtain the generator angular velocity Angular acceleration of generator and the equivalent rotational inertia of the engine and generator coaxially. Calculate the generator deceleration release power according to the following formula. : in To decelerate and release power from the generator, The equivalent moment of inertia of the engine and generator coaxially. The generator's angular velocity, For generator angular acceleration, operator ; Obtain DC bus voltage With battery current And calculate the instantaneous power of the battery according to the following formula. : ; Then extract the battery compensation power according to the following formula. : ; in This refers to the instantaneous power of the battery. To compensate for battery power, This is the DC bus voltage. This refers to the battery current. Power release by decelerating using a generator Battery compensation power and battery efficiency Construct the borrowed power term according to the following formula. : ; in For borrowed power term, For battery efficiency; Based on excitation current With the rate of change of excitation current Construct the excitation effect weight term according to the following formula. : ; in This is the weighting term for the excitation effect. It is a very small positive number; Based on forward traction power demand Current control window length Energy borrowing power item Weighting term related to excitation effect The excitation drag energy density is calculated according to the following formula. : ; in For excitation drag energy density, This is the power required for positive traction.
[0016] It should be noted that the start time of the current control window is the starting time marker of the time window corresponding to the current energy density calculation. The current control window length is the time length parameter used to calculate the integral interval of the energy density, preferably ranging from 0.02 to 0.1, which balances the ability to capture rapid changes in excitation and the ability to suppress sensor noise. The excitation current is the current parameter flowing through the excitation circuit and used to adjust the generator's magnetic field strength, and can be acquired by the excitation circuit current sensor. The excitation current change rate is the parameter of how fast the excitation current changes relative to time. The generator angular velocity is the angular velocity parameter of how fast the generator rotor rotates, and can be acquired by the rotational position sensor or the speed sensor. The generator angular acceleration is the parameter of the rate of change of the generator angular velocity relative to time. The equivalent moment of inertia of the engine and generator coaxial system is the moment of inertia parameter of the engine and generator coaxial system that is dynamically equivalent to the same shaft, preferably ranging from 0.05 to 0.5, that is, the equivalent moment of inertia of the small displacement range extender and the integrated generator coaxial system is usually within this range. Generator deceleration release power is the instantaneous power released outward by the coaxial rotational inertia of the generator during the speed reduction process. DC bus voltage is the current bus voltage parameter on the high-voltage DC bus, which can be acquired by a DC bus voltage sensor. Battery instantaneous power is the instantaneous electrical power formed by the battery current and DC bus voltage at the same moment. Battery compensation power is the positive compensation power extracted from the battery instantaneous power to support the bus power supply. The borrowed energy power term is a comprehensive borrowed energy power index composed of generator deceleration release power and battery compensation power. The excitation effect weight term is a weighted term used to characterize the correlation between the current borrowed energy phenomenon and rapid excitation changes. The minimum positive number is a positive constant used to avoid the denominator being zero and to maintain numerical stability; its preferred value is 0.000001 to 0.001, a range that avoids denominator anomalies without significantly distorting the main calculation results. Excitation drag borrowed energy density is a comprehensive characterization parameter after normalizing the rotational inertia release triggered by rapid excitation changes and the battery compensation intensity.
[0017] It should be noted that during the deceleration process of the generator, the rotational kinetic energy originally stored in the coaxial rotating components is released outward in the form of instantaneous mechanical power. Therefore, by constructing the generator deceleration release power using the generator angular velocity, generator angular acceleration, and the equivalent rotational inertia of the engine and generator coaxially, it is possible to separately identify the portion of energy released by mechanical inertia rather than that established by new fuel power, thus providing a direct basis for subsequent identification of energy-borrowing pseudo-tracking. The instantaneous power of the battery includes both charging and discharging directions. However, this invention focuses on the positive output power that the battery compensates for when the bus power supply is insufficient. Therefore, it is necessary to extract the battery compensation power from the instantaneous battery power, retaining only the positive power portion that supports the bus. This is to avoid incorrectly including regenerative energy absorption or recharging states in the energy-borrowing analysis. The energy borrowing in short-term bus power phenomena is not limited to one source; it may originate from generator deceleration power or battery compensation power. Therefore, combining the generator deceleration power and the battery compensation power (corrected for battery efficiency) into a single energy borrowing power term allows for the characterization of both energy borrowing sources under the same power caliber, thus giving energy borrowing analysis a unified physical meaning. This invention does not treat all energy borrowing phenomena indiscriminately, but specifically identifies pseudo-tracking caused by energy borrowing that is strongly correlated with rapid excitation changes. Therefore, it is necessary to construct an excitation effect weighting term using excitation current and its rate of change. Only when the excitation changes sufficiently rapidly and the current excitation level is sufficient to create significant electromagnetic drag will the energy borrowing power term be significantly amplified and included. Thus, this weighting term can distinguish between general load fluctuations and pseudo-tracking caused by excitation pre-excitation. Furthermore, the borrowed power at a single moment is easily affected by sampling noise and instantaneous disturbances. What the control system really needs to identify is the cumulative borrowed power intensity relative to the forward traction demand energy within a control window. Therefore, by using the window integral of the borrowed power term and the excitation effect weight term, and then dividing by the forward traction demand power and the current control window length to form the excitation drag borrowed power density, the instantaneous phenomenon can be converted into a comparable normalized intensity index. Therefore, this construction method can more stably characterize the degree of pseudo-tracking.
[0018] It should be noted that the excitation current change rate is calculated by reading the excitation current in two adjacent control cycles, subtracting the excitation current of the previous cycle from the excitation current of the next cycle, and dividing by the current control window length. To suppress sampling noise, the original excitation current or change rate result should be processed by first-order filtering or moving average. The generator angular acceleration is calculated by reading the generator angular velocity in two adjacent control cycles, subtracting the generator angular velocity of the previous cycle from the generator angular velocity of the next cycle, and dividing by the current control window length. To reduce the amplification error caused by encoder jitter, the generator angular velocity should be limited and filtered before differential calculation. The equivalent moment of inertia of the coaxial engine and generator is determined by summing the moments of inertia of the engine side, generator side, and related coaxial rotating parts according to the coaxial connection relationship, converted to the same rotating shaft. During the calibration phase, this equivalent moment of inertia can be corrected by speed response tests under known load step, and the corrected value is fixed in the controller. The current control window length is set by using the controller's main cycle period as the minimum time resolution, and determining the window length for energy density calculation based on the excitation current change rate and sensor noise level. When both response speed and stability need to be considered, several main cycle periods can be selected to form a fixed-length window, and this value is fixed after vehicle calibration. The current control window start time is updated by updating the end time of the current control window to the start time of the next control window after each energy density calculation. When using a fixed-step rolling window, the current control window start time should be advanced by the same step length in each control cycle, maintaining a time coverage range consistent with the current control window length. The positive and negative directions of the battery compensation power are determined by first calculating the instantaneous battery power according to the unified sign convention of battery current, then defining the power output from the battery to the DC bus as positive compensation power, and defining the power absorbed by the DC bus from the battery as non-compensation power. In the program implementation, only the battery instantaneous power greater than zero is retained as the battery compensation power. The integral discrete implementation of the excitation drag energy density is to divide the current control window length into several discrete sampling points consistent with the control cycle. At each sampling point, the product of the energy borrowing power term and the excitation effect weight term is calculated. Then, the results of each sampling point are summed using rectangular integration or trapezoidal integration. Finally, the result is divided by the positive traction demand power and the current control window length to form the discrete excitation drag energy density.
[0019] In one embodiment of the present invention, the battery state of charge (SOC) is obtained, and the SOC, forward traction power demand, battery efficiency, and excitation drag energy density are simultaneously input into a reinforcement learning policy network to generate the range extender power sharing ratio and excitation inhibition degree. The target range extender power is obtained by combining the range extender power sharing ratio and the forward traction power demand. Get the current control time Battery state of charge ;in The battery is in its state of charge. According to the battery state of charge Forward traction power demand Battery efficiency With excitation drag borrowed energy density Construct the strategy input vector according to the following formula. : ; in For the policy input vector, The battery is in its state of charge. For positive traction power demand, For battery efficiency, Energy density is borrowed for excitation drag; policy input vector Input the reinforcement learning policy into the network and generate action outputs according to the following formula. : ; in For action output, To strengthen the learning strategy network, To enhance the network parameters of the learning policy network, The power sharing ratio of the range extender, This refers to the excitation suppression degree; Based on the power sharing ratio of the range extender With forward traction power demand Calculate the target range extender power according to the following formula. : ; in For the target range extender power, The power sharing ratio of the range extender. This is the power required for positive traction.
[0020] It should be noted that the battery state of charge (SOC) is a state parameter characterizing the proportion of remaining usable charge in the power battery. The current control time is the current point in time corresponding to a single decision-making inference by the reinforcement learning policy. The policy input vector is a set of states input to the reinforcement learning policy network after combining multiple state parameters in a fixed order. The action output is the control decision result given by the reinforcement learning policy network at the current control time based on the policy input vector. The network parameters of the reinforcement learning policy network are a set of parameters characterizing the connection weights and bias relationships within the network. The range extender power sharing ratio is a parameter characterizing the proportion of forward traction demand power borne by the range extender. The excitation suppression degree is a control parameter characterizing the strength of the current control strategy's suppression of excitation enhancement behavior. The target range extender power is the target power supplied by the range extender calculated based on the range extender power sharing ratio and the forward traction demand power. Battery state of charge reflects remaining power constraints, forward traction power demand reflects current energy demand, battery efficiency reflects battery compensation costs, and excitation drag energy density reflects false tracking risk. Therefore, simultaneously inputting these four types of information into the reinforcement learning policy network allows the policy to perceive energy state, power demand, compensation costs, and borrowing risks simultaneously, rather than making decisions based solely on a single power gap. This input method is therefore more suitable for controlling the range extender amplifier. Furthermore, simply adjusting the range extender power sharing ratio only determines how much power the range extender handles and cannot directly constrain the fast response behavior of the excitation link. Conversely, simply suppressing excitation cannot complete the overall vehicle energy distribution. Therefore, by having the reinforcement learning policy network simultaneously output the range extender power sharing ratio and excitation suppression level, both the energy distribution layer and the electromagnetic execution layer can be adjusted simultaneously. This dual-action approach can directly incorporate false tracking management into the policy decision-making closed loop.
[0021] It should be noted that the battery state of charge (SOC) is obtained by the battery management system (BMS) combining coulomb integral results, terminal voltage correction results, and temperature conditions to estimate the SOC, and then sending the current SOC to the vehicle controller in each control cycle. When the vehicle is stationary for an extended period, the SOC can be recalibrated using the stationary voltage. The reinforcement learning policy network has an input layer with the same number of nodes as the policy input vector dimension, and a two-node output layer, corresponding to the range extender power sharing ratio and excitation suppression degree, respectively. Several fully connected hidden layers (e.g., three layers) are used in the intermediate layers, and bounded activation is set before the output layer to ensure the output falls within a predetermined range. It is also explicitly stated that the network only performs forward inference on the vehicle side and does not perform online training there. The reinforcement learning policy network is trained by constructing simulation conditions covering different battery SOCs, different forward traction power requirements, and different excitation drag energy densities in an offline environment. The reinforcement learning algorithm repeatedly interacts with the environment to update the network parameters. After training, the network parameters are fixed and deployed to the controller, retaining only the online inference function. The correspondence between the action output and the range extender power sharing ratio and the excitation suppression degree is defined as follows: the first component of the action output corresponds to the range extender power sharing ratio, and the second component corresponds to the excitation suppression degree. In the control program, the action output is unpacked in a fixed order, and the two components are pruned according to their respective ranges before being included in subsequent calculations. Furthermore, the range of values for the range extender power sharing ratio and the excitation suppression degree is limited to between 0 and 1. When the output of the reinforcement learning policy network exceeds this range, it is directly pruned according to the upper and lower limits, thereby ensuring that the target range extender power and flux suppression terms always have clear physical meaning.
[0022] In one embodiment of the present invention, the target engine speed is obtained by looking up the base speed based on the target range extender power and superimposing a pre-compensation term for borrowed energy generated by combining the excitation drag borrowed energy density, including: Based on the target range extender power Find the base speed And expressed as: ; in Based on the base speed, Let be the mapping function from the target range extender power to the base speed. Target range extender power; Obtaining excitation drag energy density Forward traction power demand Equivalent moment of inertia of engine and generator coaxial axis angular velocity of generator and the current control window length And calculate the borrowing pre-compensation item according to the following formula. : ; in To borrow energy pre-compensation items, For excitation drag energy density, For positive traction power demand, The current control window length, The equivalent moment of inertia of the engine and generator coaxially. The generator's angular velocity, It is a very small positive number; Based on the base speed With borrowing pre-compensation items The target engine speed is obtained according to the following formula. : ; in The target engine speed. Based on the base speed, This is a pre-compensation item for borrowing energy.
[0023] It should be noted that the base speed is the engine's base target speed obtained solely from the target range extender power. The mapping function from target range extender power to base speed is a preset correspondence that converts the target range extender power into the base speed. The borrowed energy pre-compensation term is a speed feedforward compensation amount calculated based on the excitation drag borrowed energy density, forward traction demand power, and coaxial rotational inertia state. The target engine speed is the final target engine speed obtained by superimposing the base speed and the borrowed energy pre-compensation term. The target range extender power cannot be directly issued as an engine execution command; in actual control, it needs to be converted into an executable target speed reference for the engine. Therefore, obtaining the base speed based on the target range extender power can convert the power target into an executable reference quantity on the engine speed side. Thus, this mapping method is an important intermediate step from power decision-making to mechanical execution. Furthermore, the essence of pseudo-tracking is that the excitation link responds first while the mechanical side is insufficiently prepared. Therefore, using only the base speed will still result in a lag in engine-side power establishment. By superimposing the pre-compensation term for borrowed energy generated by the excitation drag borrowed energy density onto the base speed, the engine target speed can be raised in advance when the borrowed energy risk increases. This method essentially establishes power readiness in advance on the mechanical side to shorten the time misalignment between excitation pre-positioning and mechanical response.
[0024] It should be noted that the mapping function from the target range extender power to the base speed is established during the calibration phase by creating a correspondence table between the target range extender power and the base speed based on the stable operating speed requirements of the range extender under different output power levels. During operation, the base speed is obtained by looking up or interpolating the target range extender power in this correspondence table, and boundary protection is implemented for the lookup results. The limiting method for the borrowed energy pre-compensation term is to set upper and lower limits for the borrowed energy pre-compensation term based on the engine's maximum allowable rate of change of speed and maximum safe speed. When the borrowed energy pre-compensation term calculated by the formula exceeds the upper limit, the upper limit is used; when it is lower than the lower limit, the lower limit is used, to avoid sudden changes in the target engine speed caused by abnormal spikes in borrowed energy density. The smoothing constraint method for the target engine speed is to perform slope limiting and first-order smoothing on the target engine speed after superimposing the base speed and the borrowed energy pre-compensation term. When the target engine speed changes too much in adjacent cycles, the rate of change of speed constraint is satisfied first, and then the smoothed result is sent to the engine control unit.
[0025] In one embodiment of the present invention, a DC bus reference voltage is obtained, a basic magnetic flux is generated based on the DC bus reference voltage and the target engine speed, and a target synthetic magnetic flux is obtained by performing suppression processing on the basic magnetic flux in combination with the excitation suppression degree and the excitation drag borrowed energy density. Then, a target excitation current is obtained from the target synthetic magnetic flux, including: Obtain DC bus reference voltage and generator back electromotive force coefficient And based on the target engine speed The fundamental magnetic flux is generated according to the following formula. : ; in The fundamental magnetic flux, This is the DC bus reference voltage. This is the generator back EMF coefficient. The target engine speed. It is a very small positive number; According to excitation suppression degree With excitation drag borrowed energy density Construct the flux suppression term according to the following formula. : ; in This is a flux suppression term. For excitation suppression degree, For excitation drag energy density, It is a natural exponential function; Based on the fundamental magnetic flux With flux suppression term The target composite magnetic flux is obtained according to the following formula. : ; in To synthesize the target magnetic flux, The fundamental magnetic flux, This is a flux suppression term; Obtaining magnetic flux from a permanent magnet substrate Mapping coefficient between excitation current and additional magnetic flux And synthesize magnetic flux according to the target The target excitation current is obtained according to the following formula. : ; in For the target excitation current, To synthesize the target magnetic flux, For permanent magnet substrate magnetic flux. This is the mapping coefficient from the excitation current to the additional magnetic flux.
[0026] It should be noted that the DC bus reference voltage is a target voltage preset by the control system to maintain DC bus stability, preferably ranging from 350 to 750, meaning the DC bus operating voltage of passenger cars and light commercial range-extended high-voltage platforms is typically within this range. The generator back EMF coefficient is a motor parameter characterizing the proportional relationship between generator speed and back EMF, preferably ranging from 0.05 to 0.5, meaning the back EMF proportionality coefficient of automotive generators within their commonly used speed range is typically within this range. The basic magnetic flux is the unsuppressed magnetic flux calculated based on the DC bus reference voltage and the target engine speed. The flux suppression term is a flux suppression factor constructed using excitation suppression degree and excitation drag borrowed energy density. The target synthetic magnetic flux is the final target magnetic flux formed after adjusting the basic magnetic flux using the flux suppression term. The permanent magnet base flux is the fixed magnetic flux provided by the permanent magnet itself without additional excitation, preferably ranging from 0.02 to 0.2, meaning the base flux corresponding to common magnet sizes and magnetic circuit designs in automotive permanent magnet generators is typically within this range. The mapping coefficient from excitation current to additional magnetic flux is a proportional parameter characterizing the relationship between changes in excitation current and changes in additional magnetic flux. It is preferably set between 0.001 and 0.02, meaning that the sensitivity of additional magnetic flux generated by automotive excitation windings under common magnetic circuit structures is typically within this range. The target excitation current is the target current that needs to be applied to the excitation circuit to achieve the target synthetic magnetic flux.
[0027] It should be noted that the generator output voltage is related to both magnetic flux and engine speed. Once the target engine speed is determined, a corresponding magnetic flux reference is needed to meet the DC bus reference voltage. Therefore, a basic magnetic flux is generated based on the DC bus reference voltage and the target engine speed, transforming the target bus voltage into a generator magnetic flux target. Magnetic flux suppression needs to change continuously with the excitation suppression degree and the excitation drag energy density, while avoiding abrupt discontinuous control. Therefore, a magnetic flux suppression term is constructed using a natural exponential function combined with the excitation suppression degree and the excitation drag energy density. This allows for continuous reduction of the target magnetic flux when the energy borrowing risk increases and smooth recovery when the risk decreases. Furthermore, the target synthesized magnetic flux is an electromagnetic target quantity, while the excitation circuit ultimately executes a current quantity. Therefore, the target excitation current needs to be derived from the target synthesized magnetic flux based on the mapping relationship between the permanent magnet base magnetic flux, the excitation current, and the additional magnetic flux. This allows the magnetic flux target to be converted into a practically executable excitation current target. Therefore, this inverse calculation method is a crucial step in implementing magnetic flux control into excitation execution.
[0028] It should be noted that the DC bus reference voltage is set by pre-calibrating a target value based on the rated operating voltage of the high-voltage battery, the allowable DC bus operating range of the drive system, and the voltage regulation capability of the generator rectifier side. Segmented reference values can also be set under different operating conditions, but a smooth transition should be ensured during switching. The generator back EMF coefficient is obtained by directly converting the design parameters when they are known. Alternatively, during the prototype stage, the back EMF at different speeds can be measured through no-load speed-up tests and fitted to obtain the generator back EMF coefficient. The fitting result is then calibrated and fixed into the controller. The flux suppression term is calibrated by first determining the target intensity of the excitation suppression degree and the excitation drag energy density, then normalizing and limiting the input of the natural exponential function to ensure that the flux suppression term is within a reasonable range across all operating conditions. When the energy density is low, the flux suppression term should be close to 1; when the energy density is high, the flux suppression term should decrease monotonically but not drop to zero. The permanent magnet base flux is obtained by determining the magnetic circuit design value when the motor design parameters are known. In the prototype stage, it can also be obtained through back-calculation using an open-circuit back EMF test combined with speed measurement. After obtaining the value, it should be saved as a fixed parameter in the controller and used for back-calculation of the target excitation current. The mapping coefficient from excitation current to additional magnetic flux is obtained by changing the excitation current and measuring the corresponding change in the synthetic magnetic flux in the prototype stage, and fitting the coefficient based on the ratio of the change in magnetic flux to the change in excitation current. When there is significant nonlinearity in the magnetic circuit, a segmented calibration method should be used, setting the mapping coefficient separately in different excitation intervals. The limiting method for the target excitation current is to trim the upper and lower limits of the target excitation current obtained by back-calculation of the target synthetic magnetic flux based on the maximum allowable excitation current and the minimum controllable excitation current of the excitation circuit. When the back-calculation result exceeds the allowable range, the boundary value should be directly taken, and a slope limit can be set for the change in target excitation current in adjacent cycles.
[0029] In one embodiment of the present invention, the target generator electromagnetic torque is obtained from the target range extender power and the target engine speed; the target quadrature-axis current is obtained from the target generator electromagnetic torque and the target composite magnetic flux; the target excitation duty cycle is obtained from the target excitation current; and the target generator electromagnetic torque, the target quadrature-axis current, and the target excitation duty cycle are aggregated and sent to the power amplifier to complete closed-loop control, including: Based on the target range extender power With the target engine speed The electromagnetic torque of the target generator is calculated using the following formula. : ; in For the target generator electromagnetic torque, For the target range extender power, The target engine speed, It is a very small positive number; Obtaining generator torque coefficient And based on the target generator electromagnetic torque Combined magnetic flux with target The target quadrature-axis current is obtained according to the following formula. : ; in For the target quadrature-axis current, For the target generator electromagnetic torque, This is the generator torque coefficient. To synthesize the target magnetic flux, It is a very small positive number; Obtain the maximum allowable excitation current of the excitation circuit. And according to the target excitation current The target excitation duty cycle is obtained according to the following formula. : ; in For the target excitation duty cycle, For the target excitation current, This is the maximum allowable excitation current in the excitation circuit; Electromagnetic torque of the target generator Target quadrature axis current Duty cycle relative to target excitation Aggregated into execution control results And expressed as: ; in To execute control results, For the target generator electromagnetic torque, For the target quadrature-axis current, The target excitation duty cycle; Execution control results The control is then sent to the power amplifier to complete the closed-loop control.
[0030] It should be noted that the target generator electromagnetic torque is the generator electromagnetic torque required to be output at the target engine speed to achieve the target range extender power. The generator torque coefficient is a motor parameter characterizing the generator current and magnetic flux's ability to generate electromagnetic torque, preferably ranging from 0.1 to 1.5, meaning the equivalent torque coefficient of an automotive generator under different pole pairs and winding structures is typically within this range. The target quadrature-axis current is the control current component used to generate the target generator electromagnetic torque under the target combined magnetic flux condition. The maximum allowable excitation current of the excitation circuit is the upper limit of the current that the excitation circuit can continuously or briefly output under thermal and device constraints, preferably ranging from 3 to 15, meaning the maximum operating current of the automotive excitation drive circuit and excitation winding within the thermal design allowable range is typically within this range. The target excitation duty cycle is the excitation drive duty cycle command calculated to achieve the target excitation current. The executed control result is the output control result formed by aggregating the target generator electromagnetic torque, target quadrature-axis current, and target excitation duty cycle.
[0031] It should be noted that, due to the definite physical relationship between power, speed, and torque, once the target range extender power and target engine speed are determined, the target electromagnetic torque that the generator needs to output is also determined. Therefore, obtaining the target generator electromagnetic torque from the target range extender power and target engine speed can transform the power supply target into a torque target on the motor side. This method is the fundamental relationship for the transition from power control to electromagnetic execution. Since the formation of the target generator electromagnetic torque in AC motor vector control depends on the target composite magnetic flux determining the magnitude of the current component, obtaining the target quadrature-axis current from the target generator electromagnetic torque and target composite magnetic flux can further transform the torque target into a current target that the controller can directly adjust. This method completes the execution mapping from the torque layer to the current layer. Since the power amplifier ultimately receives the control quantity for the on-time of the switching devices, rather than the abstract excitation current target, the target excitation duty cycle needs to be obtained based on the target excitation current and the maximum allowable excitation current of the excitation circuit. In this way, the target excitation current can be converted into a duty cycle command that the excitation drive circuit can execute. This method achieves the end-stage execution mapping of the excitation target. Since the target generator electromagnetic torque, target quadrature shaft current, and target excitation duty cycle correspond to mechanical power execution, current execution, and excitation execution, respectively, only by aggregating these three types of results within the same control cycle and sending them to the power amplifier can a complete execution control result be formed. This method can ensure that the range extender amplifier control system coordinates the torque link and excitation link at the same time, thereby completing closed-loop control.
[0032] It should be noted that the generator torque coefficient is obtained by establishing a theoretical value based on the generator structural parameters, and then correcting the theoretical value through the correspondence between current and torque in bench tests. After correction, the generator torque coefficient is fixed into the controller for the calculation of the target quadrature-axis current. The coordinate system corresponding to the target quadrature-axis current is established by obtaining the rotor electrical angle position through a rotary position sensor, establishing a coordinate system that rotates synchronously with the rotor magnetic field in the controller, and transforming the three-phase current into this coordinate system. Subsequently, the target quadrature-axis current is used as the quadrature-axis current command in this coordinate system to participate in the current closed-loop control. The maximum allowable excitation current of the excitation circuit is set by comprehensively determining an upper limit value based on the allowable temperature rise of the excitation winding, the current capability of the power amplifier devices, and the power supply capacity. When the system has thermal protection function, it can be further dynamically derated based on the real-time temperature on the basis of this upper limit value. The target excitation duty cycle is limited to between 0 and 1, based on the target excitation current. When there is a minimum on-time constraint on the device, a minimum non-zero duty cycle and a maximum safe duty cycle should also be set to prevent the switching device from entering the uncontrollable region. The feedback method for sending the control results to the power amplifier to complete the closed-loop control is as follows: after the power amplifier executes the target generator electromagnetic torque, target quadrature-axis current, and target excitation duty cycle, it collects the actual current, actual voltage, and actual speed in real time as feedback quantities. The controller then calculates the deviation between the actual values and the target values and updates the execution control results in the next control cycle, thus forming a complete closed-loop control.
[0033] In one embodiment of the present invention, the electric vehicle range extender amplifier control method based on reinforcement learning includes the following steps: Step S1: Obtain the forward traction power demand and battery efficiency; Step S2: Calculate the excitation drag energy density based on the excitation current change rate, generator deceleration release power, battery compensation power and battery efficiency. Step S3: Obtain the battery state of charge. Simultaneously input the battery state of charge, forward traction power demand, battery efficiency, and excitation drag energy density into the reinforcement learning policy network to generate the range extender power sharing ratio and excitation inhibition degree. Combine the range extender power sharing ratio and forward traction power demand to obtain the target range extender power. Step S4: Based on the target range extender power, the base speed is obtained, and the energy borrowing pre-compensation term generated by combining the excitation drag energy borrowing density is superimposed to obtain the target engine speed. Step S5: Obtain the DC bus reference voltage, generate the basic magnetic flux based on the DC bus reference voltage and the target engine speed, and perform suppression processing on the basic magnetic flux by combining the excitation suppression degree and the excitation drag borrowed energy density to obtain the target synthetic magnetic flux, and then obtain the target excitation current from the target synthetic magnetic flux. Step S6: Obtain the target generator electromagnetic torque from the target range extender power and the target engine speed; obtain the target quadrature axis current from the target generator electromagnetic torque and the target composite magnetic flux; obtain the target excitation duty cycle from the target excitation current; collect the target generator electromagnetic torque, the target quadrature axis current and the target excitation duty cycle and send them to the power amplifier to complete closed-loop control.
[0034] Specifically, the control method of this invention is deployed in the vehicle controller and range extender controller of the range-extended electric vehicle. After the reinforcement learning strategy network completes offline training, the solidified network parameters are written into the storage unit of the corresponding controller. After the vehicle is powered on, the controller performs full-process calculations according to a fixed control cycle, which is consistent with the main control cycle of the vehicle's power system, typically 10 to 50 milliseconds. The data acquisition process relies on the vehicle's existing onboard sensors and network architecture, without the need for additional acquisition equipment. During vehicle operation, wheel speed sensors acquire vehicle speed signals, inertial measurement units acquire longitudinal acceleration and road gradient signals, the battery management system acquires battery state of charge, battery open-circuit voltage, battery current, and battery equivalent internal resistance signals, the engine and generator controllers acquire engine speed, generator angular velocity, and excitation current signals, the high-voltage power distribution system acquires DC bus voltage signals, and the drive motor controller acquires drive motor efficiency signals. All acquired signals are transmitted to the controller executing the control method through the vehicle controller's local area network. The controller filters and synchronizes the received signals, and then completes the corresponding calculations and generates control commands according to preset steps. Each control cycle completes a full process of state acquisition, feature calculation, strategy reasoning, instruction generation and issuance, forming a continuous closed-loop control.
[0035] It should be noted that the final output of each control cycle of this invention consists of five sets of quantified control parameters that can be directly received by the vehicle's underlying actuators. These parameters are the target range extender power supply, target engine angular velocity, target excitation current, target quadrature axis current, and target excitation pulse width modulation duty cycle. All output parameters are directly executable physical quantities and can be sent to the corresponding execution units without additional conversion. For example, a 4.5-ton rated gross vehicle is a range-extended electric van used for urban trunk logistics. It is fully loaded and traveling on a continuous uphill section of an urban elevated road. The current speed is 50 km / h, the road gradient is 3 degrees, the battery charge is 22% (low charge range), and the vehicle's forward traction power requirement is 65 kW. Under this condition, after the controller completes the full-process calculation, the specific output results within one control cycle may include: target range extender power supply of 62 kW, target engine angular velocity of 314 radians per second (corresponding to an engine speed of 3000 rpm), target excitation current of 5 amps, target quadrature axis current of 180 amps, and target excitation pulse width modulation duty cycle of 0.42. After the output results are generated, they are distributed to the corresponding execution units via the vehicle controller local area network. The target engine angular velocity is sent to the engine control unit to adjust the engine throttle and fuel injection logic, establishing the corresponding mechanical power. The target excitation current and target excitation pulse width modulation duty cycle are sent to the excitation power amplifier drive unit to adjust the power supply parameters of the generator excitation winding, establishing the corresponding magnetic field. The target quadrature-axis current is sent to the generator vector controller to adjust the generator three-phase current output, achieving target electromagnetic torque control. All execution units execute synchronously upon receiving the commands, completing the range extender power supply control for the current cycle.
[0036] Specifically, at the vehicle operation level, this invention effectively improves the power response accuracy of the range extender under complex low-battery conditions, reduces bus voltage fluctuations and current ripples caused by rapid excitation regulation, enhances the operational stability of the high-voltage power supply system, and avoids speed fluctuations caused by misalignment of engine and generator dynamic responses. It also reduces vibration and noise levels during vehicle operation, improving the passenger experience. At the battery protection level, this invention reduces high-frequency short-term compensation discharge of the power battery under low-battery conditions, reduces the additional stress on the battery caused by non-stationary ripples, delays battery capacity decay and performance degradation, extends the lifespan of the power battery, and reduces battery maintenance and replacement costs throughout the vehicle's lifecycle. Furthermore, this invention is particularly suitable for typical operating conditions of urban trunk logistics vehicles, such as frequent start-stop cycles, continuous uphill and downhill driving, and slow-moving traffic congestion. It can maintain the stable operation of the range extender power supply system even under low-battery conditions, ensuring the continuous operation capability of logistics vehicles, reducing power limitations caused by low battery charge, and improving vehicle uptime efficiency and operational reliability.
[0037] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0038] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.
Claims
1. A reinforcement learning-based control system for an electric vehicle range extender amplifier, characterized in that, include: The first module obtains the forward traction power demand and battery efficiency; The second module calculates the excitation drag energy density based on the excitation current change rate, generator deceleration release power, battery compensation power and battery efficiency. The third module obtains the battery state of charge. The battery state of charge, forward traction power demand, battery efficiency and excitation drag energy density are simultaneously input into the reinforcement learning policy network to generate the range extender power sharing ratio and excitation inhibition degree. The target range extender power is obtained by combining the range extender power sharing ratio and forward traction power demand. The fourth module retrieves the base speed based on the target range extender power and superimposes the energy borrowing pre-compensation term generated by the excitation drag energy borrowing density to obtain the target engine speed. The fifth module obtains the DC bus reference voltage, generates the basic magnetic flux based on the DC bus reference voltage and the target engine speed, and performs suppression processing on the basic magnetic flux by combining the excitation suppression degree and the excitation drag borrowed energy density to obtain the target synthetic magnetic flux, and then obtains the target excitation current from the target synthetic magnetic flux. The sixth module obtains the target generator electromagnetic torque from the target range extender power and the target engine speed, obtains the target quadrature axis current from the target generator electromagnetic torque and the target composite magnetic flux, obtains the target excitation duty cycle from the target excitation current, and collects and sends the target generator electromagnetic torque, target quadrature axis current and target excitation duty cycle to the power amplifier to complete closed-loop control.
2. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 1, characterized in that, The total traction force is calculated using the vehicle's mass, longitudinal acceleration, gravitational acceleration, rolling resistance coefficient, air density, air drag coefficient, frontal area, vehicle speed, and road slope angle. The total power demand is calculated by combining the total traction force, vehicle speed, and the combined efficiency of the drive motor and inverter.
3. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 2, characterized in that, The maximum value between the total demand power and zero is taken as the positive traction demand power. The battery efficiency is obtained by using the battery open-circuit voltage, battery current, and battery equivalent internal resistance, and by performing segmented calculations based on the positive and negative states of the battery current.
4. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 1, characterized in that, Obtain the excitation current within the current control window and get the excitation current change rate; Obtain the generator angular velocity, generator angular acceleration, and equivalent moment of inertia of the engine and generator coaxially, and calculate the generator deceleration release power; Obtain the DC bus voltage and battery current, calculate the instantaneous battery power, and extract the battery compensation power from the instantaneous battery power.
5. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 4, characterized in that, Construct an energy borrowing power term by utilizing the generator deceleration release power, battery compensation power, and battery efficiency; The excitation effect weight term is constructed using the excitation current and the excitation current change rate, and the excitation drag energy density is calculated by combining the positive traction demand power, the current control window length, the minimum positive number, the borrowed energy power term, and the excitation effect weight term.
6. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 1, characterized in that, Obtain the battery state of charge at the current control moment; construct the strategy input vector using the battery state of charge, forward traction power demand, battery efficiency, and excitation drag energy density; The policy input vector is input into the reinforcement learning policy network to generate action output, and the range extender power sharing ratio and excitation inhibition degree are obtained from the action output. The target range extender power is calculated by the range extender power sharing ratio and the forward traction power demand.
7. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 1, characterized in that, Establish a mapping relationship between the target range extender power and the base speed, and look up the base speed based on the target range extender power; The excitation drag energy density, forward traction power demand, equivalent rotational inertia of the engine and generator coaxial, generator angular velocity and current control window length are obtained, and the energy borrowing pre-compensation term is calculated using the smallest positive number. The target engine speed is obtained by superimposing the base speed with the borrowed energy pre-compensation term.
8. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 1, characterized in that, Obtain the DC bus reference voltage and generator back EMF coefficient, and calculate the basic magnetic flux using the target engine speed and the minimum positive number; A flux suppression term is constructed by combining the natural exponential function with the excitation suppression degree and the excitation drag borrowed energy density; The target composite magnetic flux is obtained by using the fundamental magnetic flux and the magnetic flux suppression term; Obtain the mapping coefficients from the permanent magnet substrate magnetic flux and excitation current to the additional magnetic flux, and use the target synthetic magnetic flux to obtain the target excitation current.
9. The electric vehicle range extender amplifier control system based on reinforcement learning according to claim 1, characterized in that, The electromagnetic torque of the target generator is calculated using the target range extender power, the target engine speed, and the smallest positive number. Obtain the generator torque coefficient, and calculate the target quadrature-axis current using the target generator electromagnetic torque, the target synthetic magnetic flux, and the minimum positive number; Obtain the maximum allowable excitation current of the excitation circuit, and calculate the target excitation duty cycle using the target excitation current; The target generator electromagnetic torque, target quadrature shaft current, and target excitation duty cycle are aggregated into execution control results, and the execution control results are sent down to the power amplifier.
10. A reinforcement learning-based control method for an electric vehicle range extender amplifier, characterized in that, Implementing the reinforcement learning-based electric vehicle range extender amplifier control system as described in any one of claims 1 to 9 includes the following steps: Step S1: Obtain the forward traction power demand and battery efficiency; Step S2: Calculate the excitation drag energy density based on the excitation current change rate, generator deceleration release power, battery compensation power and battery efficiency. Step S3: Obtain the battery state of charge, and simultaneously input the battery state of charge, forward traction power demand, battery efficiency and excitation drag energy density into the reinforcement learning policy network to generate the range extender power sharing ratio and excitation inhibition degree. Combine the range extender power sharing ratio and forward traction power demand to obtain the target range extender power. Step S4: Based on the target range extender power, the base speed is obtained, and the energy borrowing pre-compensation term generated by combining the excitation drag energy borrowing density is superimposed to obtain the target engine speed. Step S5: Obtain the DC bus reference voltage, generate the basic magnetic flux based on the DC bus reference voltage and the target engine speed, and perform suppression processing on the basic magnetic flux by combining the excitation suppression degree and the excitation drag borrowed energy density to obtain the target synthetic magnetic flux, and then obtain the target excitation current from the target synthetic magnetic flux. Step S6: Obtain the target generator electromagnetic torque from the target range extender power and the target engine speed; obtain the target quadrature axis current from the target generator electromagnetic torque and the target composite magnetic flux; obtain the target excitation duty cycle from the target excitation current; collect the target generator electromagnetic torque, the target quadrature axis current and the target excitation duty cycle and send them to the power amplifier to complete closed-loop control.