A method, system, device and medium for dual-motor coupled drive of electric vehicles
By combining a deep Q-network and a deep deterministic policy gradient network, the energy distribution problem of dual-motor coupled control technology under complex dynamic conditions is solved. This enables intelligent drive mode selection and torque distribution for the dual-motor drive system of electric vehicles, improving energy efficiency management and adaptive capabilities.
Patent Information
- Application Number
- CN202411808980.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Existing dual-motor coupling control technology lacks intelligence and adaptability, making it difficult to achieve flexible and coordinated control between motors under complex dynamic conditions, resulting in problems such as unreasonable energy distribution, insufficient power redundancy, or excessive energy consumption.
An intelligent control method combining Deep Q-Network (DQN) and Deep Deterministic Policy Gradient (DDPG) network is adopted to acquire vehicle status in real time, dynamically select drive mode and torque distribution ratio, and optimize motor energy consumption through trained Q-value function to achieve intelligent coupled drive of motor.
While meeting the driving capacity required by the vehicle's condition, the energy consumption of the motor is optimized to avoid power redundancy or excessive consumption. This enables dynamic adjustment of the output of the dual motors according to the actual situation of the vehicle, improving the system's adaptability and energy efficiency management.
Smart Images

Figure CN119611096B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new energy vehicle technology, and in particular to a method, system, device and medium for dual-motor coupled drive of electric vehicles. Background Technology
[0002] With the rapid development of the new energy vehicle industry, electric drive systems have gradually become one of the key technologies determining vehicle performance. The widespread adoption of electric vehicles is not only due to their zero emissions and environmental friendliness, but also to their significant advantages in energy efficiency, power output, and driving comfort. Against this backdrop, traditional single-motor drive systems have gradually revealed certain limitations, especially in complex driving environments, where they struggle to meet increasingly diverse demands in terms of power output, energy efficiency, and vehicle stability. This is mainly manifested in the following ways: single-motor systems have relatively simple energy efficiency management, failing to flexibly adjust the vehicle's output characteristics under different operating conditions, leading to energy waste, insufficient range, and unresponsive vehicle behavior.
[0003] In contrast, dual-motor drive systems, with their multiple advantages in power output, energy efficiency management, and dynamic performance control, are gradually becoming an important development direction for future electric vehicle drive systems. Dual-motor systems not only improve vehicle power redundancy and stability but also flexibly switch operating modes under different conditions, achieving superior energy efficiency management. This makes dual-motor drive technology not only limited to high-performance models but also likely to be widely adopted in various levels of new energy vehicles in the future.
[0004] In a dual-motor drive system, two motors work together in the vehicle's transmission system. By rationally distributing the output torque of the motors, they cooperate to achieve more efficient power output, energy recovery, and vehicle dynamic balance adjustment. This system can intelligently adjust the operating state of each motor according to vehicle conditions such as speed, load, and road conditions. Through the coordinated operation of the two motors, the vehicle can save energy by relying on a single motor during low-speed starts, while both motors work together at high speeds or under high loads to ensure the vehicle's power output. Simultaneously, when the vehicle brakes or goes downhill, the two motors can convert the vehicle's kinetic energy into electrical energy through energy recovery mode, feeding it back into the battery system, thereby effectively extending the vehicle's driving range.
[0005] However, the core challenge of dual-motor drive systems lies in achieving coordinated control between the motors, i.e., how to intelligently adjust the output of the two motors according to the vehicle's real-time operating conditions to optimize energy efficiency while ensuring power performance. Currently, most existing dual-motor coupling control technologies adopt feedback control strategies. This method typically relies on pre-set control rules or operating conditions, distributing torque to the motors through simple threshold or conditional judgments. While it can provide a certain level of response speed and system stability in specific scenarios, its core problem lies in the lack of sufficient intelligence and adaptive capabilities. This means that when the vehicle is operating under complex dynamic conditions, the system struggles to respond flexibly, such as in scenarios involving varying gradients, continuous speed changes, or variable loads. It cannot dynamically adjust the motor output according to the actual situation, easily leading to problems such as unreasonable energy distribution, insufficient power redundancy, or excessive energy consumption. Summary of the Invention
[0006] The purpose of this invention is to provide a method, system, device and medium for dual-motor coupled drive of electric vehicles, which can realize the torque distribution of dual motors under complex dynamic working conditions.
[0007] To address the aforementioned technical problems, embodiments of the present invention provide a dual-motor coupled drive method for electric vehicles, comprising the following steps:
[0008] Real-time acquisition of vehicle status for electric vehicles;
[0009] The vehicle state of the electric vehicle at each moment is input into the trained deep Q network (DQN) to dynamically obtain the driving mode of the electric vehicle at each moment. The driving mode is either single motor driving or dual motor coupled driving.
[0010] The DQN network is trained in the following way: the first Q value of different driving modes is estimated based on a preset first Q value function, and the training is carried out with the goal of maximizing the first Q value of the driving mode used in each vehicle state; the first Q value is used to indicate that the energy consumption of the electric vehicle is minimized while meeting the driving capability required by the vehicle state.
[0011] If the electric vehicle's driving mode is dual-motor coupled drive, then the trained Deep Deterministic Policy Gradient (DDPG) network will dynamically output the torque distribution ratio of the two motors of the electric vehicle at each moment based on the vehicle state of the electric vehicle at each moment.
[0012] The DDPG network is trained as follows: based on a preset second Q-value function, the second Q-value for different torque distribution ratios in each vehicle state is estimated, and the training is performed with the goal of maximizing the second Q-value of the torque distribution ratio in each vehicle state; the second Q-value is used to indicate that the energy consumption of the two motors is minimized corresponding to the adopted torque distribution ratio.
[0013] Based on the driving capacity required to meet the vehicle state at each moment and the torque distribution ratio of the two motors of the electric vehicle at each moment, the torque is dynamically distributed to the two motors of the electric vehicle, so that the two motors drive with the distributed torque coupling.
[0014] Optionally, the DDPG network includes an Actor network and a Critic network;
[0015] The Actor network is used to generate the torque distribution ratio between the two motors based on the input vehicle state.
[0016] The Critic network is used to estimate the second Q value based on the second Q value function to adopt the generated torque distribution ratio under the input vehicle state;
[0017] The DDPG network updates the Actor network parameters with the objective of maximizing the second Q value estimated by the Critic network.
[0018] Optionally, the network parameters of the Critic network are updated based on the temporal difference error method, using the difference between the second Q value of the current vehicle state and the second Q value of the next vehicle state.
[0019] Optionally, the DDPG network employs a replay buffer mechanism, which stores experience consisting of several corresponding current vehicle states, torque distribution ratios, the instantaneous reward used by the second Q-value function, and the next vehicle state; the network parameters of the Actor network and the Critic network are updated by randomly sampling the experience in the replay buffer.
[0020] Optionally, the vehicle status includes: the electric vehicle's speed, acceleration, state of charge, and torque of the two motors.
[0021] Optionally, the torque distribution ratio of the two motors output by the DDPG network is in the range of [0, 1], and the motors of the electric vehicle include a first motor and a second motor;
[0022] The method of dynamically allocating torque to the two motors of an electric vehicle by satisfying the driving capability required by the vehicle's state at each moment and the torque distribution ratio of the two motors at each moment includes:
[0023] The required total vehicle torque to meet the vehicle state is determined based on the driving capability required to meet the vehicle state at each moment.
[0024] When the torque distribution ratio between the two motors is 1, all the vehicle torque required to meet the vehicle's condition will be distributed to the first motor.
[0025] When the torque distribution ratio between the two motors is 0, all the vehicle torque required to meet the vehicle's status will be distributed to the second motor.
[0026] When the torque distribution ratio of the two motors is between [0, 1], the torque is allocated to the first motor and the second motor according to the total vehicle torque required to meet the vehicle state and the torque distribution ratio.
[0027] Optionally, the method further includes:
[0028] If the current driving mode of the electric vehicle is driven by any one motor, then all the vehicle torque required to meet the vehicle's state will be distributed to any one motor, allowing any one motor to drive independently.
[0029] Embodiments of the present invention also provide a dual-motor coupled drive system for electric vehicles, comprising:
[0030] The status acquisition module is used to acquire the vehicle status of electric vehicles in real time.
[0031] The mode selection module is used to input the vehicle state of the electric vehicle at each moment into the trained deep Q network (DQN) to dynamically obtain the driving mode of the electric vehicle at each moment. The driving mode is either individual driving of any motor or coupled driving of two motors.
[0032] The DQN network is trained in the following way: the first Q value of different driving modes is estimated based on a preset first Q value function, and the training is carried out with the goal of maximizing the first Q value of the driving mode used in each vehicle state; the first Q value is used to indicate that the energy consumption of the electric vehicle is minimized while meeting the driving capability required by the vehicle state.
[0033] The torque distribution module is used to dynamically output the torque distribution ratio of the two motors of the electric vehicle at each moment based on the vehicle state at each moment when the electric vehicle is driven in the dual-motor coupled drive mode.
[0034] The DDPG network is trained as follows: based on a preset second Q-value function, the second Q-value for different torque distribution ratios in each vehicle state is estimated, and the training is performed with the goal of maximizing the second Q-value of the torque distribution ratio in each vehicle state; the second Q-value is used to indicate that the energy consumption of the two motors is minimized corresponding to the adopted torque distribution ratio.
[0035] The motor drive module is used to dynamically distribute torque to the two motors of the electric vehicle based on the driving capability required to meet the vehicle state at each moment and the torque distribution ratio of the two motors at each moment, so that the two motors are driven by torque coupling.
[0036] Embodiments of the present invention also provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described electric vehicle dual-motor coupling drive method.
[0037] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described electric vehicle dual-motor coupling drive method.
[0038] The dual-motor coupled drive method for electric vehicles provided by this invention has at least the following beneficial effects:
[0039] First, the DQN network is used to select the driving mode of the electric vehicle based on the vehicle status, determining whether it is driven by either motor alone or by a dual-motor coupled drive. If it is a dual-motor coupled drive, the DDPG network is then used to output the torque distribution ratio of the two motors in the dual-motor coupled drive mode, thereby controlling the torque distribution between the two motors.
[0040] This method utilizes real-time feedback on vehicle status and employs a hybrid control system combining DQN and DDPG networks to dynamically select the driving mode and torque distribution ratio of the electric vehicle, thereby achieving intelligent driving mode selection and torque distribution for the dual-motor drive system. Furthermore, by coupling the dual motors of the electric vehicle using this method, the energy consumed by the two motors can be optimized while meeting the driving capacity required by the vehicle's current state. This allows for dynamic adjustment of the dual motor output based on actual vehicle conditions (including dynamic or multi-condition driving), rationally distributing motor energy and avoiding insufficient power redundancy or excessive energy consumption. Attached Figure Description
[0041] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.
[0042] Figure 1 This is a flowchart of a dual-motor coupled drive method for an electric vehicle according to an embodiment of the present invention;
[0043] Figure 2 This is a network architecture diagram provided according to an embodiment of the present invention;
[0044] Figure 3 This is a network training flowchart provided according to an embodiment of the present invention. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the embodiments of the present invention to facilitate a better understanding of the invention. However, the technical solutions claimed in the present invention can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with and referenced by each other without contradiction.
[0046] Currently, the motor coordination control methods used in dual-motor drive systems generally employ rule-based control. This involves directly controlling the torque distribution and other operations of the two motors based on a pre-set set of control rules and thresholds, according to information such as the vehicle's operating status and environmental conditions. These rules are typically based on engineering experience, experimental data, or simulation analysis, possessing a clear logical structure and enabling rapid vehicle control decisions. However, this type of method has the following drawbacks:
[0047] 1. Poor flexibility: Due to the relatively fixed preset rules, it is difficult to adapt to complex or changing operating conditions. The system cannot adaptively adjust the control strategy according to the actual driving scenario, making it difficult to achieve optimal energy efficiency management.
[0048] 2. Limited optimization capability: The control rules rely on experience to set and lack global optimization capability, which may lead to suboptimal control strategies, especially under nonlinear and multi-condition conditions.
[0049] 3. Poor scalability: When the driving environment or operating conditions change, the system needs to readjust or expand the rule set, making it difficult to automatically learn new driving modes or operating conditions.
[0050] While rule-based control methods are simple and easy to implement, the fixed nature of the rules limits their adaptability and global optimization capabilities, making it difficult to provide optimal control strategies under complex and dynamic conditions. They are typically suitable for relatively simple or static conditions; for more complex and dynamically changing situations, more advanced control strategies may be required.
[0051] Furthermore, existing dual-motor control technologies are typically based on linear or simple logic control methods, and the limitations of their control strategies are particularly evident under complex nonlinear operating conditions. For example, the energy efficiency curves of a vehicle under different loads or speeds do not change linearly, but rather exhibit high complexity and time-varying characteristics. Traditional control algorithms often fail to fully capture these changing patterns, resulting in poor energy efficiency optimization performance under multiple operating conditions.
[0052] Therefore, current dual-motor coupling control technology urgently needs a more intelligent solution. This solution not only needs to find a balance between power output and energy recovery, but also needs to have the ability to adapt to complex operating conditions. An ideal dual-motor control system should be able to monitor and analyze the vehicle's dynamic operating conditions in real time, and automatically adjust the output characteristics of the two motors based on this information, thereby achieving more efficient energy management and better driving performance.
[0053] In recent years, with the development of artificial intelligence technology, intelligent control technology based on machine learning, especially reinforcement learning, has shown great potential in the control of complex nonlinear systems. Through reinforcement learning, the system can autonomously learn the optimal control strategy in an unsupervised environment, no longer relying on preset control rules. Reinforcement learning algorithms can continuously optimize the coupling control strategy of the motor based on feedback data under different operating conditions, thereby achieving dynamic optimization of energy efficiency management while ensuring power performance and stability. This provides a completely new technical path for dual-motor coupling control, significantly improving the system's intelligence level and adaptability. This intelligent dual-motor coupling control system not only improves the overall energy efficiency of the vehicle but also optimizes the vehicle's power response under complex driving conditions by dynamically adjusting output torque and speed, thus providing the driver with a smoother and more comfortable driving experience. In the future, intelligent dual-motor coupling control technology is expected to become one of the core technologies in the drive system of new energy vehicles, driving new leaps in electric vehicles in terms of performance, range, and user experience.
[0054] One embodiment of the present invention relates to a dual-motor coupled drive method for electric vehicles. The implementation details of the dual-motor coupled drive method for electric vehicles in this embodiment are described in detail below. The following content is only for the convenience of understanding and is not necessary for implementing this solution.
[0055] The specific process of the electric vehicle dual-motor coupling drive method in this embodiment is as follows: Figure 1 As shown, it includes:
[0056] Step 101: Obtain the real-time vehicle status of the electric vehicle.
[0057] The vehicle status includes: the electric vehicle's speed, acceleration, state of charge (SOC), and the torque of the two motors.
[0058] Step 102: Input the vehicle state of the electric vehicle at each moment into the trained Deep Q-Network (DQN) to dynamically obtain the driving mode of the electric vehicle at each moment. The driving mode is either single-motor driving or dual-motor coupled driving. The DQN network is trained in the following way: based on a preset first Q-value function, estimate the first Q value of different driving modes for each vehicle state, and train with the goal of maximizing the first Q value of the driving mode used in each vehicle state. The first Q value is used to indicate that the energy consumption of the electric vehicle is minimized while meeting the driving capability required by the vehicle state.
[0059] Specifically, the core of Deep Q-Network (DQN) is to use a neural network to approximate the Q-value table in traditional Q-learning. The Q-value function Q(s,a) represents the expected reward obtained by taking action a in state s. By training the neural network, the DQN network can estimate the Q-value of each state-action pair in a high-dimensional state space, thereby guiding the agent to choose the optimal action. Its basic idea is to use a target network and an estimation network (Q-Network), and to break down the correlation between data through experience replay, thereby improving the stability and efficiency of training.
[0060] The estimation network is used to calculate the Q-value of the current state. In each state s, the network outputs the Q-value of each possible action in that state. By comparing these Q-values, the action corresponding to the largest Q-value is selected as the agent's best action in that state.
[0061] The target network has the same structure as the estimation network, except that its parameters are updated more slowly. The target network is used to calculate the target Q-value, which guides the learning of the estimation network. The purpose of the target network is to improve training stability and prevent the training process from diverging due to excessively rapid network updates.
[0062] Estimated network structure and target network structure:
[0063] Input Layer: The input layer receives the feature vector of the current state s. For the control task of a dual-motor system in an electric vehicle, the input layer receives state variables such as vehicle speed, acceleration, SOC, and motor torque.
[0064] Hidden layers: Hidden layers typically consist of several fully connected layers and employ the ReLU activation function. Hidden layers are responsible for extracting deep features from the input state, helping the network learn the complex relationships between different states.
[0065] Output layer: The number of neurons in the output layer is equal to the size of the action space. In dual-motor drive mode control, each neuron in the output layer corresponds to a possible drive mode, and the output value is the Q-value of that action.
[0066] The Q-value function Q(s,a) is a value function for a state-action pair. It represents the expected reward (total reward) that the agent can obtain after choosing action a in state s. The formula is as follows:
[0067]
[0068] In the formula, r t γ is the reward obtained by the agent after performing action a, starting from state s; γ is the discount factor, which measures the degree of decay of future rewards (usually 0 < γ < 1), making the agent pay more attention to the recent reward; E represents the expected value of all possible paths.
[0069] The agent's goal is to select action a that maximizes Q(s,a). * Thus, in the long run, the highest cumulative return can be obtained:
[0070]
[0071] Therefore, the core task of the DQN network is to approximate Q(s,a) through the neural network and find the optimal action a that maximizes the Q value in each state. * .
[0072] DQN networks update their Q-values using the Temporal Difference (TD) method. The basic idea of TD error is to use the difference between the current Q-value and the target Q-value as a learning signal to update the neural network's parameters, gradually approximating the true Q-value. The target Q-value (TD target) is expressed by the following formula:
[0073]
[0074] In the formula, r is the immediate reward obtained after performing action a in the current state s; s′ is the next state after performing the action; θ - These are the parameters of the target network, used to calculate the target Q-value; It is the maximum Q value that can be obtained among all possible actions in state s′, representing the maximum expected reward that can be obtained by taking the optimal action in the next state.
[0075] By calculating the TD error, the agent can measure the difference between the current network's Q-value and the target Q-value. The formula for the TD error is:
[0076] d = (yQ(s, a; θ)) 2
[0077] In the formula, y is the target value Q; Q(s,a;θ) is the estimated Q value of choosing action a under the current network state s.
[0078] Therefore, the DQN network approximates the Q-value function of each state-action pair through a neural network. The neural network learns the relationship between the expected reward of state s and the corresponding action a, and continuously updates the network weights so that the Q-value function can more accurately reflect long-term rewards. Maximizing the Q-value and selecting the optimal action: In each state s, the DQN network calculates the Q-value Q(s,a) of all possible actions and selects the action a that maximizes the Q-value. * This process ensures that the agent can choose the optimal action in each state, thereby obtaining the highest long-term reward. Simultaneously, the DQN network introduces a target network, which, by slowly updating its parameters, avoids instability caused by Q-value fluctuations during network training. The target network ensures that, during the learning process, the estimation network can gradually approximate the true Q-value, guaranteeing the rationality and reliability of the output actions.
[0079] Step 103: If the electric vehicle's driving mode is dual-motor coupled drive, then the trained Deep Deterministic Policy Gradient (DDPG) network dynamically outputs the torque distribution ratio of the two motors of the electric vehicle at each moment, based on the vehicle state of the electric vehicle at each moment. The DDPG network is trained in the following way: it estimates the second Q value for different torque distribution ratios in each vehicle state based on a preset second Q value function, and trains with the goal of maximizing the second Q value of the torque distribution ratio adopted in each vehicle state. The second Q value is used to indicate that the energy consumption of the two motors is minimized corresponding to the adopted torque distribution ratio.
[0080] Specifically, the Deterministic Policy Gradient (DDPG) network combines the ideas of policy gradient and Q-learning, and makes decisions through two neural networks:
[0081] Actor networks are used to generate continuous actions a = π(s|θ) π ), representing choosing an action a, θ in state s. π These are the parameters of the policy network. This network directly outputs a continuous action, suitable for tasks requiring precise control, such as torque proportional distribution.
[0082] Critic networks are used to evaluate the quality of action selection by actor networks, i.e., to estimate the state-action value function Q(s,a|θ). Q ), θ Q These are the parameters of the Critic network. The Critic network provides the Q-value for taking action a in the current state s, representing the expected reward that action can obtain.
[0083] To maintain the stability of the learning process, DDPG also uses two target networks:
[0084] Target Actor Network: Has the same parameters as the Actor Network, but updates more slowly, ensuring a smooth policy transition. Target Critic Network: Has the same parameters as the Critic Network, also updates more slowly, ensuring the stability of Q-value estimation.
[0085] The goal of the DDPG network is to generate the optimal action a = π(s|θ) through the Actor network. π The Q-value of the action is evaluated using a Critic network. This process can be represented by the following formula:
[0086] Actor network generates actions:
[0087] a=π(s|θ π )
[0088] In each state s, the Actor network generates a continuous action 'a' based on the input state. This action can be represented as selecting a certain proportion or distribution ratio of motor torque in the current state.
[0089] Critic network evaluation Q-value:
[0090] Q(s,a|θ Q )
[0091] The Critic network receives state s and action a generated by the Actor network, and calculates the Q-value of taking that action in that state. The Q-value represents the cumulative reward that can be obtained in the future, so the higher the Q-value, the better the action.
[0092] Policy Update: The goal is to maximize the Q-value evaluated by the Critic network; therefore, the update objective of the Actor network is:
[0093]
[0094] By adjusting the parameters θ of the Actor network π This allows the actions generated by the Actor network to maximize the Q-value given by the Critic network.
[0095] The Critic network is updated similarly to Q-learning, adjusting the network's parameters θ by minimizing the Temporal Difference (TD) error. Q The target Q-value is calculated using the target network and is expressed as:
[0096] y=r+gQ′(s′,π′(s′|θ π′ )|θ Q′ )
[0097] In the formula, r is the immediate reward, representing the feedback obtained by the agent after performing the action; γ is the discount factor, representing the importance of future rewards; Q′(s′,π′(s′|θ) π′ )|θ Q′ ) is the Q-value corresponding to the next state s′ calculated by the target network; θ π′ and θ Q′ These are the parameters of the target Actor and the target Critic network.
[0098] Then, the parameters of the Critic network are updated by minimizing the TD error:
[0099] L(θ Q )=E(yQ(s,a|θ Q )) 2
[0100] By learning the Q-values of the current state and actions, the Critic network is able to evaluate the quality of different actions, thereby helping the Actor network generate the optimal actions.
[0101] To improve the stability and efficiency of learning, DDPG employs the following two mechanisms:
[0102] Experience replay: A replay buffer stores the agent's experience (s, a, r, s′) in the environment. During each update, a small batch of experience is randomly sampled to update the Actor and Critic networks, breaking the temporal correlation between samples and improving training stability.
[0103] Target Networks: DDPG uses target Actor and target Critic networks. These networks have the same structure as the main network, but their parameters are updated more slowly (using soft updates). By controlling the update rate of the target networks, drastic fluctuations in Q-value estimation can be avoided, improving training stability. The update formula for the target networks is:
[0104] tθ+(1-t)θ′→θ′
[0105] In the formula, t is a very small value (such as 0.001), which represents the update rate of the target network parameters.
[0106] Therefore, the Actor network of the DDPG network generates a continuous action 'a' based on the electric vehicle's state (such as speed, SOC, motor torque, etc.), representing the torque distribution ratio between the two motors. This ratio is directly used to control the motor output. For example, the torque distribution ratio 'a' ∈ 0, 1 output by the Actor network can represent the power distribution relationship between motor 1 and motor 2. The Critic network is responsible for evaluating the rationality of the torque distribution strategy output by the Actor network, determining whether the strategy can achieve optimal energy management and power output under the current state. By evaluating the Q-value, the Critic network helps the Actor network generate a better torque distribution strategy. The advantage of the DDPG network lies in its ability to handle continuous action spaces, making the motor torque distribution more refined and flexible. Combined with the Q-value evaluation by the Critic network, the direction of strategy adjustment becomes clearer, ultimately achieving optimal energy management and torque distribution.
[0107] Step 103: Based on the driving capability required to meet the vehicle state of the electric vehicle at each moment and the torque distribution ratio of the two motors of the electric vehicle at each moment, dynamically distribute torque to the two motors of the electric vehicle so that the two motors drive with the distributed torque coupling.
[0108] In the specific implementation, the torque distribution ratio of the two motors output by the DDPG network ranges from [0, 1], and the electric vehicle's motors include a first motor and a second motor. Therefore, in the process of dynamically distributing torque to the two motors of the electric vehicle by satisfying the total vehicle torque required for the vehicle state at each moment and the torque distribution ratio of the two motors at each moment, when the torque distribution ratio of the two motors is 1, all the total vehicle torque required for the vehicle state is distributed to the first motor; when the torque distribution ratio of the two motors is 0, all the total vehicle torque required for the vehicle state is distributed to the second motor; when the torque distribution ratio of the two motors is between [0, 1], torque is distributed to the first motor and the second motor according to the total vehicle torque required for the vehicle state and the torque distribution ratio.
[0109] In one example, if the electric vehicle's drive mode is either motor-driven, then all the vehicle torque required to meet the vehicle's state will be distributed to either motor, allowing either motor to drive independently.
[0110] In this embodiment, the DQN network is first used to select the driving mode of the electric vehicle based on its vehicle status, determining whether it is driven by either a single motor or by a dual-motor coupled drive. If it is a dual-motor coupled drive, the DDPG network is then used to output the torque distribution ratio of the two motors in the dual-motor coupled drive mode, controlling the torque distribution between the two motors. This method, through real-time feedback of the vehicle status and using a hybrid control of the DQN and DDPG networks, dynamically selects the driving mode and torque distribution ratio of the electric vehicle, thereby realizing intelligent driving mode selection and torque distribution of the dual-motor drive system of the electric vehicle. Furthermore, by using this method for dual-motor coupled drive of the electric vehicle, the energy consumed by the two motors can be optimized while meeting the driving capacity required by the vehicle status. This allows for dynamic adjustment of the dual motor output according to the actual vehicle conditions (including dynamic or multi-condition driving), rationally distributing motor energy and avoiding insufficient vehicle power redundancy or excessive energy consumption.
[0111] In one embodiment, the network architecture of the DQN network and DDPG network in this invention is as follows: Figure 2 As shown, the network training process is as follows: Figure 3 As shown, the training process includes:
[0112] 1. System initialization:
[0113] (1) After the vehicle starts, the system initializes relevant parameters, including vehicle status (speed, acceleration, SOC, torque of the two motors, etc.).
[0114] (2) These states are input into the DQN network and DDPG network for subsequent mode selection and torque distribution.
[0115] 2. Status Input:
[0116] (1) Input parameters include: speed (v), acceleration (a), SOC, T_motor1 (motor 1 torque), T_motor2 (motor 2 torque).
[0117] (2) The state space is input into the DQN and DDPG networks as the basis for subsequent action selection. The state is defined as follows:
[0118] state = {v t ,T_motor it ,N_motor it |i={1,2},t={1,2,L n|n=steps max}}
[0119] 3. DQN network mode selection:
[0120] (1) The three working modes are discrete: 0 indicates that motor 1 is working, 1 indicates that motor 2 is working, and 2 indicates that both motors are working together. The DQN motion space is defined as follows:
[0121] action DQN ={0,1,2}
[0122] (2) The DQN network selects the vehicle's driving mode based on the input state.
[0123] mode=action DQN
[0124] (3) The DQN network selects an optimal driving mode based on the current state at each time step.
[0125] 4. Torque distribution in DDPG network:
[0126] (1) If the DQN is selected as the dual motor mode (mode 2), the DDPG network is started to distribute torque proportionally.
[0127] action DDPG ∈0,1
[0128] (2) The DDPG network outputs a continuous distribution value to control the torque distribution ratio between motor 1 and motor 2.
[0129] distribution = action DDPG
[0130] (3) Calculate the torque borne by motor 1 as T_motor based on the torque distribution ratio. 1t The torque borne by motor 2 is T_motor 2t .
[0131] T_motor 1t =T_req*distribution
[0132] T_motor 2t = T_req*(1-distribution)
[0133] 5. Environmental Updates and Reward Calculation:
[0134] (1) The system applies the decision results of DQN and DDPG to the model vehicle and performs vehicle state updates, including actual speed, SOC, torque, etc.
[0135] (2) Calculate the reward value for the current time step based on the updated state. tThis reward value is used to guide network learning in order to optimize the energy management and drive efficiency of the motor.
[0136] reward t =-(a*|Dsoc|+b*|Dv|)
[0137]
[0138] 6. Experience storage:
[0139] (1) Store the current time step's state, action, reward, next state, and whether it has terminated into the experience pool (Replay Buffer).
[0140] (2) The data in the experience pool will be used for subsequent network updates to achieve policy optimization.
[0141] 7. DQN and DDPG network updates:
[0142] (1) By sampling a batch of data from the experience pool, the DQN network updates the Q-value network and optimizes the working mode selection strategy.
[0143] (2) At the same time, the DDPG network uses data from the experience pool to update the Actor and Critic networks to optimize the torque distribution strategy.
[0144] (3) The updates of the two networks are performed in parallel, which ensures that the system can gradually improve the performance of the strategy in each round of training.
[0145] 8. Termination condition check:
[0146] (1) After each execution of the control strategy, the system will check whether the current step count has reached the maximum step size requirement (equivalent to the end of a working cycle).
[0147] (2) If the maximum step size requirement is reached, a working cycle ends, the system stores the accumulated experience and updates the network model.
[0148] (3) After the training rounds reach the maximum set number of rounds, the system will stop training and save the final model.
[0149] Based on the model obtained through this training process, drive mode selection and torque ratio allocation can be performed:
[0150] First, the system collects the current vehicle status information, including vehicle speed (v), remaining battery charge (soc), motor torque (T_motor), and rotational speed (N_motor). This information serves as the input to the DQN and DDPG networks. The DQN network outputs a discrete action (mode) based on the current state, representing the vehicle's current operating mode. There are three operating modes: Mode 0 (motor 1 driven alone), Mode 1 (motor 2 driven alone), and Mode 2 (dual-motor coupled drive). In Modes 0 and 1, the system only controls the torque output of one motor (distributing all the vehicle's required torque to a single motor), and the DDPG network does not need to be activated.
[0151] The DDPG network is activated when the DQN network selects mode 2, and is responsible for calculating the torque distribution ratio between the two motors. The DDPG network outputs a continuous value, `distribution`, ranging from [0,1], representing the torque distribution ratio between the two motors. Specifically, when `distribution = 1`, all torque is distributed to motor 1; when `distribution = 0`, all torque is distributed to motor 2; when `distribution` is an intermediate value, the two motors distribute torque proportionally (motor 1 receives the product of the total required torque and the proportionality coefficient, and motor 2 receives the product of the total required torque and 1 minus the proportionality coefficient). Then, the two motors output torque according to their assigned torque and transmit it to the vehicle dynamics system to drive the vehicle forward. The specific distribution method is as follows:
[0152] Motor 1's torque capacity: T_motor 1t =T_req*distribution
[0153] Motor 2's torque capacity: T_motor 2t = T_req*(1-distribution)
[0154] Through this intelligent torque distribution strategy, the system can achieve optimal energy allocation under different operating conditions, ensuring vehicle power output while maximizing energy efficiency.
[0155] Furthermore, at each time step, the system updates the vehicle's state based on the mode selection result and torque distribution, and calculates the reward value for the current step. Through an experience playback mechanism, the parameters of the DQN and DDPG networks are continuously updated to adapt to complex operating conditions and gradually learn the optimal control strategy.
[0156] The dual-motor coupled drive method for electric vehicles of the present invention has the following beneficial effects:
[0157] 1. Improve adaptability to complex working conditions
[0158] (1) Introduce a more intelligent control method, combining machine learning or deep reinforcement learning technology, and automatically adjust the torque distribution and working mode of the dual motors by learning the dynamic behavior of the vehicle and changes in the external environment in real time, so as to solve the problem of poor adaptability of rule-based methods to different working conditions.
[0159] (2) Improve the system’s responsiveness and adaptive adjustment capabilities in complex and ever-changing driving environments through a real-time data-driven learning mechanism.
[0160] 2. Achieve globally optimal control strategy
[0161] (1) To address the problem that rule-based control methods cannot achieve global optimization, my solution combines adaptive optimization methods such as reinforcement learning. By dynamically analyzing driving conditions and energy status, it optimizes multiple objectives (such as power performance, energy efficiency, and comfort) to achieve global optimization among multiple objectives.
[0162] (2) In environments with high system complexity and multiple variables, it can weigh various objectives and seek the overall optimal control strategy, rather than being limited to preset rules.
[0163] 3. Reduce computational complexity and improve real-time response capabilities
[0164] (1) To address the issues of high computational complexity and poor real-time performance of optimization algorithms, my solution combines lightweight deep reinforcement learning strategies, such as neural network-based control strategies, with traditional optimization algorithms to avoid cumbersome iterative calculations and improve the response speed and computational efficiency of the control system.
[0165] (2) By introducing technologies such as approximate dynamic programming and online learning algorithms, the high dependence on computing resources in real-time control is reduced, and an effective balance between real-time performance and control accuracy is achieved.
[0166] 4. Improve robustness to changes in the external environment
[0167] (1) The problem of optimization algorithms being highly dependent on the system model and lacking the ability to cope with environmental changes will be solved by introducing data-driven deep learning technology. My solution will use real-time data to build a dynamic model and combine it with external sensor information to dynamically update the model parameters, making the system more robust to changes in the external environment.
[0168] (2) The system can automatically adjust under complex conditions such as road slope, wind resistance changes, and vehicle speed fluctuations, avoiding control failures caused by inaccurate models.
[0169] 5. Achieve intelligent torque distribution and energy management
[0170] To address the limitations of existing torque distribution and energy management strategies, my technical solution employs deep reinforcement learning to learn the optimal dual-motor torque distribution strategy. Through intelligent control, it achieves efficient motor collaboration across different driving modes, maximizing energy recovery and battery utilization efficiency, thereby improving overall energy efficiency.
[0171] 6. Provides scalability and system upgrade capabilities.
[0172] Not only is it applicable to current dual-motor coupling systems, but it will also have strong scalability, allowing it to be applied to more types of drive systems in the future (such as three-motor or hybrid systems). Through modular design, it can be flexibly expanded to different types of electric vehicle architectures, and reserves space for subsequent system upgrades.
[0173] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this invention. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the protection scope of this invention.
[0174] Another embodiment of the present invention relates to a dual-motor coupled drive system for an electric vehicle. The implementation details of this dual-motor coupled drive system are described below. The following details are provided for ease of understanding and are not essential for implementing this solution. The dual-motor coupled drive system for an electric vehicle in this embodiment includes:
[0175] The status acquisition module is used to acquire the vehicle status of electric vehicles in real time.
[0176] The mode selection module is used to input the vehicle state of the electric vehicle at each moment into the trained deep Q network (DQN) to dynamically obtain the driving mode of the electric vehicle at each moment. The driving mode is either single motor driving or dual motor coupled driving.
[0177] The DQN network is trained in the following way: the first Q value of different driving modes is estimated based on a preset first Q value function, and the training is carried out with the goal of maximizing the first Q value of the driving mode used in each vehicle state; the first Q value is used to indicate that the energy consumption of the electric vehicle is minimized while meeting the driving capability required by the vehicle state.
[0178] The torque distribution module is used to dynamically output the torque distribution ratio of the two motors of the electric vehicle at each moment based on the vehicle state at each moment when the electric vehicle is driven in the dual-motor coupled drive mode.
[0179] The DDPG network is trained as follows: based on a preset second Q-value function, the second Q-value for different torque distribution ratios in each vehicle state is estimated, and the training is performed with the goal of maximizing the second Q-value of the torque distribution ratio in each vehicle state; the second Q-value is used to indicate that the energy consumption of the two motors is minimized corresponding to the adopted torque distribution ratio.
[0180] The motor drive module is used to dynamically distribute torque to the two motors of the electric vehicle based on the driving capability required to meet the vehicle state at each moment and the torque distribution ratio of the two motors at each moment, so that the two motors are driven by torque coupling.
[0181] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0182] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0183] Another embodiment of the present invention relates to a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the electric vehicle dual-motor coupling drive method of the above embodiments.
[0184] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0185] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0186] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.
[0187] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0188] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A dual-motor coupled drive method for an electric vehicle, characterized in that, include: Real-time acquisition of vehicle status for electric vehicles; The vehicle state of the electric vehicle at each moment is input into the trained deep Q network (DQN) to dynamically obtain the driving mode of the electric vehicle at each moment. The driving mode is either single motor driving or dual motor coupled driving. The DQN network is trained in the following way: the first Q value of different driving modes is estimated based on a preset first Q value function, and the training is carried out with the goal of maximizing the first Q value of the driving mode used in each vehicle state; the first Q value is used to indicate that the energy consumption of the electric vehicle is minimized while meeting the driving capability required by the vehicle state. If the electric vehicle's driving mode is dual-motor coupled drive, then the trained Deep Deterministic Policy Gradient (DDPG) network will dynamically output the torque distribution ratio of the two motors of the electric vehicle at each moment based on the vehicle state of the electric vehicle at each moment. The DDPG network is trained as follows: based on a preset second Q-value function, the second Q-value for different torque distribution ratios in each vehicle state is estimated, and the training is performed with the goal of maximizing the second Q-value of the torque distribution ratio in each vehicle state; the second Q-value is used to indicate that the energy consumption of the two motors is minimized corresponding to the adopted torque distribution ratio. Based on the driving capacity required to meet the vehicle state at each moment and the torque distribution ratio of the two motors of the electric vehicle at each moment, the torque is dynamically distributed to the two motors of the electric vehicle, so that the two motors drive with the distributed torque coupling.
2. The electric vehicle dual-motor coupled drive method according to claim 1, characterized in that, The DDPG network includes an Actor network and a Critic network; The Actor network is used to generate the torque distribution ratio between the two motors based on the input vehicle state. The Critic network is used to estimate the second Q value based on the second Q value function to adopt the generated torque distribution ratio under the input vehicle state; The DDPG network updates the Actor network parameters with the objective of maximizing the second Q value estimated by the Critic network.
3. The electric vehicle dual-motor coupled drive method according to claim 2, characterized in that, The network parameters of the Critic network are updated based on the temporal difference error method, using the difference between the second Q value of the current vehicle state and the second Q value of the next vehicle state.
4. The electric vehicle dual-motor coupled drive method according to claim 3, characterized in that, The DDPG network employs a replay buffer mechanism, which stores experience consisting of several corresponding current vehicle states, torque distribution ratios, the instantaneous reward used by the second Q-value function, and the next vehicle state. The network parameters of the Actor network and Critic network are updated by randomly sampling the experience in the replay buffer.
5. The electric vehicle dual-motor coupled drive method according to any one of claims 1 to 4, characterized in that, The vehicle status includes: the electric vehicle's speed, acceleration, state of charge, and the torque of the two motors.
6. The electric vehicle dual-motor coupled drive method according to claim 1, characterized in that, The torque distribution ratio of the two motors output by the DDPG network ranges from [0, 1], and the electric vehicle's motors include a first motor and a second motor. The method of dynamically allocating torque to the two motors of the electric vehicle based on the driving capability required to meet the vehicle state at each moment and the torque distribution ratio of the two motors at each moment includes: By satisfying the driving capability required for the electric vehicle's state at every moment, the required total vehicle torque to satisfy the vehicle state is determined. When the torque distribution ratio between the two motors is 1, all the vehicle torque required to meet the vehicle's condition will be distributed to the first motor. When the torque distribution ratio between the two motors is 0, all the vehicle torque required to meet the vehicle's status will be distributed to the second motor. When the torque distribution ratio of the two motors is between [0, 1], the torque is allocated to the first motor and the second motor according to the total vehicle torque required to meet the vehicle state and the torque distribution ratio.
7. The electric vehicle dual-motor coupled drive method according to claim 1, characterized in that, The method further includes: If the electric vehicle is driven by any one motor, then all the torque required to meet the vehicle's needs will be distributed to that motor, allowing it to drive independently.
8. A dual-motor coupled drive system for an electric vehicle, characterized in that, include: The status acquisition module is used to acquire the vehicle status of electric vehicles in real time. The mode selection module is used to input the vehicle state of the electric vehicle at each moment into the trained deep Q network (DQN) to dynamically obtain the driving mode of the electric vehicle at each moment. The driving mode is either individual driving of any motor or coupled driving of two motors. The DQN network is trained in the following way: the first Q value of different driving modes is estimated based on a preset first Q value function, and the training is carried out with the goal of maximizing the first Q value of the driving mode used in each vehicle state; the first Q value is used to indicate that the energy consumption of the electric vehicle is minimized while meeting the driving capability required by the vehicle state. The torque distribution module is used to dynamically output the torque distribution ratio of the two motors of the electric vehicle at each moment based on the vehicle state at each moment when the electric vehicle is driven in the dual-motor coupled drive mode. The DDPG network is trained as follows: based on a preset second Q-value function, the second Q-value for different torque distribution ratios in each vehicle state is estimated, and the training is performed with the goal of maximizing the second Q-value of the torque distribution ratio in each vehicle state; the second Q-value is used to indicate that the energy consumption of the two motors is minimized corresponding to the adopted torque distribution ratio. The motor drive module is used to dynamically distribute torque to the two motors of the electric vehicle based on the driving capability required to meet the vehicle state at each moment and the torque distribution ratio of the two motors at each moment, so that the two motors are driven by torque coupling.
9. A computer device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the electric vehicle dual-motor coupling drive method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the electric vehicle dual-motor coupling drive method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Motor torque control method, device and equipment for dual-motor electric vehicle and vehicle
CN113147429A
Optimal torque distribution method and system for dual-motor four-wheel drive automobile
CN119058430A