Vehicle motor anti-shake control method and system, vehicle and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DEEPAL AUTOMOBILE NANJING RESEARCH INSTITUTE CO LTD
- Filing Date
- 2025-07-30
- Publication Date
- 2026-08-07
AI Technical Summary
[0008](1)对于防抖模型,现有方法通常需要针对不同车型、负载、工况逐次调整参数,导致标定周期长、成本高
[0046](1)实现电机防抖功能
Smart Images

Figure CN120811202B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of motor anti-shake technology, specifically to a vehicle motor anti-shake control method, system, vehicle, and electronic equipment. Background Technology
[0002] With increasing global focus on environmental protection and sustainable energy development, the electric vehicle industry has ushered in unprecedented development opportunities, with a continuously expanding market size and rapid technological advancements. Among the many performance indicators of electric vehicles, the user driving experience has become a key area of focus for major automakers and research institutions. A good driving experience can not only improve user satisfaction and loyalty but also give products a differentiated advantage in fierce market competition, thereby promoting the healthy development of the entire industry.
[0003] Driving experience is a comprehensive concept, encompassing multiple aspects such as vehicle acceleration performance, handling stability, and ride comfort. Among these, ride comfort, as a crucial component of the driving experience, directly impacts the user's physical and mental well-being during long drives or daily commutes. A comfortable and quiet driving environment can effectively reduce driver fatigue and improve travel quality. Therefore, improving the ride comfort of electric vehicles has become a critical issue that the industry urgently needs to address.
[0004] In the core component of electric vehicles—the motor system—the drive shaft and wheels are typically connected by a rigid system. While this design ensures efficient power transmission, it also introduces some potential problems. Within certain torque ranges, the motor may experience speed fluctuations during operation. Due to the rigid connection between the drive shaft and wheels, these speed fluctuations are transmitted directly to the vehicle body without attenuation, causing vibrations and noise.
[0005] Vibrations in the vehicle body cause noticeable bumps and swaying to passengers, especially at low speeds or during frequent start-stop operations, where this discomfort is more pronounced. Simultaneously, the noise generated by engine speed fluctuations disrupts the quiet atmosphere inside the vehicle, interfering with passenger conversation and rest, and severely impacting driving comfort. Prolonged exposure to such a poor driving environment may also have negative health effects on passengers, such as causing dizziness and nausea. Therefore, effectively suppressing engine speed fluctuations in the motor system is crucial for improving the driving comfort of electric vehicles.
[0006] Currently, the most common approach to controlling motor system speed fluctuations is through equipment, specifically by adding shock absorber components to the motor system. These components absorb and attenuate vibration energy, reducing the transmission of speed fluctuations to the vehicle body and thus improving driving comfort to some extent. However, adding shock absorbers requires a redesign and reconfiguration of the motor system structure. This not only increases the number and complexity of components but also places higher demands on system installation space and assembly processes. Furthermore, the high cost of shock absorbers themselves, including material costs, manufacturing costs, and R&D costs, directly leads to a significant increase in the design cost of electric vehicles. In today's increasingly competitive market, high costs undoubtedly weaken a product's market competitiveness and limit its large-scale application.
[0007] Besides equipment-based methods, many existing speed jitter control schemes also heavily rely on high-precision speed sensors and road surface sensors. These sensors can monitor motor speed, torque, and road conditions in real time, feeding this data back to the control system's anti-jitter model so that the control strategy can be adjusted promptly to suppress speed jitter. However, this also raises a series of problems:
[0008] (1) For anti-shake models, existing methods usually require adjusting parameters one by one for different vehicle models, loads and operating conditions, resulting in long calibration cycles and high costs.
[0009] (2) Existing methods may introduce additional torque fluctuations due to excessive suppression of jitter, resulting in increased copper loss and energy consumption of the motor.
[0010] (3) The anti-shake model trained by pure simulation is easily affected by motor parameter drift, mechanical wear, etc., and requires periodic manual calibration.
[0011] Therefore, it is necessary to develop a new method, system, vehicle, and electronic equipment for vehicle motor anti-shake control. Summary of the Invention
[0012] The purpose of this invention is to provide a vehicle motor anti-shake control method, system, vehicle, and electronic equipment that can achieve motor anti-shake function, reduce calibration difficulty and cost, reduce losses, and allow online fine-tuning to adapt to the actual vehicle, ensuring control effect.
[0013] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0014] In a first aspect, the vehicle motor anti-vibration control method of the present invention includes the following steps:
[0015] The motor speed and required torque are acquired in real time, and the time-series and frequency-domain indices of the motor speed are calculated.
[0016] The collected motor speed, required torque, timing parameters, and frequency domain parameters are combined into a state vector;
[0017] The state vector is input into a pre-trained neural network model, and the compensation torque is output.
[0018] The compensated torque is superimposed with the required torque to generate a total torque, which is then transmitted to the motor controller for torque control.
[0019] The training process of the neural network model includes:
[0020] Set up an offline simulation environment: Simulate the dynamic behavior of the motor through a motor simulation model to generate simulation training data;
[0021] Offline training of reinforcement learning model: Based on the simulation training data, the neural network model in the reinforcement learning framework is trained in an offline simulation environment, with the motor simulation model as environmental feedback, until the neural network model meets the first preset requirement;
[0022] Online fine-tuning of the model: The offline-trained neural network model is deployed to a real vehicle environment, and a real motor controller is used as environmental feedback to fine-tune the neural network model until the second preset requirement is met.
[0023] One possible implementation is that the simulation training data includes demand torque, motor speed, and time-series and frequency-domain indices of motor speed.
[0024] One possible implementation is that the required torque is the required torque at the current time t; and the motor speed is the historical motor speed data from time tn to time t-1. Historical motor speed data reflects the changing trend and historical state of the motor speed, providing rich information for analyzing motor vibration patterns; the required torque at the current time is directly related to the motor's current power output demand. By acquiring these two types of key data, a solid foundation is laid for the subsequent accurate calculation of timing and frequency domain indicators and the generation of reasonable compensation torque, which helps improve the accuracy and effectiveness of the entire anti-vibration control method.
[0025] One possible implementation is that the time-series indicators include speed distribution, rate of change, and direction of fluctuation; the frequency-domain indicators include low-frequency ratio, high-frequency ratio, and frequency distribution. Distribution statistics clearly present the frequency distribution of motor speed within a period, reflecting the frequency of occurrence of speed within different value ranges, which helps to understand the central tendency and dispersion of motor speed. The rate of change, by calculating the sum of the absolute values of the differences in speed between adjacent sampling points, can intuitively reflect the speed of change of motor speed. The direction of fluctuation defines the upward and downward trends and amplitude within a period, accurately grasping the fluctuation direction of motor speed. These comprehensive time-series indicators characterize the time-series features of motor speed from multiple perspectives, providing rich and detailed information for the reinforcement learning strategy network, enabling it to more accurately judge the state of motor speed jitter and thus generate more suitable compensation torque. The low-frequency ratio and high-frequency ratio respectively statistically analyze the energy proportion of low-frequency and high-frequency components in the speed signal, helping to understand the energy distribution of motor speed jitter in different frequency bands and distinguish the main jitter frequency components; the frequency distribution statistics statistically analyze the energy distribution of each frequency component in the speed signal, comprehensively presenting the frequency domain overview of motor speed jitter. By analyzing these frequency domain indicators, reinforcement learning policy networks can gain a deeper understanding of the frequency domain characteristics of motor speed jitter and take more targeted compensation measures for jitter at different frequencies, thereby improving the effectiveness of anti-jitter control.
[0026] One possible implementation is that the reinforcement learning framework includes a policy network and a value network;
[0027] The policy network has a three-layer fully connected structure, including a first input layer, a first hidden layer, and a first output layer. The input is a state vector, and the output is the compensation torque. The first input layer receives a state vector composed of motor speed, required torque, time-series indicators, and frequency-domain indicators, enabling comprehensive acquisition of various information about motor operation. The first hidden layer can perform nonlinear transformations and feature extraction on the input information to uncover potential patterns in the data. The first output layer directly outputs the compensation torque, which is closely linked to the subsequent torque control stage. This three-layer fully connected structure ensures a certain level of network complexity and learning ability while avoiding excessive computation and training difficulties caused by overly complex structures, thus improving the network's training efficiency and practical application effectiveness.
[0028] The value network has a three-layer fully connected structure, including a second input layer, a second hidden layer, and a second output layer. The input is a state vector, and the output is a state value. The role of the value network is to evaluate the value of different actions taken by the policy network in different states. By calculating the state value, it measures the quality of the policy network generating compensating torque in the current state. The three-layer fully connected structure can effectively process and analyze the input state vector, generating accurate state values and providing a reliable basis for policy updates in the reinforcement learning process.
[0029] One possible implementation, during the offline training of the reinforcement learning model, also includes: setting the target loss to minimize the error of the state value, and training the neural network model until the loss function stabilizes or a stopping condition is triggered.
[0030] One possible approach is to reset the weights of the first output layer of the policy network and freeze the weights of other layers during online model fine-tuning. This approach has significant advantages: freezing the weights of other layers preserves the general features and knowledge learned by the model during the pre-training phase, which have a certain universality for analyzing motor speed jitter under different vehicle models and operating conditions; while resetting the weights of the first output layer allows the model to quickly adapt to changes in jitter characteristics caused by factors such as motor parameter drift and mechanical wear in the real vehicle environment. Fine-tuning of the compensation torque can be achieved simply by adjusting the weights of the first output layer, without the need for large-scale retraining of the entire model, greatly improving the efficiency and flexibility of online fine-tuning and reducing the cost and time of adapting the model to the real vehicle environment.
[0031] One possible implementation, during the online fine-tuning of the model, includes: when the required torque changes, obtaining speed feedback and dynamically adjusting the reward score in the reinforcement learning framework, retraining the neural network model until the loss function stabilizes or a stopping condition is triggered. In practical applications, the required torque of a vehicle constantly changes with driving conditions, driving operations, and other factors. This dynamic adjustment mechanism enables the neural network model to promptly perceive changes in required torque and adjust the reward score according to the new situation, guiding the neural network model to generate a compensation torque more suitable for the current operating conditions. By retraining the neural network model, it is ensured that the neural network model can quickly adapt to changes in required torque, maintaining good anti-shake control performance at all times, thus improving the adaptability and robustness of the entire vehicle motor anti-shake control method.
[0032] One possible implementation is that the reward score in the reinforcement learning framework is a weighted sum of a trend score, a stability score, and an energy consumption score, wherein the trend score, stability score, and energy consumption score are calculated as follows:
[0033] Trend score: The closer the motor speed is to its filter value, the higher the score;
[0034] Stability score: The smaller the differential fluctuation of the motor speed, the higher the score;
[0035] Energy Consumption Score: The smaller the compensation torque value, the higher the score. The reward score is a weighted sum of the trend score, stability score, and energy consumption score. The trend score increases the closer the motor speed is to its filter value, which helps guide the motor speed towards a more stable trend and reduces abnormal speed fluctuations. The stability score aims to minimize the differential fluctuations in motor speed, directly encouraging a reduction in the amplitude of motor speed fluctuations and improving motor operation stability. The energy consumption score increases the score by minimizing the compensation torque value, prompting the model to minimize the use of compensation torque while ensuring anti-shake effectiveness, thereby reducing motor energy consumption. This comprehensive reward score fully considers the goals of anti-shake control, ensuring stable motor operation while also taking energy-saving requirements into account, enabling the reinforcement learning policy network to be trained and optimized in a direction that is more conducive to improving motor performance.
[0036] Secondly, the vehicle motor anti-shake control system of the present invention includes:
[0037] Data acquisition module: used to acquire motor speed and required torque in real time;
[0038] Indicator calculation module: connected to the data acquisition module, used to calculate the time-series index and frequency domain index of motor speed based on the real-time acquired motor speed;
[0039] State vector generation module: connected to the data acquisition module and the index calculation module respectively, used to combine the acquired motor speed, demand torque, time series index and frequency domain index into a state vector;
[0040] The reinforcement learning policy network module contains a pre-trained neural network model. The reinforcement learning policy network module is connected to the state vector generation module and is used to receive the state vector and output compensation torque based on the input state vector.
[0041] Torque processing and control module: connected to the reinforcement learning policy network module and the data acquisition module respectively, used to superimpose the compensation torque with the demand torque acquired by the data acquisition module to generate a total torque, and transmit the total torque to the motor controller to perform torque control.
[0042] Thirdly, the vehicle described in this invention employs a vehicle motor anti-shake control system as described in this invention.
[0043] Fourthly, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores at least one computer program, and when the at least one computer program is loaded and executed by the processor, it can implement the steps of the vehicle motor anti-shake control method provided in the above-described method embodiments.
[0044] It should be noted that any of the possible implementations of any of the above aspects can be combined, provided that the solutions do not contradict each other.
[0045] The present invention has the following beneficial effects:
[0046] (1) Implement motor anti-shake function
[0047] This invention acquires the motor speed and required torque in real time, calculates the time-series and frequency-domain indices of the motor speed, combines these data into a state vector, inputs it into a pre-trained neural network model, and outputs the compensation torque, thereby realizing the motor anti-shake function.
[0048] (2) Applicable to various working conditions and vehicle models, reducing calibration difficulty
[0049] In this invention, an offline simulation environment is first established. A motor simulation model is used to simulate the dynamic behavior of the motor, generating simulation training data containing motor speed, required torque, and their time-series and frequency-domain indices under various typical operating conditions. A neural network model within a reinforcement learning framework is trained in this offline simulation environment. This allows the neural network model to access data from multiple operating conditions during the pre-training phase, learning the characteristics and patterns under different conditions. During online fine-tuning, only a small amount of real-vehicle data is needed to adapt to different vehicle models, eliminating the need to recalibrate PID parameters. This is because the model already possesses a certain degree of versatility during the pre-training phase, enabling it to quickly adapt to the vibration characteristics of different vehicle models in the real-vehicle environment. This significantly reduces engineering adaptation costs and solves the problems of long calibration cycles and high costs associated with existing methods.
[0050] (3) Reduce losses caused by large compensation torque
[0051] This invention acquires the motor speed and required torque in real time, calculates the time-series and frequency-domain indices of the motor speed, and combines these data into a state vector, which is then input into a pre-trained reinforcement learning policy network to output the compensation torque. Through real-time frequency domain analysis, the reinforcement learning policy network can more accurately understand the characteristics of motor speed fluctuations and generate just the right compensation torque based on the actual situation, avoiding over-compensation. This avoids introducing additional torque fluctuations, reduces current changes in the motor windings, thereby reducing motor copper losses and energy consumption, and solves the loss problem caused by excessive suppression of fluctuations in existing methods.
[0052] (4) Online fine-tuning to adapt to the real vehicle environment
[0053] In this invention, after deploying the offline-trained neural network model to a real vehicle environment, a real motor controller is used as environmental feedback to reset the weights of the first output layer of the neural network model and freeze the weights of other layers. The reward score is then dynamically optimized based on the frequency and amplitude of the vehicle's vibrations, and the neural network model is retrained. This online fine-tuning method allows the model to quickly perceive changes in vibration characteristics caused by factors such as motor parameter drift and mechanical wear in the real vehicle environment. It can quickly adapt to the vibration characteristics of the real vehicle by fine-tuning only the first output layer without destroying the general features learned during pre-training. This eliminates the need for periodic manual calibration, reduces maintenance costs, ensures the model's control performance in practical applications, and solves the problems associated with purely simulation-trained models. Attached Figure Description
[0054] Figure 1 This is a flowchart of the vehicle motor anti-shake control method described in the embodiments of this application;
[0055] Figure 2 This is a flowchart of the training process of the neural network model in the embodiments of this application;
[0056] Figure 3 This is a schematic diagram of the structure of the neural network model in the reinforcement learning framework in the embodiments of this application;
[0057] Figure 4 This is a comparison chart of the rotational speeds with and without compensated torque in the embodiments of this application;
[0058] Figure 5 This is a comparison chart of the compensation torque obtained through the PID method and the compensation torque obtained through reinforcement learning;
[0059] Figure 6 This is a schematic block diagram of the vehicle motor anti-shake control system described in the embodiments of this application;
[0060] Figure 7 This is a schematic block diagram of the electronic device described in the embodiments of this application;
[0061] In the diagram: 1. Data acquisition module, 2. Index calculation module, 3. State vector generation module, 4. Reinforcement learning policy network module, 5. Torque processing and control module, 6. Memory, 7. Processor. Detailed Implementation
[0062] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.
[0063] In the embodiments of this application, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different. The technical features described by "first" and "second" have no sequential or size order.
[0064] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0065] like Figure 1 As shown in the embodiments of this application, a vehicle motor anti-shake control method includes the following steps:
[0066] The motor speed and required torque are acquired in real time, and the time-series and frequency-domain indices of the motor speed are calculated.
[0067] The collected motor speed, required torque, timing parameters, and frequency domain parameters are combined into a state vector.
[0068] The state vector is input into the pre-trained neural network model, and the compensation torque is output.
[0069] The compensated torque is superimposed with the required torque to generate the total torque, which is then transmitted to the motor controller for torque control.
[0070] like Figure 2 As shown in the embodiments of this application, the training process of the neural network model includes the following steps:
[0071] Build an offline simulation environment: Simulate the dynamic behavior of the motor through a motor simulation model to generate simulation training data including demand torque, motor speed and time-series and frequency domain indices of motor speed.
[0072] Offline training of reinforcement learning model: Based on simulation training data, the neural network model in the reinforcement learning framework is trained in an offline simulation environment, using the motor simulation model as environmental feedback, until the neural network model meets the first preset requirement.
[0073] Online fine-tuning of the model: The offline-trained neural network model is deployed to the real vehicle environment, and the real motor controller is used as environmental feedback to fine-tune the neural network model until the second preset requirement is met.
[0074] In this embodiment, an offline simulation environment is first established. A motor simulation model is used to simulate the dynamic behavior of the motor, generating simulation training data containing motor speed, required torque, and their time-series and frequency-domain indices under various typical operating conditions. A neural network model within a reinforcement learning framework is trained in the offline simulation environment. This allows the neural network model to access data from multiple operating conditions during the pre-training phase, learning the characteristics and patterns under different conditions. During online fine-tuning, only a small amount of real-vehicle data is needed to adapt to different vehicle models, eliminating the need to recalibrate PID parameters. This is because the neural network model, after the pre-training phase, already possesses a certain degree of versatility and can quickly adapt to the vibration characteristics of different vehicle models in the real-vehicle environment, thereby significantly reducing engineering adaptation costs and solving the problems of long calibration cycles and high costs associated with existing methods. This method acquires motor speed and required torque in real time, calculates the time-series and frequency-domain indices of the motor speed, combines these data into a state vector, inputs it into the pre-trained neural network model, and outputs the compensation torque. Through real-time frequency-domain analysis, the neural network model can more accurately understand the characteristics of motor speed vibration and generate appropriate compensation torque based on actual conditions, avoiding over-compensation. This approach avoids introducing additional torque fluctuations, reduces current variations in the motor windings, and thus lowers motor copper losses and energy consumption, resolving the loss issues caused by excessive jitter suppression in existing methods. This method deploys the offline-trained neural network model to a real vehicle environment, uses a real motor controller as environmental feedback, fine-tunes the neural network model, and then retrains it. This online fine-tuning method allows the neural network model to quickly perceive jitter characteristics changes caused by factors such as motor parameter drift and mechanical wear in the real vehicle environment. This eliminates the need for periodic manual calibration, reduces maintenance costs, ensures the control performance of the neural network model in practical applications, and solves the problems associated with purely simulation-trained models.
[0075] In one possible embodiment, motor speed and required torque are acquired in real time, and time-series and frequency-domain indices of motor speed are calculated. The motor speed is historical data from time tn to time t-1; the required torque is the torque required at the current time t. Historical motor speed data reflects the trend and historical state of motor speed changes, providing rich information for analyzing motor vibration patterns; the required torque at the current moment is directly related to the motor's current power output demand. Acquiring these two types of key data lays a solid foundation for the subsequent accurate calculation of time-series and frequency-domain indices and the generation of reasonable compensation torque, contributing to improving the accuracy and effectiveness of the entire anti-vibration control method.
[0076] For example, time-series indicators include rotational speed distribution, rate of change, and direction of fluctuation, and the specific calculation method is as follows:
[0077] Speed distribution: Frequency distribution of motor speed within a statistical period.
[0078] Rate of change: Calculate the sum of the absolute values of the differences in rotational speed between adjacent sampling points.
[0079] Fluctuation direction: Defines the upward and downward trends and magnitude within the cycle.
[0080] Distribution statistics clearly present the frequency distribution of motor speed within a period, reflecting the frequency of occurrence of speed within different value ranges, which helps to understand the central tendency and dispersion of motor speed. The rate of change, calculated by summing the absolute differences in speed between adjacent sampling points, intuitively reflects the speed of change in motor speed. The direction of fluctuation defines the upward and downward trends and amplitude within a period, accurately grasping the fluctuation direction of motor speed. These comprehensive time-series indicators characterize the temporal features of motor speed from multiple perspectives, providing rich and detailed information for reinforcement learning policy networks, enabling them to more accurately determine the state of motor speed jitter and thus generate more appropriate compensation torque.
[0081] For example, frequency domain metrics include low-frequency proportion, high-frequency proportion, and frequency distribution, and the specific calculation methods are as follows:
[0082] Low-frequency ratio: The proportion of energy of low-frequency components in the statistical speed signal.
[0083] High-frequency ratio: The proportion of energy of high-frequency components in the statistical speed signal.
[0084] Frequency distribution: The energy distribution of each frequency component in a statistical rotational speed signal.
[0085] The low-frequency and high-frequency proportions, respectively, statistically analyze the energy percentages of low-frequency and high-frequency components in the speed signal. This helps to understand the energy distribution of motor speed jitter across different frequency bands and distinguish the main jitter frequency components. Frequency distribution statistics of the energy distribution of each frequency component in the speed signal provide a comprehensive picture of the motor speed jitter in the frequency domain. By analyzing these frequency domain indicators, reinforcement learning policy networks can gain a deeper understanding of the frequency domain characteristics of motor speed jitter and implement more targeted compensation measures for jitter at different frequencies, thereby improving the effectiveness of anti-jitter control.
[0086] In one possible embodiment, the reinforcement learning framework includes a policy network and a value network. The policy network has a three-layer fully connected structure, comprising a first input layer, a first hidden layer, and a first output layer. The input is a state vector, and the output is the compensation torque. The first input layer receives a state vector composed of motor speed, required torque, time-series indicators, and frequency-domain indicators, enabling comprehensive acquisition of various information about motor operation. The first hidden layer can perform nonlinear transformations and feature extraction on the input information to uncover potential patterns in the data. The first output layer directly outputs the compensation torque, which is closely linked to subsequent torque control stages. This three-layer fully connected structure ensures a certain level of complexity and learning capability while avoiding excessive computation and training difficulties caused by overly complex structures, thus improving the training efficiency and practical application effectiveness of the neural network model. The value network also has a three-layer fully connected structure, comprising a second input layer, a second hidden layer, and a second output layer. The input is a state vector, and the output is the state value. The value network evaluates the value of different actions taken by the policy network in different states, measuring the quality of the compensation torque generated by the policy network in the current state by calculating the state value. The three-layer fully connected structure can effectively process and analyze the input state vector, generate accurate state values, and provide a reliable basis for policy updates in the reinforcement learning process.
[0087] In one possible embodiment, the offline training of the reinforcement learning model further includes: setting the target loss to minimize the error of the state value, and training the neural network model until the loss function stabilizes or a stopping condition is triggered (i.e., a first preset requirement).
[0088] In one possible implementation, during online model fine-tuning, the weights of the first output layer of the policy network are reset, while the weights of other layers are frozen. This approach has significant advantages: freezing the weights of other layers preserves the general features and knowledge learned by the model during the pre-training phase, which have a certain degree of universality for analyzing motor speed jitter under different vehicle models and operating conditions; while resetting the weights of the first output layer allows the model to quickly adapt to changes in jitter characteristics caused by factors such as motor parameter drift and mechanical wear in the real vehicle environment. Fine-tuning of the compensation torque can be performed simply by adjusting the weights of the first output layer, without the need to retrain the entire neural network model, greatly improving the efficiency and flexibility of online fine-tuning and reducing the cost and time of adapting the neural network model to the real vehicle environment.
[0089] In one possible embodiment, during the online fine-tuning of the model, when the required torque changes, speed feedback is obtained, and the reward score is dynamically adjusted. The neural network model is then retrained until a second preset requirement is met (e.g., loss function stability or triggering a stopping condition). In practical applications, the vehicle's required torque constantly changes with driving conditions, driving operations, and other factors. This dynamic adjustment mechanism enables the neural network model to promptly perceive changes in required torque and adjust the reward score according to the new situation, guiding the neural network model to generate a compensation torque more suitable for the current operating conditions. By retraining the neural network model, it is ensured that the neural network model can quickly adapt to changes in required torque, maintaining good anti-shake control performance at all times, thus improving the adaptability and robustness of the entire vehicle motor anti-shake control method.
[0090] In one possible implementation, the reward score in the reinforcement learning framework is a weighted sum of the trend score, stability score, and energy consumption score, which are calculated as follows:
[0091] Trend score (i.e. filter reward): The closer the motor speed is to its filter value, the higher the score, with the goal of driving more and more stably.
[0092] Stability score (i.e. differential bonus): The smaller the differential fluctuation of the motor speed, the higher the score, and the goal is to drive more and more stably.
[0093] Energy consumption score (i.e. torque bonus): The smaller the compensation torque value, the higher the score, and the goal is to reduce energy consumption.
[0094] In this method, the reward score is a weighted sum of the trend score, stability score, and energy consumption score. The trend score increases as the motor speed closely matches its filtered value, which helps guide the motor speed towards a more stable trend and reduces abnormal speed fluctuations. The stability score aims to minimize differential fluctuations in motor speed, directly encouraging a reduction in the amplitude of speed fluctuations and improving motor operational stability. The energy consumption score increases by minimizing the value of the compensation torque, prompting the model to minimize the use of compensation torque while maintaining anti-shake effectiveness, thereby reducing motor energy consumption. This comprehensive reward score fully considers the goals of anti-shake control, ensuring stable motor operation while also taking energy-saving requirements into account, enabling the reinforcement learning policy network to be trained and optimized in a direction that better improves motor performance.
[0095] The following provides a detailed explanation of the process of building and training neural network models:
[0096] (1) Define the parameters and neural network model in the reinforcement learning framework.
[0097] like Figure 3 As shown, the reinforcement learning framework includes a motor model and two neural network models (i.e., a policy network and a value network).
[0098] like Figure 3 As shown, the policy network is defined as a fully connected network structure of n_states1 × n_hiddens1 × n_actions, where n_states1 represents the state dimension of the first input layer in the policy network, n_hiddens1 represents the number of neurons in the first hidden layer in the policy network, and n_actions represents the action parameter dimension of the first output layer in the policy network. The input of the policy network is the state, and the output of the policy network is the normal distribution parameters mu and std of the action (i.e., compensation torque), where mu is the mean and std is the standard deviation. In the PPO (Proximal Policy Optimization) type reinforcement learning framework, the action is a continuous value (represented in the form of a normal distribution, and randomly sampled from it as the unique numerical result). It can be understood that the reinforcement learning outputs a compensation torque range of a normal distribution (mu, std), and sets an expert threshold range. Within this compensation torque range & expert threshold range, a value is randomly sampled as the output value of the compensation torque for this time.
[0099] like Figure 3As shown, the value network is defined as a fully connected network structure of n_states2 × n_hiddens2 × action, where n_states2 represents the state dimension of the second input layer in the value network, n_hiddens2 represents the number of neurons in the second hidden layer in the value network, and action represents the output dimension of the second output layer in the value network. The input of the value network is the state, and the output is the state value.
[0100] Wherein, state = {rotational speed W at time t-1, required torque Tq at time t, rotational speed distribution from tn to t-1 (25th quantile / 50th quantile / 75th quantile), rotational speed fluctuation direction W_trd at time t-1, rotational speed fluctuation rate W_vol at time t-1, low-frequency ratio ratio_low at time t-1, high-frequency ratio ratio_high at time t-1, power spectral density psd at time t-1}. Here, fluctuation rate characterizes the rate of change, and power spectral density characterizes the frequency distribution.
[0101] Set the output action, action = {compensation torque ctrl_Tq at time t}, dimension n_actions.
[0102]
[0103] Where psd is the power spectral density, fft() is the Fourier transform function, and W t-n Let n be the rotational speed at time tn. s where fs is the signal length, fs is the sampling rate (Hz), power is the power spectral density including both positive and negative frequencies, and ratio is the ratio. low For low frequency proportion, w c For a custom cutoff frequency, n is the window length, and ratio is... high For high frequency proportion, w c This is the cutoff frequency. 0:n / / 2 indicates that the first half of the spectrum (i.e., the positive frequency and DC component) is extracted, ignoring the redundant part of the negative frequency.
[0104] (2) Reinforcement learning to train the neural network model.
[0105] Initialize the network: Initialize the policy network and the value network.
[0106] Action selection: Input the state into the policy network to obtain the normal distribution parameters mu and std of the compensation torque. Set the expert threshold [lower_thr, upper_thr], and randomly sample values within the expert threshold in the normal distribution space normal(mu, std) as the compensation torque ctrl_Tq at time t, specifically:
[0107] ctrl_Tq = Normal(mu, std)
[0108] ctrl_Tq∈[lower_thr,upper_thr]
[0109] Environmental Interaction: Input the compensation torque ctrl_Tq at time t and other driving data into the motor simulation model to obtain the speed at time t. At this time, the next state can be obtained: next_state = {speed W at time t, required torque Tq at time t+1, speed distribution from t-n+1 to time t (25th percentile / 50th percentile / 75th percentile), speed fluctuation direction W_trd at time t, speed fluctuation rate W_vol at time t, low frequency ratio ratio_low at time t, high frequency ratio ratio_high at time t, power spectral density psd at time t}.
[0110] Performance score: Calculate the reward score for this state transition based on the current state, the current action, and the next state.
[0111]
[0112] Where reward_filter is the trend score, w is the rotational speed obtained from the motor simulation model, n is the window length, abs() is the absolute value function, sg() is the SG filter function, and freq_mdl is the frequency of the motor simulation model (as the window value required for SG filtering). 100 is the upper limit of the trend score.
[0113]
[0114] Where reward_diff is the stable score; This represents the first-order difference of W within the window [tn, t-1]; std() is the standard deviation, which measures the degree of dispersion of the difference.
[0115] reward_Tq = 50 / ctrl_Tq
[0116] Here, reward_Tq represents the torque reward, 50 indicates that the maximum score for the torque reward is 50 points (meaning that the torque reward has a lower priority than the trend score and the stability score), and ctrl_Tq represents the compensation torque at time t.
[0117] reward=reward_filter+reward_diff+reward_Tq
[0118] Here, reward is the score for the current action in the current state. State value is the expected score of all future actions in all future states, given the current state.
[0119] Loss calculation: The advantage function is calculated based on the value network, and the policy ratio is calculated based on the policy network. The loss of the policy network is set as the state transition advantage, and the loss of the value network is set as the state transition.
[0120] value = critic_net(state)
[0121] Where, `value` represents the predicted value (expectation) of the state at time t, i.e., the state value at time t, which serves as the benchmark when calculating the advantage function. `critic_net()` is the value network;
[0122] target_value=reward+ε*critic_net(next_state)
[0123] Where target_value is the actual value (expected value) of the state at time t, critic_net(next_state) is the predicted value (expected value) of the state at time t+1, that is, the state value at time t+1; ε is the discount factor.
[0124] Calculate the advantage function:
[0125] advantage=∑(λε) n (target_value-value)
[0126] Here, advantage is the advantage function, used to evaluate the merits of taking an action in the current state relative to the average performance. λ is a hyperparameter. When Advantage > 0, it indicates better performance than the average; otherwise, it indicates worse performance than the average.
[0127] ratio=exp(log_prob_new-log_prob)
[0128] The ratio is the probability ratio of choosing an action under the new policy and the old policy in the same state. Here, log_prob is the logarithmic probability of choosing action under the old policy in the current state, and log_prob_new is the corresponding value for the new policy.
[0129] Based on the aforementioned Advantage, ratio, and target_value, the loss functions for the policy network and value network can be calculated respectively:
[0130] Actor_loss=avg(-min(ratio*advantage,clip(ratio,1±eps)*advantage))
[0131] The policy network loss, Actor_loss, employs a proximal policy optimization scheme to limit the changes in actions from being too large or too small, thus preventing training crashes. Here, clip() is the clipping function, eps are the clipping parameters, and avg() is the averaging function.
[0132] Critic_loss=avg(MSE(critic(states),target_value))
[0133] The value network loss Critic_loss approximates the true value (expectation) estimate and guides the updates of the policy network actor_net. Here, MSE() is the mean squared error function, and avg() is the average value function.
[0134] Offline training: Gradient updates continue until the reward score stabilizes and the loss decreases to a stable value, then the policy network is output. The rotational speed of the simulated motor is compared with the speed after reinforcement learning. If there is an anti-shake effect, the training is considered effective, and the policy network is stored.
[0135] (3) Online fine-tuning: The policy network has a three-layer fully connected structure (a first input layer, a first hidden layer and a first output layer). The weights of the first output layer are reset and the weights of other layers are frozen.
[0136] Acquire real-world driving data under different operating conditions, and optimize the reward score based on the frequency and amplitude of real-world vehicle vibrations. For example, update the filtering calculation in the trend score to reflect the real-world vehicle frequency.
[0137]
[0138] Here, `reward_filter_online` represents the trend score based on the actual vehicle frequency; after replacing the offline `reward_filter` in this stage, the weights of the first output layer are retrained using the new reward. `freq_car` represents the actual vehicle frequency.
[0139] Embed the trainable files of the model in a real vehicle or bench. Use the real motor controller as the environmental feedback in the reinforcement learning framework. Acquire speed feedback and calculate a new reward score every time the required torque changes. Retrain until the preset requirements are met, and store the policy network.
[0140] Finally, the trained policy network is converted into an application layer model and deployed in the vehicle controller.
[0141] Read the motor speed and required torque data during the vehicle's driving process within the time period tn to t, and combine the time-series index and frequency domain index of motor speed and required torque into a state vector.
[0142] The input to the policy network (i.e., the pre-trained neural network model) outputs the compensation torque at time T. This compensation torque is then superimposed on the required torque to generate the total torque, total_Tq, which is transmitted to the motor controller for torque control, thus achieving anti-jitter control. The formula for calculating the total torque, total_Tq, is as follows:
[0143] total_Tq = Tq + ctrl_Tq
[0144] Where total_Tq is the total torque, Tq is the required torque, and ctrl_Tq is the compensation torque.
[0145] like Figure 4 The image shows a comparison of the rotational speeds without added compensation torque (i.e., anti-shake torque) and without added compensation torque. The horizontal axis represents the number of recorded points, and the vertical axis represents rotational speed. The blue line represents the rotational speed obtained from the motor simulation model with the required torque input, and the red line represents the rotational speed obtained from the motor simulation model with the required torque plus reinforcement learning anti-shake torque input. Figure 4 As can be seen, after incorporating the anti-shake torque through reinforcement learning, the rotation speed becomes smoother, indicating that the anti-shake torque is effective.
[0146] like Figure 5 The image shows a comparison between the compensation torque obtained through the PID method and the compensation torque obtained through reinforcement learning. The horizontal axis represents the number of recording points, and the vertical axis represents the anti-shake torque. Figure 5 The green line represents the compensation torque (i.e., anti-shake torque) obtained by the PID method installed on the vehicle, while the black line represents the compensation torque obtained by reinforcement learning. Figure 5 As can be seen, reinforcement learning has smaller numerical fluctuations and lower energy consumption compared to the PID method used in vehicles to compensate for torque.
[0147] like Figure 6As shown in the embodiment of this application, a vehicle motor anti-shake control system includes a data acquisition module 1, an index calculation module 2, a state vector generation module 3, a reinforcement learning strategy network module 4, and a torque processing and control module 5. The data acquisition module 1 is used to acquire the motor speed and required torque in real time. The index calculation module 2 is connected to the data acquisition module 1 and is used to calculate the time-series index and frequency domain index of the motor speed based on the real-time acquired motor speed. The state vector generation module 3 is connected to both the data acquisition module 1 and the index calculation module 2, and is used to combine the acquired motor speed, required torque, time-series index, and frequency domain index into a state vector. The reinforcement learning strategy network module 4 contains a pre-trained reinforcement learning strategy network. The reinforcement learning strategy network module 4 is connected to the state vector generation module 3 and is used to receive the state vector and output the compensation torque based on the input state vector. The torque processing and control module 5 is connected to both the reinforcement learning strategy network module 4 and the data acquisition module 1, and is used to superimpose the compensation torque with the required torque acquired by the data acquisition module 1 to generate a total torque, and transmit the total torque to the motor controller to execute torque control.
[0148] In this embodiment of the application, a vehicle employs a vehicle motor anti-shake control system as described in this embodiment of the application.
[0149] like Figure 7 As shown, in another aspect, an electronic device is provided, which includes a memory 6 and a processor 7. The memory 6 stores at least one computer program. When the at least one computer program is loaded and executed by the processor 7, it can implement the steps of the vehicle motor anti-shake control method provided in the above-described method embodiments.
[0150] In another aspect, a computer-readable storage medium is provided, which stores at least one computer program. When the at least one computer program is loaded and executed by a processor, it can implement the vehicle motor anti-shake control method provided in the above-described method embodiments.
[0151] It should be noted that when one or more instructions in the computer-readable storage medium or computer program product are executed by the processor of an electronic device, they implement the various processes of the above method embodiments and achieve the same technical effect as the above method. To avoid repetition, they will not be described again here.
[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0153] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for controlling vibration reduction in a vehicle motor, characterized in that, Includes the following steps: The motor speed and required torque are acquired in real time, and the time-series and frequency-domain indices of the motor speed are calculated. The collected motor speed, required torque, timing parameters, and frequency domain parameters are combined into a state vector; The state vector is input into a pre-trained neural network model, and the compensation torque is output. The compensated torque is superimposed with the required torque to generate a total torque, which is then transmitted to the motor controller for torque control. The training process of the neural network model includes: Set up an offline simulation environment: Simulate the dynamic behavior of the motor through a motor simulation model to generate simulation training data; Offline training of reinforcement learning model: Based on the simulation training data, the neural network model in the reinforcement learning framework is trained in an offline simulation environment, with the motor simulation model as environmental feedback, until the neural network model meets the first preset requirement; Online fine-tuning of the model: The offline-trained neural network model is deployed to a real vehicle environment, and a real motor controller is used as environmental feedback to fine-tune the neural network model until the second preset requirement is met.
2. The vehicle motor anti-vibration control method according to claim 1, characterized in that, The simulation training data includes the required torque, motor speed, and time-series and frequency-domain indices of the motor speed.
3. The vehicle motor anti-vibration control method according to claim 2, characterized in that, The required torque is the required torque at the current time t; the motor speed is the historical motor speed data during the period from time tn to time t-1.
4. The vehicle motor anti-vibration control method according to claim 2, characterized in that, The time-series indicators include rotational speed distribution, rate of change, and direction of fluctuation; the frequency domain indicators include low-frequency ratio, high-frequency ratio, and frequency distribution.
5. The vehicle motor anti-vibration control method according to claim 1, characterized in that, The reinforcement learning framework includes a policy network and a value network; The policy network has a three-layer fully connected structure, including a first input layer, a first hidden layer and a first output layer. The input is a state vector and the output is the compensation torque. The value network has a three-layer fully connected structure, including a second input layer, a second hidden layer, and a second output layer. The input is a state vector, and the output is a state value.
6. The vehicle motor anti-vibration control method according to claim 5, characterized in that, The offline training of reinforcement learning models also includes: setting the target loss to minimize the error of the state value, and training the neural network model until the loss function stabilizes or the stopping condition is triggered.
7. The vehicle motor anti-vibration control method according to claim 5, characterized in that, During the online fine-tuning of the model, the weights of the first output layer of the policy network are reset, and the weights of other layers are frozen.
8. The vehicle motor anti-vibration control method according to claim 5, characterized in that, The online fine-tuning process also includes: when the required torque changes, obtaining speed feedback, dynamically adjusting the reward score in the reinforcement learning framework, and retraining the neural network model until the loss function stabilizes or the stopping condition is triggered.
9. The vehicle motor anti-vibration control method according to claim 8, characterized in that, The reward score is a weighted sum of the trend score, stability score, and energy consumption score. The trend score, stability score, and energy consumption score are calculated as follows: Trend score: The closer the motor speed is to its filter value, the higher the score; Stability score: The smaller the differential fluctuation of the motor speed, the higher the score; Energy consumption score: The smaller the compensation torque value, the higher the score.
10. A vehicle motor anti-shake control system, characterized in that, include: Data acquisition module (1): used to acquire motor speed and required torque in real time; Index calculation module (2): connected to the data acquisition module (1), used to calculate the time-series index and frequency domain index of motor speed based on the real-time acquired motor speed; State vector generation module (3): connected to the data acquisition module (1) and the index calculation module (2) respectively, used to combine the acquired motor speed, demand torque, time series index and frequency domain index into a state vector; Reinforcement learning strategy network module (4): It has a pre-trained neural network model in its memory. The reinforcement learning strategy network module (4) is connected to the state vector generation module (3) and is used to receive the state vector and output the compensation torque according to the input state vector. Torque processing and control module (5): It is connected to the reinforcement learning policy network module (4) and the data acquisition module (1) respectively, and is used to superimpose the compensation torque with the demand torque acquired by the data acquisition module (1) to generate a total torque, and transmit the total torque to the motor controller to perform torque control.
11. A vehicle, characterized in that, The vehicle motor anti-shake control system as described in claim 10 is adopted.
12. An electronic device, characterized in that, The electronic device includes a memory (6) and a processor (7), wherein the memory (6) stores at least one computer program, and when the at least one computer program is loaded and executed by the processor (7), it is capable of performing the steps of the vehicle motor anti-shake control method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Driving motor active anti-shake control method and system
CN116961494A
Vehicle-mounted motor anti-shake control method and device, vehicle and program product
CN118386867A