Vehicle motor anti-shake control method and system, vehicle and electronic equipment

The speed jitter problem of electric vehicle motor system is solved by generating compensation torque in real time through neural network model, which reduces calibration cost and energy consumption and improves the adaptability and robustness of motor anti-shake control.

CN120811202AActive Publication Date: 2025-10-17DEEPAL AUTOMOBILE NANJING RESEARCH INSTITUTE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511057765.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-10-17
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

The speed jitter problem of existing electric vehicle motor systems causes body vibration and noise, affecting ride comfort. Existing solutions increase cost and complexity, and rely on high-precision sensors, resulting in long calibration cycles and increased energy consumption.

Method used

A neural network model is used to obtain the motor speed and required torque in real time. Compensation torque is generated through offline simulation and online fine-tuning training to achieve motor anti-shake control and reduce calibration difficulty and loss.

Benefits of technology

The motor anti-shake function is realized, which is suitable for various working conditions and vehicle models, reduces calibration costs and energy consumption, and improves the adaptability and robustness of the control effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811202A_ABST
    Figure CN120811202A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of motor anti-shake, in particular to a vehicle motor anti-shake control method and system, a vehicle and electronic equipment.The method comprises the steps that the rotating speed and the required torque of a motor are obtained in real time, and the time sequence index and the frequency domain index of the rotating speed of the motor are calculated; combining the collected rotating speed of the motor, the required torque, the time sequence index and the frequency domain index into a state vector; inputting the state vector into a pre-trained neural network model, and outputting a compensation torque; superposing the compensation torque and the required torque to generate a total torque, and transmitting the total torque to a motor controller to execute torque control; wherein the training process of the neural network model comprises the steps of establishing an offline simulation environment, performing offline training on a reinforcement learning model and performing online fine tuning on the model. According to the invention, the anti-shake function of the motor can be realized, the calibration difficulty and cost are reduced, the loss is reduced, online fine adjustment and adaptation to a real vehicle can be realized, and the control effect is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of motor anti-shake, and particularly relates to a vehicle motor anti-shake control method and system, a vehicle and an electronic device. BACKGROUND

[0002] With the increasing attention to environmental protection and energy sustainable development around the world, the electric vehicle industry has ushered in an unprecedented development opportunity, with a continuously expanding market size and rapidly iterating technologies. Among the many performance indicators of electric vehicles, user driving experience has become a key area of focus for major automakers and research institutions. Good driving experience not only improves user satisfaction and loyalty to the product, but also helps the product gain a differentiated advantage in the fierce market competition, thereby promoting the healthy development of the entire industry.

[0003] Driving experience is a comprehensive concept that encompasses acceleration performance, handling stability, ride comfort, and other aspects of a vehicle. Ride comfort, as an important part of driving experience, is directly related to the physical and mental well-being of users during long-distance driving or daily commuting. A comfortable and quiet driving environment can effectively reduce driving fatigue and improve travel quality. Therefore, how to improve the ride comfort of electric vehicles has become an important problem that needs to be solved by the industry.

[0004] In the electric motor system, which is a core component of electric vehicles, the drive shaft and the wheels are usually connected through a rigid system. While this design ensures efficient power transmission, it also poses some potential problems. Within certain torque ranges, the electric motor may experience speed fluctuation during operation. Due to the rigid connection between the drive shaft and the wheels, this speed fluctuation is directly transmitted to the vehicle body without any attenuation, causing vibrations and noise in the vehicle body.

[0005] The vibrations of the vehicle body can cause passengers to feel significant jolts and shaking, especially during low-speed driving or frequent start-stop conditions, making the discomfort more pronounced. At the same time, the noise generated by the speed fluctuation can disrupt the quiet atmosphere inside the vehicle, interfere with passengers' communication and rest, and seriously affect the user's driving comfort. Prolonged exposure to such an unfavorable driving environment can also have a negative impact on passengers' physical health, such as causing dizziness, nausea, and other discomforts. Therefore, effectively suppressing the speed fluctuation of the electric motor system is of great significance to improving the driving comfort of electric vehicles.

[0006] Currently, for the problem of motor system speed jitter, the commonly used method is to control by adding a shock absorber element in the motor system. The shock absorber element can absorb and attenuate vibration energy, reduce the degree of speed jitter transmitted to the vehicle body, and thus improve the driving comfort to a certain extent. However, adding a shock absorber element requires redesigning and laying out the structure of the motor system, which not only increases the number and complexity of parts, but also puts higher requirements on the installation space and assembly process of the system. In addition, the cost of the shock absorber element itself is relatively high, including material cost, manufacturing cost and research and development cost, etc., which will directly lead to a substantial increase in the design cost of electric vehicles. In today's increasingly competitive market, high cost will undoubtedly weaken the market competitiveness of products and limit their large-scale popularization and application.

[0007] In addition to the equipment method, many existing speed jitter control schemes also highly depend on high-precision speed sensors and road sensors. These sensors can monitor the speed, torque of the motor and the condition of the road in real time, and feed these data back to the anti-jitter model of the control system to adjust the control strategy in time and suppress the speed jitter. However, this also raises a series of problems:

[0008] (1) For the anti-jitter model, the existing method usually needs to adjust the parameters for different vehicle models, loads and working conditions one by one, resulting in long calibration period and high cost.

[0009] (2) The existing method may introduce additional torque fluctuations due to excessive suppression of jitter, resulting in increased motor copper loss and energy consumption.

[0010] (3) The anti-jitter model trained purely by simulation is easily affected by motor parameter drift and mechanical wear, and needs to be calibrated periodically by manual.

[0011] Therefore, it is necessary to develop a new vehicle motor anti-jitter control method, system, vehicle and electronic equipment. SUMMARY

[0012] The purpose of the present application is to provide a vehicle motor anti-jitter control method, system, vehicle and electronic equipment, which can realize the motor anti-jitter function, reduce the calibration difficulty and cost, reduce the loss, and can be online fine-tuned and adapted to the real vehicle to ensure the control effect.

[0013] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0014] In a first aspect, the vehicle motor anti-jitter control method comprises the following steps:

[0015] Real-time acquisition of motor speed and required torque, and calculation of time sequence index and frequency domain index of motor speed;

[0016] The collected motor speed, demand torque, time sequence index and frequency domain index are combined into a state vector;

[0017] The state vector is input into a pre-trained neural network model, and a compensation torque is output;

[0018] The compensation torque is superimposed with the demand torque to generate a total torque and transmitted to the motor controller for torque control;

[0019] The training process of the neural network model includes:

[0020] An offline simulation environment is built: the dynamic behavior of the motor is simulated through a motor simulation model to generate simulation training data;

[0021] The reinforcement learning model is trained offline: based on the simulation training data, the neural network model in the reinforcement learning framework is trained in the offline simulation environment, the motor simulation model is used as the environment feedback, and the neural network model is trained until the first preset requirement is met;

[0022] The model is fine-tuned online: the offline trained neural network model is deployed to the real vehicle environment, the real motor controller is used as the environment feedback, and the neural network model is fine-tuned until the second preset requirement is met.

[0023] In a possible implementation, the simulation training data includes demand torque, motor speed, time sequence index and frequency domain index of the motor speed.

[0024] In a possible implementation, the demand torque is the demand torque at the current time t, and the motor speed is the historical motor speed data within the time period from time t-n to time t-1. The historical motor speed data can reflect the trend and historical state of the motor speed, providing rich information for analyzing the motor jitter rule. The demand torque at the current time is directly related to the current power output demand of the motor. By obtaining these two types of key data, a solid foundation is laid for subsequent accurate calculation of time sequence index, frequency domain index and generation of reasonable compensation torque, which helps to improve the accuracy and effectiveness of the entire anti-jitter control method.

[0025] In a possible implementation, the time sequence indicators include a speed distribution, a change speed, and a fluctuation direction; and the frequency domain indicators include a low frequency proportion, a high frequency proportion, and a frequency distribution. The distribution statistics can clearly present the frequency distribution of the motor speed in a period, reflect the frequency of the motor speed in different value ranges, and help understand the concentration trend and dispersion degree of the motor speed; the change speed can intuitively reflect the speed of the change of the motor speed by calculating the sum of the absolute values of the differences between adjacent sampling points; and the fluctuation direction can accurately grasp the fluctuation trend of the motor speed. These comprehensive time sequence indicators describe the time sequence characteristics of the motor speed from multiple angles, provide rich and detailed information for the reinforcement learning strategy network, enable the reinforcement learning strategy network to more accurately judge the state of the motor speed jitter, and further generate a more appropriate compensation torque. The low frequency proportion and the high frequency proportion respectively count the energy proportions of the low frequency and high frequency components in the motor speed signal, help understand the energy distribution of the motor speed jitter in different frequency bands, and distinguish the main jitter frequency components; and the frequency distribution counts the energy distribution of each frequency component in the motor speed signal, and can comprehensively present the frequency domain of the motor speed jitter. Through analysis of these frequency domain indicators, the reinforcement learning strategy network can deeply understand the frequency domain characteristics of the motor speed jitter, take more targeted compensation measures for jitter of different frequencies, and improve the effect of the anti-jitter control.

[0026] In a possible implementation, the reinforcement learning framework includes a strategy network and a value network.

[0027] The strategy network has a three-layer fully connected structure, includes a first input layer, a first hidden layer, and a first output layer, the input is a state vector, and the output is a compensation torque; the first input layer receives a state vector composed of the motor speed, the demand torque, the time sequence indicators, and the frequency domain indicators, and can comprehensively obtain various information of the motor operation; the first hidden layer can perform nonlinear transformation and feature extraction on the input information, and mine potential rules in the data; and the first output layer directly outputs the compensation torque, and is closely connected with a subsequent torque control link. The three-layer fully connected structure can ensure that the network has a certain complexity and learning ability, avoid problems such as excessive calculation and difficult training caused by an excessively complex structure, and be beneficial to improving the training efficiency and actual application effect of the network.

[0028] The structure of the value network is a three-layer fully connected structure, including a second input layer, a second hidden layer and a second output layer. The input is a state vector, and the output is a state value. The role of the value network is to evaluate the value of different actions taken by the policy network in different states. By calculating the state value, the performance of the policy network in generating compensation torque in the current state is measured. The three-layer fully connected structure can effectively process and analyze the input state vector to generate accurate state values, providing a reliable basis for policy updating in the reinforcement learning process.

[0029] In a possible implementation, during the process of offline training of the reinforcement learning model, the target loss is set to be the minimum error of the state value, and the neural network model is trained until the loss function is stable or a stop condition is triggered.

[0030] In a possible implementation, during the process of online fine-tuning of the model, the weights of the first output layer of the policy network are reset, and the weights of other layers are frozen. This approach has significant advantages. Freezing the weights of other layers can preserve the general features and knowledge learned by the model in the pre-training stage, which are applicable to the analysis of motor speed jitter in different vehicle models and working conditions. Resetting the weights of the first output layer allows the model to quickly adapt to changes in jitter characteristics caused by factors such as motor parameter drift and mechanical wear in real-world environments. By adjusting only the weights of the first output layer, the compensation torque can be fine-tuned without the need for large-scale retraining of the entire model, greatly improving the efficiency and flexibility of online fine-tuning and reducing the cost and time of adapting the model to real-world environments.

[0031] In a possible implementation, during the process of online fine-tuning of the model, the process further includes: when the demand torque changes, obtaining the speed feedback, dynamically adjusting the reward score in the reinforcement learning framework, and retraining the neural network model until the loss function is stable or a stop condition is triggered. In actual applications, the demand torque of a vehicle changes constantly due to factors such as driving conditions and driving operations. This dynamic adjustment mechanism allows the neural network model to perceive changes in demand torque in a timely manner and adjust the reward score based on the new situation, guiding the neural network model to generate compensation torque that is more suitable for the current working condition. By retraining the neural network model, it ensures that the neural network model can quickly adapt to changes in demand torque and maintain good anti-jitter control effect at all times, improving the adaptability and robustness of the entire vehicle motor anti-jitter control method.

[0032] In a possible implementation, the reward score in the reinforcement learning framework is the weighted sum of the trend score, the stability score and the energy consumption score, and the calculation method of the trend score, the stability score and the energy consumption score is as follows:

[0033] Trend score: the higher the closeness between the motor speed and its filtered value, the higher the score.

[0034] Stability score: the smaller the differential fluctuation of the motor speed, the higher the score;

[0035] Energy consumption score: the smaller the value of the compensation torque, the higher the score. The reward score is the weighted sum of the trend score, the stability score and the energy consumption score. The trend score makes the motor speed closer to its filtered value, which helps to guide the motor speed to develop towards a more stable trend and reduce abnormal fluctuations in the motor speed; the stability score directly encourages to reduce the fluctuation amplitude of the motor speed and improve the stability of the motor operation by aiming at the smaller the differential fluctuation of the motor speed, the higher the score; the energy consumption score promotes the model to minimize the use of compensation torque on the premise of ensuring the anti-shake effect, thereby reducing the energy consumption of the motor by making the smaller the value of the compensation torque, the higher the score. This comprehensive reward score can comprehensively consider the goal of anti-shake control, ensure the stable operation of the motor, and take into account the energy saving requirement, so that the reinforcement learning strategy network is trained and optimized in a direction more conducive to improving the performance of the motor.

[0036] In a second aspect, the vehicle motor anti-shake control system comprises:

[0037] A data acquisition module is configured to acquire the motor speed and the demand torque in real time.

[0038] An index calculation module is connected to the data acquisition module and configured to calculate the time series index and the frequency domain index of the motor speed according to the real-time acquired motor speed.

[0039] A state vector generation module is connected to the data acquisition module and the index calculation module, respectively, and configured to combine the acquired motor speed, demand torque, time series index and frequency domain index into a state vector.

[0040] A reinforcement learning strategy network module has a pre-trained neural network model in its memory, and is connected to the state vector generation module and configured to receive the state vector and output the compensation torque according to the input state vector.

[0041] A torque processing and control module is connected to the reinforcement learning strategy network module and the data acquisition module, respectively, and configured to superimpose the compensation torque and the demand torque acquired by the data acquisition module to generate a total torque, and transmit the total torque to the motor controller for torque control.

[0042] In a third aspect, the vehicle adopts the vehicle motor anti-shake control system as described in the present application.

[0043] In a fourth aspect, the electronic device comprises a processor and a memory, and the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the steps of the motor anti-shake control method provided by the above method embodiments.

[0044] It should be noted that the various possible implementations of any one of the above aspects can be combined as long as the schemes are not contradictory.

[0045] The present application has the following advantages:

[0046] (1) Realize motor anti-shake function

[0047] After the present application obtains the motor speed and the demand torque in real time, the time sequence index and the frequency domain index of the motor speed are calculated, and these data are combined into a state vector and input into a pre-trained neural network model, and a compensation torque is output, so that the motor anti-shake function is realized.

[0048] (2) Suitable for various working conditions and vehicle models, and reduce calibration difficulty

[0049] In the present application, an offline simulation environment is first built, the dynamic behavior of the motor is simulated through a motor simulation model, and simulation training data containing the motor speed, the demand torque, the time sequence index and the frequency domain index of the motor under various typical working conditions are generated. The neural network model in the reinforcement learning framework is trained in the offline simulation environment, so that the neural network model can be exposed to data of various working conditions in the pre-training stage and learn the characteristics and rules under different working conditions. During online fine-tuning, only a small amount of real vehicle data is needed to adapt to different vehicle models, and there is no need to recalibrate the PID parameters. This is because the model in the pre-training stage has a certain universality and can quickly adapt to the shaking characteristics of different vehicle models in the real vehicle environment, thereby greatly reducing the engineering adaptation cost and solving the problems of long calibration period and high cost of the existing methods.

[0050] (3) Reduce the loss problem caused by large compensation torque

[0051] After the present application obtains the motor speed and the demand torque in real time, the time sequence index and the frequency domain index of the motor speed are calculated, and these data are combined into a state vector and input into a pre-trained reinforcement learning strategy network to output a compensation torque. Through real-time frequency domain analysis, the reinforcement learning strategy network can more accurately understand the characteristics of the motor speed shaking, generate a just-right compensation torque according to the actual situation, and avoid the occurrence of excessive compensation. In this way, no additional torque fluctuation is introduced, the current change in the motor winding is reduced, the motor copper loss and energy consumption are reduced, and the loss problem caused by excessive suppression of shaking of the existing method is solved.

[0052] (4) Online fine-tuning adapts to real vehicle environment

[0053] In the present application, after deploying the offline trained neural network model to the real vehicle environment, using the real motor controller as the environment feedback, resetting the first output layer weight of the neural network model, freezing other layer weights, dynamically optimizing the reward score according to the frequency and amplitude of the real vehicle jitter, and retraining the neural network model. This online fine-tuning method enables the model to quickly perceive the jitter feature changes caused by factors such as motor parameter drift and mechanical wear in the real vehicle environment, and quickly adapt to the real vehicle jitter feature by fine-tuning only the first output layer, without destroying the general features learned by pre-training. This way, periodic manual calibration is not needed, reducing maintenance costs and ensuring the control effect of the model in actual application, solving the problem of pure simulation training model. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 is a flowchart of the vehicle motor anti-jitter control method described in the embodiments of the present application;

[0055] Figure 2 is a flowchart of the training process of the neural network model in the embodiments of the present application;

[0056] Figure 3 is a structural diagram of the neural network model in the reinforcement learning framework in the embodiments of the present application;

[0057] Figure 4 is a comparison chart of the rotational speed with and without compensation torque in the embodiments of the present application;

[0058] Figure 5 is a comparison chart of the compensation torque obtained by the PID method and the compensation torque obtained by reinforcement learning;

[0059] Figure 6 is a principle block diagram of the vehicle motor anti-jitter control system described in the embodiments of the present application;

[0060] Figure 7 is a principle block diagram of the electronic device described in the embodiments of the present application;

[0061] In the figure: 1, data acquisition module, 2, index calculation module, 3, state vector generation module, 4, reinforcement learning strategy network module, 5, torque processing and control module, 6, memory, 7, processor. DETAILED DESCRIPTION

[0062] Other advantages and embodiments of the application will be or will become apparent to one with skill in the art upon examination of the following figures and detailed description. It is intended that any such

[0063] In the embodiments of the present application, in order to clearly describe the technical solutions of the embodiments of the present application, the terms of "first", "second", etc. are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms of "first", "second", etc. do not limit the quantity and execution order, and the terms of "first", "second", etc. also do not limit the difference. The technical features described by the terms of "first", "second" have no sequence or size order.

[0064] In the embodiments of the present application, the words of "exemplary" or "for example" are used to represent an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. In fact, the words of "exemplary" or "for example" are intended to present the related concept in a specific manner, and facilitate understanding.

[0065] As shown in Figure 1 In the embodiments of the present application, a vehicle motor anti-shake control method comprises the following steps:

[0066] Real-time acquisition of motor speed and demand torque, and calculation of time sequence index and frequency domain index of motor speed.

[0067] Combination of the collected motor speed, demand torque, time sequence index and frequency domain index into a state vector.

[0068] Inputting the state vector into a pre-trained neural network model, and outputting a compensation torque.

[0069] Superimposing the compensation torque and the demand torque to generate a total torque and transmit it to a motor controller to perform torque control.

[0070] As shown in Figure 2 In the embodiments of the present application, the training process of the neural network model comprises the following steps:

[0071] Building an offline simulation environment: simulating the dynamic behavior of the motor through a motor simulation model to generate simulation training data including demand torque, motor speed and time sequence index and frequency domain index of motor speed.

[0072] Offline training of reinforcement learning model: based on simulation training data, train the neural network model in the reinforcement learning framework in the offline simulation environment, use the motor simulation model as the environment feedback, train the neural network model until the first preset requirement is met.

[0073] Online fine-tuning of model: deploy the offline trained neural network model to the real vehicle environment, use the real motor controller as the environment feedback, fine-tune the neural network model until the second preset requirement is met.

[0074] In the embodiments of the application, an offline simulation environment is first built, the dynamic behavior of the motor is simulated through the motor simulation model, and simulation training data containing motor speed, demand torque, time sequence indicators and frequency domain indicators under multiple typical working conditions are generated. The neural network model in the reinforcement learning framework is trained in the offline simulation environment, so that the neural network model can access data under multiple working conditions in the pre-training stage and learn the characteristics and rules under different working conditions. During online fine-tuning, only a small amount of real vehicle data is needed to adapt to different vehicle models, and there is no need to recalibrate the PID parameters. This is because the neural network model after the pre-training stage has a certain universality and can quickly adapt to the jitter characteristics of different vehicle models in the real vehicle environment, thereby greatly reducing the engineering adaptation cost and solving the problem of long calibration period and high cost of the existing method. After the motor speed and demand torque are obtained in real time, the time sequence indicators and frequency domain indicators of the motor speed are calculated, and these data are combined into a state vector and input into the pre-trained neural network model, and a compensation torque is output. Through real-time frequency domain analysis, the neural network model can more accurately understand the characteristics of the motor speed jitter, generate the appropriate compensation torque according to the actual situation, and avoid the occurrence of excessive compensation. In this way, additional torque fluctuations are not introduced, the current change in the motor winding is reduced, the motor copper loss and energy consumption are reduced, and the loss problem caused by excessive suppression of jitter in the existing method is solved. After the offline trained neural network model is deployed to the real vehicle environment, the neural network model is retrained after fine-tuning using the real motor controller as the environment feedback. This online fine-tuning method enables the neural network model to quickly perceive the jitter characteristic changes caused by factors such as motor parameter drift and mechanical wear in the real vehicle environment. In this way, periodic manual calibration is not required, maintenance costs are reduced, the control effect of the neural network model in actual application is guaranteed, and the problems of the pure simulation training model are solved.

[0075] In one possible embodiment, the motor speed and the demand torque are acquired in real time, and the time-domain index and the frequency-domain index of the motor speed are calculated. The motor speed is historical motor speed data in a time period from time t-n to time t-1; and the demand torque is the demand torque at the current time t. The historical motor speed data can reflect the change trend and historical state of the motor speed, and provide rich information for analyzing the motor jitter rule; and the demand torque at the current time is directly related to the current power output demand of the motor. By acquiring the two types of key data, a solid foundation is laid for subsequent accurate calculation of the time-domain index and the frequency-domain index, and generation of a reasonable compensation torque, which helps to improve the accuracy and effectiveness of the entire anti-jitter control method.

[0076] For example, the time-domain index includes the speed distribution, the change speed, and the fluctuation direction, and the specific calculation method is as follows:

[0077] Speed distribution: frequency distribution of the motor speed in the statistical period.

[0078] Change speed: sum of the absolute values of the differences between adjacent sampling points.

[0079] Fluctuation direction: definition of the uplink, downlink trend and amplitude in the period.

[0080] The distribution statistics can clearly present the frequency distribution of the motor speed in the period, reflect the frequency of the motor speed in different value ranges, and help to understand the concentration trend and dispersion degree of the motor speed. The change speed can intuitively reflect the speed of the motor speed change by calculating the sum of the absolute values of the differences between adjacent sampling points. The fluctuation direction can accurately grasp the fluctuation trend of the motor speed by defining the uplink, downlink trend and amplitude in the period. These comprehensive time-domain indexes characterize the time-domain characteristics of the motor speed from multiple angles, provide rich and detailed information for the reinforcement learning strategy network, enable it to more accurately judge the state of the motor speed jitter, and further generate a more suitable compensation torque.

[0081] For example, the frequency-domain index includes the low-frequency proportion, the high-frequency proportion, and the frequency distribution, and the specific calculation method is as follows:

[0082] Low-frequency proportion: statistics of the energy proportion of the low-frequency component in the speed signal.

[0083] High-frequency proportion: statistics of the energy proportion of the high-frequency component in the speed signal.

[0084] Frequency distribution: statistics of the energy distribution of each frequency component in the speed signal.

[0085] The low frequency proportion and the high frequency proportion respectively count the energy proportion of the low frequency component and the high frequency component in the speed signal, which is helpful to understand the energy distribution of the motor speed jitter in different frequency bands and distinguish the main jitter frequency component. The frequency distribution counts the energy distribution of each frequency component in the speed signal, which can comprehensively present the frequency domain of the motor speed jitter. Through the analysis of these frequency domain indicators, the reinforcement learning strategy network can deeply understand the frequency domain characteristics of the motor speed jitter, take more targeted compensation measures for different frequency jitter, and improve the effect of anti-jitter control.

[0086] In a possible embodiment, the reinforcement learning framework includes a strategy network and a value network. The structure of the strategy network is a three-layer fully connected structure, including a first input layer, a first hidden layer and a first output layer. The input is a state vector, and the output is a compensation torque. The first input layer receives a state vector composed of motor speed, demand torque, time domain indicators and frequency domain indicators, which can comprehensively obtain various information of the motor operation. The first hidden layer can perform nonlinear transformation and feature extraction on the input information to mine the potential rules in the data. The first output layer directly outputs the compensation torque, which is closely connected with the subsequent torque control link. The three-layer fully connected structure can ensure that the network has a certain complexity and learning ability, avoid the problems of excessive calculation and training difficulty caused by too complex structure, and is conducive to improving the training efficiency and practical application effect of the neural network model. The structure of the value network is a three-layer fully connected structure, including a second input layer, a second hidden layer and a second output layer. The input is a state vector, and the output is a state value. The role of the value network is to evaluate the value of different actions taken by the strategy network in different states, and to measure the advantages and disadvantages of the compensation torque generated by the strategy network in the current state by calculating the state value. The three-layer fully connected structure can effectively process and analyze the input state vector to generate accurate state value, providing reliable basis for policy update in the reinforcement learning process.

[0087] In a possible embodiment, in the process of training the reinforcement learning model offline, the target loss is set to be the minimum error of the state value, and the neural network model is trained until the loss function is stable or the stop condition is triggered (i.e. the first preset requirement).

[0088] In one possible embodiment, during the online fine-tuning of the model, the weights of the first output layer of the strategy network are reset and the weights of other layers are frozen. This approach has significant advantages. Freezing the weights of other layers can retain the common features and knowledge learned by the model during the pre-training phase. These features have a certain universality for motor speed jitter analysis under different vehicle models and working conditions. Resetting the weights of the first output layer can enable the model to quickly adapt to changes in jitter characteristics caused by factors such as motor parameter drift and mechanical wear in the actual vehicle environment. The compensation torque can be fine-tuned simply by adjusting the weights of the first output layer, without the need to retrain the entire neural network model. This greatly improves the efficiency and flexibility of online fine-tuning and reduces the cost and time of adapting the neural network model to the actual vehicle environment.

[0089] In one possible embodiment, during the online fine-tuning of the model, when the required torque changes, speed feedback is obtained, the reward score is dynamically adjusted, and the neural network model is retrained until the second preset requirement is met (for example, the loss function is stable or the stop condition is triggered). In actual applications, the vehicle's required torque will continue to change with factors such as driving conditions and driving operations. This dynamic adjustment mechanism enables the neural network model to promptly perceive changes in the required torque and adjust the reward score according to the new situation, guiding the neural network model to generate a compensation torque that is more suitable for the current working conditions. By retraining the neural network model, it is ensured that the neural network model can quickly adapt to changes in the required torque, always maintain a good anti-shake control effect, and improve the adaptability and robustness of the entire vehicle motor anti-shake control method.

[0090] In one possible embodiment, the reward score in the reinforcement learning framework is a weighted sum of the trend score, the stability score, and the energy consumption score. The trend score, the stability score, and the energy consumption score are calculated as follows:

[0091] Trend score (i.e., filter reward): The closer the motor speed is to its filtered value, the higher the score. The goal is to achieve more stable driving.

[0092] Stability score (i.e. differential reward): The smaller the differential fluctuation of the motor speed, the higher the score. The goal is to make driving more stable.

[0093] Energy consumption score (i.e. torque bonus): The smaller the value of the compensation torque, the higher the score. The goal is to reduce energy consumption.

[0094] In the method, the reward score is a weighted sum of the trend score, the stability score and the energy consumption score, the trend score is higher when the motor speed is closer to the filtered value, which helps to guide the motor speed to develop towards a more stable trend and reduce abnormal fluctuations in the motor speed. The stability score aims to reduce the fluctuation amplitude of the motor speed and improve the stability of the motor operation. The energy consumption score is higher when the compensation torque is smaller, which encourages the model to reduce the use of compensation torque as much as possible under the premise of ensuring the anti-shake effect, thereby reducing the energy consumption of the motor. This comprehensive reward score can comprehensively consider the goals of anti-shake control, ensure stable operation of the motor, and take into account energy saving requirements, so that the reinforcement learning strategy network is trained and optimized in a direction that is more conducive to improving the performance of the motor.

[0095] The building and training process of the neural network model are described in detail as follows:

[0096] (1) Define the parameters in the reinforcement learning framework and the neural network model.

[0097] As shown in Figure 3 , the reinforcement learning framework includes a motor model and two neural network models (i.e., a policy network and a value network).

[0098] As shown in Figure 3 , the policy network is defined as a fully connected network structure of n_states1x n_hiddens1x n_actions, where n_states1 represents the state dimension of the first input layer in the policy network, n_hiddens1 represents the number of neurons in the first hidden layer in the policy network, and n_actions represents the action parameter dimension of the first output layer in the policy network. The input of the policy network is the state state, and the output of the policy network is the normal distribution parameters mu and std of the action action (i.e., the compensation torque), where mu is the mean and std is the standard deviation. In the PPO (Proximal Policy Optimization) type of reinforcement learning framework, the action aciton is a continuous value (represented in the form of a normal distribution, and a random sample is taken as the only numerical result), which can be understood as the reinforcement learning outputting a compensation torque range of a normal distribution (mu, std), and setting an expert threshold range. Within the compensation torque range & expert threshold range, a value is randomly sampled as the output value of the compensation torque for this time.

[0099] As shown in Figure 3As shown, the value network is defined as a fully connected network structure of n_states2x n_hiddens2x action, where n_states2 represents the state dimension of the second input layer in the value network, n_hiddens2 represents the number of neurons in the second hidden layer in the value network, and action represents the output dimension of the second output layer in the value network. The input of the value network is the state state, and the output is the state value.

[0100] where state = {speed W at t-1 moment, demand torque Tq at t moment, speed distribution (25% percentile / 50% percentile / 75% percentile) at t-n~t-1 moment, speed fluctuation direction W_trd at t-1 moment, speed fluctuation rate W_vol at t-1 moment, low frequency ratio ratio_low at t-1 moment, high frequency ratio ratio_high at t-1 moment, power spectral density psd at t-1 moment}. Where the change speed is represented by the fluctuation rate, and the frequency distribution is represented by the power spectral density.

[0101] The output action action is set, action = {compensation torque ctrl_Tq at t moment}, dimension n_actions.

[0102]

[0103] where psd is the power spectral density, fft() is the Fourier transform function, W t-n is the speed at t-n moment, n s is the signal length, fs is the sampling rate (Hz), power is the power spectral density containing positive and negative frequencies, ratio low is the low frequency ratio, w c is the self-defined cutoff frequency, n is the window length, ratio high is the high frequency ratio, w c is the cutoff frequency. 0:n / / 2 represents the first half of the extracted frequency spectrum (i.e. positive frequency and direct current component), and the redundant part of the negative frequency is ignored.

[0104] (2) Reinforcement learning to train neural network model.

[0105] Initialize network: initialize policy network and value network.

[0106] Action selection: input state state into policy network to get normal distribution parameters mu, std of compensation torque. Set expert threshold [lower_thr, upper_thr], randomly sample a value within the expert threshold in the normal distribution space normal(mu, std) as the compensation torque ctrl_Tq at t moment, specifically:

[0107] ctrl_Tq = Normal(mu, std)

[0108] ctrl_Tq ∈ [lower_thr, upper_thr]

[0109] Environment Interaction: The compensation torque ctrl_Tq at time t and other driving data are input into the motor simulation model to obtain the speed at time t. At this time, the next state next_state = {speed W at time t, demand torque Tq at time t+1, speed distribution (25% percentile / 50% percentile / 75% percentile) from time t-n+1 to t, speed fluctuation direction W_trd at time t, speed fluctuation rate W_vol at time t, low frequency ratio ratio_low at time t, high frequency ratio ratio_high at time t, power spectral density psd at time t} can be obtained.

[0110] Performance Score: According to the current state state, the current action action, and the next state next_state, the reward score reward of the state transition is calculated.

[0111]

[0112] wherein reward_filter is the trend score, w is the speed obtained by the motor simulation model, n is the window length, abs() is the absolute value function, sg() is the sg filter function, freq_mdl is the frequency of the motor simulation model (as the window value input for sg filtering). 100 is the score upper limit of the trend score, which is 100 points.

[0113]

[0114] wherein reward_diff is the stability score; represents the first-order difference of W in the window [t-n, t-1]; std() is the standard deviation, which measures the dispersion degree of the difference.

[0115] reward_Tq = 50 / ctrl_Tq

[0116] wherein reward_Tq is the torque reward, 50 represents that the score upper limit of the torque reward is 50 points (indicating that the priority of the torque reward is lower than that of the trend score and the stability score), and ctrl_Tq represents the compensation torque at time t.

[0117] reward = reward_filter + reward_diff + reward_Tq

[0118] reward = score of current action in current state state. State value = expected score of action in all future states state from current state state.

[0119] Loss calculation: advantage function is calculated based on value network, and policy ratio is calculated based on policy network. Set the loss of policy network as state transition advantage, and the loss of value network as state transition.

[0120] value = critic_net(state)

[0121] Where value is the predicted value (expectation) of the state at time t, i.e. the state value at time t, which is used as a reference point when calculating the advantage function advantage. critic_net() is the value network;

[0122] target_value = reward + epsilon * critic_net(next_state)

[0123] Where target_value is the actual value (expectation) of the state at time t, and critic_net(next_state) is the predicted value (expectation) of the state at time t+1, i.e. the state value at time t+1; epsilon is the discount factor.

[0124] Calculate the advantage function advantage:

[0125] advantage = ∑(lambda * e) n (target_value - value)

[0126] Where advantage is the advantage function, which is used to evaluate the performance of taking action in the current state state relative to the average performance. Where lambda is a hyperparameter. When Advantage > 0, it means better than average performance, otherwise worse than average performance.

[0127] ratio = exp(log_prob_new - log_prob)

[0128] The ratio is the probability ratio of the new policy and the old policy selecting the action in the same state. Where log_prob is the logarithmic probability of the old policy selecting the action action in the current state state, and log_prob_new is the corresponding value of the new policy.

[0129] Based on the above Advantage, ratio, target_value, the loss function of the policy network and the value network can be calculated respectively:

[0130] Actor_loss=avg(-min(ratio*advantage,clip(ratio,1±eps)*advantage))

[0131] The policy network loss Actor_loss uses the proximal policy optimization scheme, which limits the change of action to avoid training collapse. Wherein, clip() is a clipping function, eps is a clipping parameter, avg() is an average function.

[0132] Critic_loss=avg(MSE(critic(states),target_value))

[0133] The value network loss Critic_loss approximates the true value (expectation) estimate, which guides the policy network actor_net update. Wherein, MSE() is a mean square error function, avg() is an average function.

[0134] Offline training: gradient update until the reward score performance is stable and the loss decreases to a stable value, then output the policy network. Compare the motor simulation model speed with the reinforcement learning speed, if there is anti-shake effect, consider the training effective, and store the policy network.

[0135] (3) Online fine-tuning: the structure of the policy network is a three-layer fully connected structure (one first input layer, one first hidden layer and one first output layer), reset the first output layer weight, freeze other layer weight.

[0136] Get real car driving data under different working conditions, and optimize the reward score according to the frequency and amplitude of real car shaking. For example, the filter calculation in the trend score is updated to the real car frequency.

[0137]

[0138] reward_filter_online represents the trend score based on the real car frequency; replace the offline reward_filter in this stage, and use the new reward to retrain the weight of the first output layer. freq_car represents the real car frequency.

[0139] In the real vehicle or bench, the trainable file of the model is embedded. The real motor controller is taken as the environment feedback in the reinforcement learning framework, the speed feedback is obtained at each demand torque change, the new reward score is calculated, retraining is carried out until the preset requirement is met, and the strategy network is stored.

[0140] Finally, the trained strategy network is converted into an application layer model and deployed in the vehicle-mounted controller.

[0141] The motor speed and demand torque data in the whole vehicle driving process in the t-n~t period are read, and the motor speed, demand torque, time sequence index and frequency domain index of the motor speed are combined into a state vector.

[0142] In the input strategy network (i.e. pre-trained neural network model), the compensation torque at T time is output through the strategy network. The compensation torque and the demand torque are superimposed to generate the total torque total_Tq and are transmitted to the motor controller to execute torque control, i.e. to realize the anti-shake control, and the calculation formula of the total torque total_Tq is as follows:

[0143] total_Tq = Tq + ctrl_Tq

[0144] Wherein, total_Tq is the total torque, Tq is the demand torque, and ctrl_Tq is the compensation torque.

[0145] As shown in Figure 4 , it is a comparison diagram of the corresponding speed after adding the compensation torque (i.e. anti-shake torque) and the corresponding speed without adding the compensation torque, wherein the horizontal axis is the recording point number, and the vertical axis is the speed, wherein the blue line is the speed obtained by inputting the demand torque into the motor simulation model, and the red line is the speed obtained by inputting the demand torque + reinforcement learning anti-shake torque into the motor simulation model. Through Figure 4 It can be seen that after adding the reinforcement learning anti-shake torque, the speed is more smooth, which shows that the anti-shake torque is effective.

[0146] As shown in Figure 5 , it is a comparison diagram of the compensation torque obtained by the PID method and the compensation torque obtained by the reinforcement learning, wherein the horizontal axis represents the recording point number, and the vertical axis represents the anti-shake torque, Figure 5 The green line in the green line is the compensation torque (i.e. anti-shake torque) obtained by the PID method carried on the vehicle, and the black line is the compensation torque obtained by the reinforcement learning. Through Figure 5 It can be seen that the compensation torque obtained by the reinforcement learning is smaller in value fluctuation and lower in energy consumption than the compensation torque obtained by the PID method carried on the vehicle.

[0147] As shown in Figure 6As shown, in the embodiment of the present application, a vehicle motor anti-shake control system includes a data acquisition module 1, an index calculation module 2, a state vector generation module 3, a reinforcement learning strategy network module 4, and a torque processing and control module 5. The data acquisition module 1 is used to acquire the motor speed and the required torque in real time. The index calculation module 2 is connected with the data acquisition module 1, and is used to calculate the time sequence index and the frequency domain index of the motor speed according to the real-time acquired motor speed. The state vector generation module 3 is connected with the data acquisition module 1 and the index calculation module 2 respectively, and is used to combine the acquired motor speed, the required torque, the time sequence index and the frequency domain index into a state vector. The reinforcement learning strategy network module 4 stores a pre-trained reinforcement learning strategy network, and is connected with the state vector generation module 3, used to receive the state vector and output the compensation torque according to the input state vector. The torque processing and control module 5 is connected with the reinforcement learning strategy network module 4 and the data acquisition module 1 respectively, used to superimpose the compensation torque and the required torque acquired by the data acquisition module 1 to generate a total torque, and transmit the total torque to the motor controller to perform torque control.

[0148] In the embodiment of the present application, a vehicle adopts the vehicle motor anti-shake control system as in the embodiment of the present application.

[0149] As shown in the above Figure 7 In another aspect, an electronic device is provided, the electronic device includes a memory 6 and a processor 7, the memory 6 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 7, and the steps of the vehicle motor anti-shake control method provided by the above method embodiments can be realized.

[0150] In another aspect, a computer readable storage medium is provided, the computer readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor, and the vehicle motor anti-shake control method provided by the above method embodiments can be realized.

[0151] It should be noted that the instructions in the above computer readable storage medium or one or more instructions in the computer program product are executed by the processor of the electronic device to realize the processes of the above method embodiments, and the same technical effects as the above method can be achieved. To avoid repetition, it will not be described here.

[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete the above described full classification part or part of the function.

[0153] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A vehicle motor anti-shake control method, characterized in that: The following steps are involved: Obtain motor speed and required torque in real time, and calculate the timing and frequency domain indicators of motor speed; Combine the collected motor speed, required torque, timing indicators and frequency domain indicators into a state vector; Inputting the state vector into a pre-trained neural network model and outputting a compensation torque; Adding the compensation torque to the required torque to generate a total torque and transmitting the total torque to the motor controller for torque control; The training process of the neural network model includes: Build an offline simulation environment: Use the motor simulation model to simulate the dynamic behavior of the motor and generate simulation training data; Offline training of the reinforcement learning model: Based on the simulation training data, training a neural network model in the reinforcement learning framework in an offline simulation environment, using the motor simulation model as environmental feedback, and training the neural network model until a first preset requirement is met; Online fine-tuning model: deploy the offline trained neural network model into a real vehicle environment, use the real motor controller as environmental feedback, and fine-tune the neural network model until it meets the second preset requirement.

2. The vehicle motor anti-shake control method according to claim 1, characterized in that: The simulation training data includes required torque, motor speed, and timing indicators and frequency domain indicators of the motor speed.

3. The vehicle motor anti-shake control method according to claim 2, characterized in that: The required torque is the required torque at the current time t; and the motor speed is the historical motor speed data from time tn to time t-1.

4. The vehicle motor anti-shake control method according to claim 2, characterized in that: The time series indicators include rotation speed distribution, change speed and fluctuation direction; the frequency domain indicators include low frequency ratio, high frequency ratio and frequency distribution.

5. The vehicle motor anti-shake control method according to claim 1, characterized in that: The reinforcement learning framework includes a policy network and a value network; The structure of the strategy network is a three-layer fully connected structure, including a first input layer, a first hidden layer and a first output layer, the input is a state vector, and the output is a compensation torque; The structure of the value network is a three-layer fully connected structure, including a second input layer, a second hidden layer and a second output layer. The input is a state vector and the output is a state value.

6. The vehicle motor anti-shake control method according to claim 5, characterized in that: The process of offline training of the reinforcement learning model also includes: setting the target loss to the minimum error of the state value, and training the neural network model until the loss function stabilizes or the stopping condition is triggered.

7. The vehicle motor anti-shake control method according to claim 5, characterized in that: During the online fine-tuning of the model, the weights of the first output layer of the policy network are reset and the weights of other layers are frozen.

8. The vehicle motor anti-shake control method according to claim 5, characterized in that: The online fine-tuning model process also includes: when the required torque changes, obtaining speed feedback, dynamically adjusting the reward score in the reinforcement learning framework, and retraining the neural network model until the loss function stabilizes or a stopping condition is triggered.

9. The vehicle motor anti-shake control method according to claim 8, characterized in that: The bonus score is a weighted sum of the trend score, stability score, and energy consumption score, which are calculated as follows: Trend score: The closer the motor speed is to its filtered value, the higher the score; Stability score: The smaller the differential fluctuation of the motor speed, the higher the score; Energy consumption score: The smaller the compensation torque value, the higher the score.

10. A vehicle motor anti-shake control system, characterized in that: include: Data acquisition module (1): used to obtain motor speed and required torque in real time; An index calculation module (2) is connected to the data acquisition module (1) and is used to calculate the timing index and frequency domain index of the motor speed based on the motor speed acquired in real time; A state vector generating module (3) is connected to the data acquisition module (1) and the index calculating module (2) respectively, and is used to combine the collected motor speed, required torque, timing index and frequency domain index into a state vector; Reinforcement learning strategy network module (4): a pre-trained neural network model is stored therein, and the reinforcement learning strategy network module (4) is connected to the state vector generation module (3) and is used to receive the state vector and output a compensation torque according to the input state vector; A torque processing and control module (5) is connected to the reinforcement learning strategy network module (4) and the data acquisition module (1) respectively, and is used to superimpose the compensation torque with the required torque collected by the data acquisition module (1) to generate a total torque, and transmit the total torque to the motor controller to perform torque control.

11. A vehicle, characterized in that: The vehicle motor anti-shake control system as claimed in claim 10 is adopted.

12. An electronic device, characterized in that: The electronic device comprises a memory (6) and a processor (7), wherein the memory (6) stores at least one computer program, and when the at least one computer program is loaded and executed by the processor (7), the steps of the vehicle motor anti-shake control method according to any one of claims 1 to 9 can be executed.

Citation Information

Patent Citations

  • Vehicle anti-shake control method and device, storage medium and vehicle

    CN116638981A

  • Driving motor active anti-shake control method and system

    CN116961494A

  • Vehicle-mounted motor anti-shake control method and device, vehicle and program product

    CN118386867A

  • Vehicle motor torque ripple suppression method and device

    CN119341434A

  • Vibration compensation controller with neural network band-pass filters for bearingless permanent magnet synchronous motor

    US20230008153A1