Automatic driving control method and device, medium and electronic equipment

Through reinforcement learning model and multi-objective cost function, the suspension, braking and steering control of the autonomous driving system are dynamically adjusted, which solves the problems of insufficient response capabilities and coordinated optimization of the existing system in emergencies, and achieves higher safety, comfort and economy.

CN120440071APending Publication Date: 2025-08-08VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510726293.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing autonomous driving systems are difficult to cope with sudden scenarios. The suspension system, braking system and steering system lack collaborative optimization, and the vehicle control method cannot be dynamically adjusted according to the physiological characteristics of the driver and passengers.

Method used

The reinforcement learning model is adopted to balance safety, comfort and economy through a multi-objective cost function, combine passenger physiological characteristics and usage scenario data, and dynamically adjust the control parameters such as suspension damping coefficient, longitudinal acceleration and steering angular rate to coordinately optimize the suspension, braking and steering system.

Benefits of technology

It improves the ability of the autonomous driving system to respond in emergencies, improves the safety, comfort and economy of the vehicle, and meets the personalized needs of different passengers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120440071A_ABST
    Figure CN120440071A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving control method and device, a medium and electronic equipment. The method comprises the steps that use scene data and / or user preference data of a vehicle are / is acquired; updating a multi-target cost function of a pre-established reinforcement learning model according to the use scene and / or the user preference data; acquiring environmental parameters and preset control parameters of the vehicle, and inputting the environmental parameters and the preset control parameters into the updated reinforcement learning model to output target control parameters; and controlling the vehicle to perform automatic driving according to the target control parameter. According to the invention, the emergency scene of the vehicle in the automatic driving process can be better dealt with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular, to an autonomous driving control method, device, medium, and electronic device. Background Art

[0002] An autonomous driving system refers to a system that uses advanced perception systems, decision-making algorithms and control technologies to enable a vehicle to complete environmental perception, path planning and dynamic control without human intervention.

[0003] However, existing autonomous driving systems have the following defects:

[0004] (1) Using fixed rules to control the chassis makes it difficult to cope with unexpected scenarios;

[0005] (2) When controlling the vehicle chassis, the suspension system, braking system, and steering system are controlled independently, lacking coordinated optimization;

[0006] (3) The vehicle control method cannot be dynamically adjusted according to the physiological characteristics of the vehicle driver and passengers. Summary of the Invention

[0007] The embodiments of the present application provide an autonomous driving control method, device, medium, and electronic device for solving the technical problem that existing autonomous driving systems use fixed rules to control the chassis and are difficult to cope with emergency scenarios.

[0008] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.

[0009] According to a first aspect of the present application, there is provided an automatic driving control method, the method comprising:

[0010] Obtaining vehicle usage scenario data and / or user preference data;

[0011] updating a multi-objective cost function of a pre-established reinforcement learning model based on the usage scenario and / or user preference data, wherein the reinforcement learning model is a framework for interaction between the vehicle and the environment, and its model structure is described by a multi-tuple, the multi-tuple including at least the multi-objective cost function, wherein the multi-objective cost function is used to balance the safety, comfort, and economy of the vehicle during autonomous driving;

[0012] Obtaining environmental parameters and preset control parameters of the vehicle, and inputting the environmental parameters and the preset control parameters into the updated reinforcement learning model to output target control parameters, wherein the preset control parameters include a suspension damping coefficient, longitudinal jerk, and steering angular rate, and the target control parameters include a suspension damping force, a front wheel steering angle, and a brake cylinder pressure;

[0013] The vehicle is controlled to perform automatic driving according to the target control parameters.

[0014] In some embodiments, based on the aforementioned solution, the multi-objective cost function is obtained by weighted summation of a first cost function, a second cost function, and a third cost function, wherein the first cost function is used to characterize the safety of the vehicle, the second cost function is used to characterize the comfort of the vehicle, and the third cost function is used to characterize the economy of the vehicle. Updating the multi-objective cost function of the pre-established reinforcement learning model according to the usage scenario and / or user preference data includes:

[0015] Based on the scenario adjustment mapping relationship, a weight adjustment set is obtained according to the usage scenario data, and / or based on the preference weight mapping relationship, a target weight set is obtained according to the user preference data, wherein the scenario adjustment mapping relationship includes: multiple scenarios and an adjustment set corresponding to each scenario, the adjustment set includes adjusting weights of the first cost function, the second cost function, and the third cost function, and the preference weight mapping relationship includes: multiple preferences and a weight set corresponding to each preference, the weight set includes weights of the first cost function, the second cost function, and the third cost function;

[0016] The weights of the first cost function, the second cost function and the third cost function are adjusted according to the weight adjustment set, and / or the weights of the first cost function, the second cost function and the third cost function are updated to the target weight set.

[0017] In some embodiments, based on the foregoing solution, updating the weights of the first cost function, the second cost function, and the third cost function according to the usage scenario data and / or the user preference data includes:

[0018] If the usage scenario data is an emergency obstacle avoidance scenario, increase the weight of the first cost function and decrease the weight of the second cost function;

[0019] If the usage scenario data is a steady cruising driving scenario, reducing the weight of the first cost function and increasing the weight of the third cost function;

[0020] If the usage scenario data is a bumpy road scenario or a congested following vehicle scenario, increasing the weight of the second cost function;

[0021] If the user preference data is a comfort preference, the weight of the second cost function is greater than the sum of the weight of the first cost function and the weight of the third cost function, and the weight of the first cost function is greater than the weight of the third cost function;

[0022] If the user preference data is an energy-saving preference, the weight of the second cost function is less than or equal to the sum of the weight of the first cost function and the weight of the third cost function, the weight of the second cost function is greater than the weight of the first cost function, and the weight of the first cost function is greater than the weight of the third cost function;

[0023] If the user preference data is motion preference, the weight of the third cost function is greater than or equal to the sum of the weight of the first cost function and the weight of the second cost function, and the weight of the second cost function is greater than or equal to the weight of the first cost function.

[0024] In some embodiments, based on the above solution, the preset control parameters are obtained through the following steps:

[0025] Acquiring multiple physiological characteristics of passengers in the vehicle and a current characteristic weight dataset, wherein the multiple physiological characteristics of passengers include heart rate variability, skin surface conductivity, three-dimensional head acceleration, and facial expression, and the current characteristic dataset includes current weights of the multiple physiological characteristics of passengers;

[0026] For each of the passenger physiological characteristics, obtain a first characteristic comfort score according to the passenger physiological characteristic, and obtain a second characteristic comfort score by performing the calculation on the first characteristic comfort score based on the weight of the passenger physiological characteristic;

[0027] Summing each of the second characteristic comfort scores to obtain a target comfort score;

[0028] Based on a pre-established score control mapping relationship, the preset control parameter is obtained according to the target comfort score, and the comfort score control mapping relationship includes: multiple comfort scores and control parameters corresponding to each comfort score.

[0029] In some embodiments of the present application, based on the above solution, the current feature weight dataset is obtained through the following steps:

[0030] Obtaining current passenger information of the vehicle;

[0031] Based on the passenger weight mapping relationship, a current feature weight data set is obtained according to the current passenger information. The passenger weight mapping relationship includes: multiple passenger information and a weight data set corresponding to each passenger information. The weight data set includes: weights of multiple passenger physiological characteristics.

[0032] In some embodiments, based on the aforementioned solution, the reinforcement learning model further includes a multi-objective reward function, and the method further includes:

[0033] The multi-objective reward function is used to replace the multi-objective cost function. The multi-objective reward function is the difference between the weighted sum of the first reward function, the second reward function, and the third reward function and the penalty function, wherein the first reward function reflects the safety of the vehicle through the collision risk and lane keeping condition of the vehicle, the second reward function characterizes the comfort of the vehicle through the lateral comfort, longitudinal comfort and road bumps of the vehicle, and the third reward function characterizes the traffic efficiency of the vehicle through the speed maintenance and energy efficiency of the vehicle.

[0034] In some embodiments, based on the above solution, the method further includes:

[0035] If the passenger of the vehicle is identified as a target passenger, obtaining a passenger reward function of the target passenger, wherein the passenger reward function is used to characterize a change in comfort of the target passenger and a change in energy consumption of the vehicle;

[0036] The passenger reward function is used to replace the multi-objective reward function.

[0037] In some embodiments, based on the aforementioned solution, controlling the vehicle to perform automatic driving according to the target control parameter includes:

[0038] Inputting the target control parameters into a pre-established chassis control model to obtain chassis control parameters, wherein the chassis control model is established based on a dynamic matrix control algorithm, and the chassis control parameters include longitudinal acceleration, lateral acceleration, vertical acceleration, vehicle body roll angle, and vehicle body pitch angle;

[0039] The suspension system, braking system and steering system of the vehicle are controlled according to the chassis control parameters.

[0040] According to a second aspect of the present application, there is provided an automatic driving control device, the device comprising:

[0041] A first acquisition unit acquires vehicle usage scenario data and / or user preference data;

[0042] a first updating unit configured to update a multi-objective cost function of a pre-established reinforcement learning model based on the usage scenario and / or user preference data, wherein the reinforcement learning model is a framework for interaction between the vehicle and the environment, and its model structure is described by a multi-tuple, the multi-tuple including at least the multi-objective cost function, wherein the multi-objective cost function is used to balance safety, comfort, and economy of the vehicle during autonomous driving;

[0043] a second acquisition unit, configured to acquire environmental parameters and preset control parameters of the vehicle, and input the environmental parameters and the preset control parameters into the updated reinforcement learning model to output target control parameters, wherein the preset control parameters include a suspension damping coefficient, a longitudinal jerk, and a steering angular rate, and the target control parameters include a suspension damping force, a front wheel steering angle, and a brake cylinder pressure;

[0044] A first control unit controls the vehicle to perform automatic driving according to the target control parameters.

[0045] In some embodiments, based on the aforementioned solution, the multi-objective cost function is obtained by weighted summation of a first cost function, a second cost function, and a third cost function, wherein the first cost function is used to characterize the safety of the vehicle, the second cost function is used to characterize the comfort of the vehicle, and the third cost function is used to characterize the economy of the vehicle. The first updating unit is configured as follows:

[0046] A first obtaining unit is configured to obtain a weight adjustment set according to the usage scenario data based on a scenario adjustment mapping relationship, and / or obtain a target weight set according to the user preference data based on a preference weight mapping relationship, wherein the scenario adjustment mapping relationship includes: a plurality of scenarios and an adjustment set corresponding to each scenario, the adjustment set includes adjusting weights of the first cost function, the second cost function, and the third cost function, and the preference weight mapping relationship includes: a plurality of preferences and a weight set corresponding to each preference, the weight set includes weights of the first cost function, the second cost function, and the third cost function;

[0047] The second updating unit adjusts the weights of the first cost function, the second cost function and the third cost function according to the weight adjustment set, and / or updates the weights of the first cost function, the second cost function and the third cost function to the target weight set.

[0048] In some embodiments, based on the above solution, the apparatus further includes a third acquisition unit, the third acquisition unit being configured to acquire the preset control parameter, and the third acquisition unit is configured to:

[0049] a fourth acquisition unit, configured to acquire a plurality of physiological characteristics of passengers in the vehicle and a current characteristic weight dataset, wherein the plurality of physiological characteristics of passengers include heart rate variability, skin surface conductivity, three-dimensional head acceleration, and facial expression, and the current characteristic dataset includes current weights of the plurality of physiological characteristics of passengers;

[0050] a second obtaining unit, for each of the passenger physiological characteristics, obtaining a first characteristic comfort score according to the passenger physiological characteristic, and obtaining a second characteristic comfort score by performing the calculation on the first characteristic comfort score based on the weight of the passenger physiological characteristic;

[0051] a third obtaining unit, summing the comfort scores of the second characteristics to obtain a target comfort score;

[0052] The fourth obtaining unit obtains the preset control parameter according to the target comfort score based on a pre-established score control mapping relationship, where the comfort score control mapping relationship includes: multiple comfort scores and control parameters corresponding to each comfort score.

[0053] In some embodiments, based on the foregoing solution, the apparatus further includes a fifth obtaining unit, the fifth obtaining unit being configured to obtain the current feature weight dataset, and the fifth obtaining unit being configured to:

[0054] a fifth acquiring unit, for acquiring current passenger information of the vehicle;

[0055] The sixth obtaining unit obtains the current feature weight data set based on the passenger weight mapping relationship and the current passenger information. The passenger weight mapping relationship includes: multiple passenger information and a weight data set corresponding to each passenger information. The weight data set includes: weights of multiple passenger physiological characteristics.

[0056] In some embodiments, based on the aforementioned solution, the reinforcement learning model further includes a multi-objective reward function, and the apparatus further includes:

[0057] A first replacement unit adopts the multi-objective reward function to replace the multi-objective cost function, wherein the multi-objective reward function is the difference between the weighted sum of the first reward function, the second reward function, and the third reward function and the penalty function, wherein the first reward function reflects the safety of the vehicle through the collision risk and lane keeping condition of the vehicle, the second reward function characterizes the comfort of the vehicle through the lateral comfort, longitudinal comfort and road bumps of the vehicle, and the third reward function characterizes the traffic efficiency of the vehicle through the speed maintenance and energy efficiency of the vehicle.

[0058] In some embodiments, based on the above solution, the device further includes:

[0059] a sixth acquiring unit, configured to acquire a passenger reward function of the target passenger if the passenger of the vehicle is identified as a target passenger, the passenger reward function being used to characterize a change in comfort of the target passenger and a change in energy consumption of the vehicle;

[0060] The second replacing unit adopts the passenger reward function to replace the multi-objective reward function.

[0061] In some embodiments, based on the aforementioned solution, the first control unit is configured as follows:

[0062] a seventh obtaining unit, inputting the target control parameters into a pre-established chassis control model to obtain chassis control parameters, wherein the chassis control model is established based on a dynamic matrix control algorithm, and the chassis control parameters include longitudinal acceleration, lateral acceleration, vertical acceleration, vehicle body roll angle, and vehicle body pitch angle;

[0063] The second control unit controls the suspension system, the braking system and the steering system of the vehicle according to the chassis control parameters.

[0064] According to a third aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program includes executable instructions. When the executable instructions are executed by a processor, the method described in any embodiment of the first aspect of the present application is implemented.

[0065] According to the fourth aspect of the present application, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions of the processors, wherein when the executable instructions are executed by the one or more processors, the one or more processors implement the method described in any embodiment of the first aspect of the present application.

[0066] The beneficial effects of this application are as follows:

[0067] When a vehicle encounters an emergency scenario, the multi-objective cost function can be updated according to the vehicle's usage scenario. The updated reinforcement learning model matches the usage scenario, and the target control parameters obtained can better cope with the emergency scenario.

[0068] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 A flowchart of an automatic driving control method according to an embodiment of the present application is shown;

[0070] Figure 2 A block diagram of an automatic driving control device according to an embodiment of the present application is shown;

[0071] Figure 3 A schematic diagram showing a computer-readable storage medium in an embodiment of the present application is shown;

[0072] Figure 4 A schematic diagram showing the system structure of an electronic device in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0073] To better understand the embodiments of this application, the following description of the autonomous driving related technologies is provided:

[0074] (1) Automatic deceleration of the vehicle when turning: The forward-looking camera detects the turning radius of the lane ahead and, combined with the vehicle's set speed and yaw rate, controls the speed of the vehicle when it is about to pass the curve. If the vehicle's lateral acceleration exceeds the set threshold, the resulting roll will cause a sense of oppression to the occupants (driver or passengers), and the vehicle will automatically reduce its speed to minimize the adverse driving effects caused by the roll.

[0075] (2) Scenarios with limited perception capabilities: In scenarios where the vehicle's radar and camera perception capabilities are limited, such as in heavy fog, rain, or snow, when the perception sensors cannot achieve their maximum capabilities, the autonomous driving system will autonomously reduce the vehicle's speed based on the actual driving scenario to ensure safe passage.

[0076] (3) Magic Carpet Function: When a vehicle is equipped with a CDC active suspension system, the magic carpet function can be realized through the air suspension, stereo camera, multi-axis speed sensor, acceleration sensor, vehicle height sensor, active hydraulic suspension system and its electronic control. That is, the magic carpet function can be realized by sensing the unevenness of the road in advance and actively adjusting the suspension height and hardness to improve ride comfort.

[0077] (4) Pre-stored multiple driving modes: The vehicle generates multiple driving modes for users to choose from, such as standard, aggressive, and efficiency-first. Different driving modes have different parameters that affect the driving experience, such as suspension height, maximum cruising speed, and deceleration for steep turns.

[0078] (5) Comfortable reclining function: The vehicle provides the driver or passenger with optimal riding comfort through electric adjustment, leg rest extension, lumbar support, heating and ventilation, headrest adjustment, intelligent control and safety logic. The driver or passenger can also adjust the seat automatically or manually to ensure comfort and safety, which is suitable for a variety of scenarios;

[0079] (6) Multiple scenarios: The vehicle provides a variety of scenario modes for users to choose from, including driving mode, comfort mode, environmental mode, safety mode, smart mode and energy-saving mode. Driving modes include economy, sports, comfort and off-road. Comfort modes include reclining, zero gravity and massage. Environmental modes include winter and summer. Safety modes include night and rainy days. Smart modes include automatic and personalized. Energy-saving modes include pure electric and hybrid.

[0080] Figure 1 A flowchart of an automatic driving control method in an embodiment of the present application is shown. Figure 1, this application provides an automatic driving control method, which includes at least S1 to S4, which are described in detail as follows:

[0081] In step S1 , vehicle usage scenario data and / or user preference data are obtained.

[0082] In step S2, a multi-objective cost function of a pre-established reinforcement learning model is updated based on the usage scenario and / or user preference data. The reinforcement learning model is a framework for the interaction between the vehicle and the environment, and its model structure is described by a multi-tuple that includes at least the multi-objective cost function. The multi-objective cost function is used to balance the safety, comfort, and economy of the vehicle during autonomous driving. In other words, the reinforcement learning model models driving decisions as a partially observable Markov decision process.

[0083] In some embodiments, the multi-tuple is a seven-tuple, comprising a state space, an action space, a multi-objective cost function, a state transition function, a discount factor, an initial state distribution, and a time step. The state space is used to constrain the vehicle's dynamic state, environment, and user preferences; the action space is used to represent the vehicle's executable actions; and the control strategy is the vehicle's decision-making mechanism. In this application, the vehicle can be understood as an intelligent agent.

[0084] In some embodiments, the reinforcement learning model further includes a state encoder, which includes a cross-modal attention layer that multiplies weights with sensor features and sums them to generate a fusion result, and dynamically allocates sensor weights using the following formula: The meaning is the final feature vector after weighted fusion, H senor It means the concatenation matrix of multi-sensor feature vectors, where each column corresponds to the feature vector of a sensor (such as lidar point cloud features, previous camera image features, millimeter wave radar target information). t The meaning of is the context feature vector of the current time step (t), which represents the current global state of the autonomous driving system (such as vehicle speed, road curvature, and comfort score). q and W k The meaning score is represented by a learnable query and key weight matrix, W q h t Mapped into a query vector, which is used to “ask” which sensors need to be paid attention to at present, W k H senor Mapped to a key vector, which is used to "answer" the importance of each sensor.

[0085] Thus, through W q and W kThe synergy of the sensors can automatically adjust the credibility of each sensor according to environmental changes, such as sensor failure and weather interference.

[0086] In some embodiments, the calculation function of the state space is as follows:

[0087] s t =[v,a x ,a y ,a z ,θ roll ,θ pitch ,d front ,ρ road ,C vib ,P user ],s t is the function value of the state space at time t, v is the vehicle speed, a x is the vehicle longitudinal acceleration, a y is the vehicle lateral acceleration, a z is the vertical acceleration, θ roll is the vehicle roll angle, θ pitch is the vehicle pitch angle, d front is the distance between the vehicle and the preceding vehicle or the time to collision (TTC) between the vehicle and the preceding vehicle, ρ road is the curvature of the road in the traversable area perceived in front of the vehicle, C vib Vehicle vibration comfort score, P user is the user preference coefficient.

[0088] In some embodiments, the calculation function of the action space is as follows:

[0089] a t =[Δδ steer , ΔP brake , ΔT throttle , m lane ], a t is the function value of the action space at time t, Δδ steer is the vehicle front wheel turning angle increment, ΔP brake is the brake cylinder pressure increment, ΔT throttle is the accelerator pedal opening increment, m lane is the discrete lane change decision. steer It can also be understood as the continuous value of the vehicle's steering EPS angle increment.

[0090] In some embodiments, the method further comprises: if m lane is 0, the vehicle is in lane keeping state; if m lane is 1, the vehicle is changing lanes to the left; if m lane is 2, the vehicle is in the state of changing lanes to the right.

[0091] In step S3, the environmental parameters and preset control parameters of the vehicle are obtained, and the environmental parameters and the preset control parameters are input into the updated reinforcement learning model to output target control parameters, wherein the preset control parameters include the suspension damping coefficient, longitudinal acceleration, and steering angle rate, and the target control parameters include the suspension damping force, front wheel angle, and brake cylinder pressure.

[0092] In some embodiments, the environmental parameters include an environmental semantic map and a road roughness index. These environmental parameters are obtained by: obtaining the environmental semantic map based on detection data from a forward-looking camera, a lidar, and a millimeter-wave radar; and obtaining the road roughness index by acquiring detection data from the vehicle's speed, vertical acceleration, and vibration sensor. The forward-looking camera's detection data consists of a video feed of the road ahead, with a frame rate of 30 fps. The lidar's detection data consists of point cloud data of obstacles ahead, with a sampling frequency of 20 Hz. The millimeter-wave radar's detection data consists of the distance, azimuth, pitch angle, and Doppler velocity of obstacles ahead, with a resolution of 1°×1°. The road vibration sensor's detection data includes displacement data from a multi-axis IMU and a suspension displacement sensor, both of which have an accuracy of 0.1 mm and a response time of less than 1 ms.

[0093] In some embodiments, the method further includes aligning detection data from a forward-looking camera, a lidar, and a millimeter-wave radar. This data alignment involves hardware-level synchronization and software compensation. Hardware-level synchronization utilizes the PTP (IEEE 1588) protocol to send synchronization pulse signals to each sensor. Software compensation aligns asynchronous data differences, using a cubic spline interpolation method to compensate for time deviations.

[0094] In some embodiments, the method further includes: expressing the detection data of the forward-looking camera, the laser radar, and the millimeter-wave radar in a unified coordinate system. The expression can be a calibration parameter storage, i.e., a pre-stored sensor extrinsic parameter matrix, such as the transformation matrix from the laser radar to the vehicle body coordinate system: R = RotationMatrix (yaw = 0°, pitch = 2°, roll = -1°). This can also be used for dynamic calibration of online algorithms based on feature point matching, such as minimizing reprojection errors of lane lines and traffic signs: π cam For the camera to affect the projection model, and are the matching feature points of the lidar and the front-view camera.

[0095] In some embodiments, the road roughness index is determined by the following formula: v is the vehicle speed, az (t) is the vertical acceleration of the vehicle at time t, β is the attenuation coefficient, which can be 0.2, and w k is the weight coefficient.

[0096] In some embodiments, the method further includes: health detection and degradation fusion of the detection data of the front-view camera, the laser radar and the millimeter wave radar. The health detection of the laser radar can be the effective point cloud ratio, i.e. N points is the number of valid points in the LiDAR point cloud of the current frame, N max is the maximum theoretical number of points in a single LiDAR frame, P vaild If the value is less than 50 and lasts for 5 frames, it is considered as failure. The health detection of the front camera can be based on the image entropy mutation H analysis. p i is the probability distribution of gray value i in the pseudo-color image, if at time t |H t -H t-1 If |>2.0, an abnormality is detected. If the LiDAR fails, the weight of the millimeter-wave radar detection data is increased. The detection data of the millimeter-wave radar is coupled with the detection data of the forward-looking camera, and the 3D position of the target is calculated based on the EPnP algorithm. If the forward-looking camera fails, the LiDAR reflection intensity pseudo-color imaging is activated to replace the visual semantic information, namely the detection data of the forward-looking camera.

[0097] In some embodiments, the preset control parameters also include air conditioning air supply intensity.

[0098] In step S4, the vehicle is controlled to perform automatic driving according to the target control parameters.

[0099] In some embodiments, based on the aforementioned scheme, the multi-objective cost function is obtained by weighted summation of a first cost function, a second cost function and a third cost function, the first cost function is used to characterize the safety of the vehicle, the second cost function is used to characterize the comfort of the vehicle, and the third cost function is used to characterize the economy of the vehicle. The multi-objective cost function of the pre-established reinforcement learning model is updated according to the usage scenario and / or user preference data, including: updating the weights of the first cost function, the second cost function and the third cost function according to the usage scenario data and / or user preference data.

[0100] In some embodiments, the method further includes: taking the first moment as the starting point and the second moment as the end point, selecting multiple target moments in chronological order, and for each target moment, obtaining the lateral deviation between the actual driving trajectory of the vehicle and the expected trajectory planned by automatic driving, as well as the square of the longitudinal acceleration of the vehicle, taking the sum of the various lateral deviations as the function value of the first cost function at the second moment, taking the sum of the squares of the longitudinal accelerations of the various vehicles as the function value of the second cost function at the second moment, and taking the vehicle braking energy consumption from the first moment to the second moment as the function value of the third cost function at the second moment.

[0101] Exemplarily, the second moment is the current moment, the first moment is 2s before the current moment, the function value of the first cost function at the current moment can be understood as the sum of the lateral deviations within 2s from the first moment to the second moment, the function value of the second cost function at the current moment can be understood as the sum of the vehicle longitudinal accelerations within 2s from the first moment to the second moment, and the function value of the third cost function at the current moment can be understood as the vehicle braking energy consumption within 2s from the first moment to the second moment.

[0102] In this way, the safety of the vehicle is reflected by the lateral deviation. If the lateral deviation is large, such as greater than 0.3m, the vehicle may be at risk of collision or lane departure during autonomous driving; the vehicle's longitudinal jerk measures the rate of change of the vehicle's acceleration, characterizes the severity of the acceleration change, and reflects the comfort of the driver and passengers of the vehicle. For example, during sudden acceleration or braking, the vehicle's longitudinal jerk will increase sharply, which can easily cause passenger discomfort such as motion sickness; the vehicle's braking energy consumption reflects the vehicle's fuel economy or power battery consumption, including mechanical braking energy loss and energy recovery. For electric vehicles, energy recovery systems such as regenerative braking can partially offset energy consumption. In actual application, it reduces the vehicle's braking energy consumption. During autonomous driving, it can reduce unnecessary braking operations and improve energy utilization.

[0103] It should be noted that the purpose of squaring the vehicle's longitudinal acceleration is to punish violent fluctuations so that autonomous driving tends to adopt a smoother acceleration and deceleration strategy.

[0104] In some embodiments, the calculation formula of the multi-objective cost function is as follows:

[0105] J=w1·J1+w2·J2+w3·J3, J is the function value of the multi-objective cost function, J1 is the first cost function, w1 is the weight of the first cost function, J2 is the second cost function, w2 is the weight of the second cost function, J3 is the third cost function, and w3 is the weight of the third cost function.

[0106] In some embodiments, the calculation formula of the first cost function is as follows: J 1t is the function value of the first cost function at time t, t = nt0, It is the lateral deviation between the actual vehicle trajectory at time it0 and the expected trajectory planned by autonomous driving.

[0107] In some embodiments, the calculation formula of the second cost function is as follows: J 2t is the function value of the second cost function at time t, t = nt0, jerk i is the vehicle longitudinal acceleration at time it0.

[0108] In some embodiments, the calculation formula of the third cost function is as follows: 3t =E break , J 3t is the function value of the third cost function at time t, E break The braking energy consumption of the vehicle from time t0 to time t, t=nt0.

[0109] In some embodiments, the multi-objective cost function is obtained by weighted summation of a first cost function, a second cost function, and a third cost function, wherein the first cost function is used to characterize the safety of the vehicle, the second cost function is used to characterize the comfort of the vehicle, and the third cost function is used to characterize the economy of the vehicle. The updating of the multi-objective cost function of the pre-established reinforcement learning model according to the usage scenario and / or user preference data includes: based on a scenario adjustment mapping relationship, obtaining a weight adjustment set according to the usage scenario data, and / or based on a preference weight mapping relationship, obtaining a target weight set according to the user preference data, wherein the scenario adjustment mapping relationship includes: multiple scenarios and an adjustment set corresponding to each scenario, the adjustment set includes adjusting the weights of the first cost function, the second cost function, and the third cost function, the preference weight mapping relationship includes: multiple preferences and a weight set corresponding to each preference, the weight set includes the weights of the first cost function, the second cost function, and the third cost function; adjusting the weights of the first cost function, the second cost function, and the third cost function according to the weight adjustment set, and / or updating the weights of the first cost function, the second cost function, and the third cost function to the target weight set. The weights of the first cost function, the second cost function, and the third cost function may be adjusted to increase, decrease, or remain unchanged.

[0110] In some embodiments, the scene mapping relationship is: if the scene is an emergency obstacle avoidance scene, the adjustment set is to increase the weight of the first cost function and reduce the weight of the second cost function; if the scene is a smooth cruising driving scene, reduce the weight of the first cost function and increase the weight of the third cost function; if the scene is a bumpy road scene or a congested following scene, increase the weight of the second cost function; if the scene is a highway navigation scene or a lane-centering driving scene, increase the weight of the first cost function; if the scene is a high-speed cruising scene, increase the weight of the third cost function.

[0111] For example, the scenario is a congested following scenario. During the autonomous driving process, comfort is prioritized over safety and economy. The weight of the first cost function is increased, and the vehicle longitudinal acceleration is limited to less than or equal to 1.5 km / s. 3 , avoid the frustration caused by frequent starting and stopping, dynamically increase the safe distance, and adjust to increase the vehicle's longitudinal acceleration ≥ 2.2km / s 3 , ensure a safe distance from the vehicle in front, reduce the weight of the third cost function, allow moderate energy loss to prioritize comfort needs; the scenario is a highway pilot scenario. During autonomous driving, economy is equal to safety and greater than comfort. Increase the weight of the third cost function, optimize the motor output curve, reduce the number of braking times, improve energy recovery efficiency, maintain the weight of the first cost function, ensure lane centering accuracy, make the lateral deviation less than 0.2m, slightly increase the vehicle longitudinal acceleration, make the vehicle longitudinal acceleration less than or equal to 2.0km / s 3 , to improve traffic efficiency; the scenario is a bumpy road scenario. During autonomous driving, comfort is prioritized over safety and economy. The weight of the first cost function is greatly increased. Through active adjustment of the suspension damping, vibration is suppressed from being transmitted to the vehicle cabin. By suppressing and constraining the lateral swing amplitude, the vehicle's longitudinal acceleration is kept less than or equal to 1.0 km / s. 3 , reduce the vehicle speed to reduce bumpy impact, ignore energy consumption optimization, make the weight of the third cost function 0, and give priority to passenger comfort.

[0112] In this way, increasing the weight of the first cost function can force the vehicle to strictly follow the trajectory and ensure the safety of the vehicle; increasing the weight of the second cost function can significantly suppress acceleration and improve ride comfort; increasing the weight of the third cost function can optimize the motor output curve and improve energy utilization.

[0113] It should be noted that excessive energy conservation should be avoided, which may result in a decrease in the comfort or safety of the vehicle, such as extending the following distance to reduce braking.

[0114] In some embodiments, the preference weight mapping relationship is: if the preference is comfort, the weight set satisfies the following conditions: the weight of the second cost function is greater than the sum of the weight of the first cost function and the weight of the third cost function, and the weight of the first cost function is greater than the weight of the third cost function; if the preference is energy saving, the weight set satisfies the following conditions: the weight of the second cost function is less than or equal to the sum of the weight of the first cost function and the weight of the third cost function, the weight of the second cost function is greater than the weight of the first cost function, and the first cost function is greater than the weight of the third cost function; if the preference is exercise, the weight set satisfies the following conditions: the weight of the third cost function is greater than or equal to the sum of the weight of the first cost function and the weight of the second cost function, and the weight of the second cost function is greater than or equal to the weight of the first cost function.

[0115] For example, the preference is comfort, which can also be understood as the driver selecting the "comfort mode", the weight of the first cost function is 0.3, the weight of the second cost function is 0.6, and the weight of the third cost function is 0.1; the preference is energy saving, which can also be understood as the driver selecting the "energy saving mode", the weight of the first cost function is 0.3, the weight of the second cost function is 0.5, and the weight of the third cost function is 0.2; the preference is sports, which can also be understood as the driver selecting the "sports mode", the weight of the first cost function is 0.2, the weight of the second cost function is 0.2, and the weight of the third cost function is 0.6.

[0116] In some embodiments, the preset control parameters are obtained by the following steps: obtaining multiple passenger physiological characteristics of the vehicle and a current feature weight data set, the multiple passenger physiological characteristics including heart rate variability, skin surface conductivity, three-dimensional head acceleration and facial expressions, the current feature data set including: current weights of multiple passenger physiological characteristics; for each of the passenger physiological characteristics, obtaining a first feature comfort score based on the passenger physiological characteristic, and obtaining a second feature comfort score based on the weight of the passenger physiological characteristic by performing the first feature comfort score; summing up the respective second feature comfort scores to obtain a target comfort score; based on a pre-established score control mapping relationship, obtaining the preset control parameters according to the target comfort score, the comfort score control mapping relationship including: multiple comfort scores and control parameters corresponding to each comfort score.

[0117] In some embodiments, the current feature weight dataset is obtained by the following steps: obtaining the current passenger information of the vehicle; based on the passenger weight mapping relationship, obtaining the current feature weight dataset according to the current passenger information, the passenger weight mapping relationship includes: multiple passenger information and a weight dataset corresponding to each passenger information, and the weight dataset includes: weights of multiple passenger physiological characteristics.

[0118] In some embodiments, the target comfort score is calculated by the following formula: comfort =k HRV ×HRV Score +k GSR ×GSR Score +k Motion ×Motion Score +

[0119] k Vision ×Vision Score , where k HRV is the weight of heart rate variability, HRV Score is the first characteristic comfort score of heart rate variability, k GSR is the weight of skin surface conductivity, GSR Score is the first characteristic comfort score of skin surface conductivity, k Motion is the weight of the head's three-dimensional acceleration, Motion Score is the first characteristic comfort score of the three-dimensional acceleration of the head, k Vision is the weight of facial expression, Vision Score Score the comfort of the first characteristic of the facial expression.

[0120] In some embodiments, the first characteristic comfort score of heart rate variability is calculated by the following formula: Among them, R d It is the time domain parameter of heart rate variability, namely the RMSSD parameter.

[0121] It should be noted that the RMSSD parameter represents the root mean square value of the difference between adjacent heartbeat intervals (RR intervals), and its calculation formula is as follows: Among them, RR i is the i-th RR interval, i.e., the heartbeat interval, and n is the total number of RR intervals. d The normal range of R is 20ms to 120ms, which is often used to reflect the activity of the passenger's parasympathetic nerve, namely the vagus nerve. d The higher the value, the greater the heart rate variability and the stronger the autonomic nervous system regulation ability, which usually indicates that the passenger is in a relaxed state. d Lower values indicate lower heart rate variability, which may indicate a passenger is stressed or fatigued. If R d = 120ms, the passenger performs at his best. When R d = 20ms, the passenger shows stress. The heart rate sensor can output real-time heart rate HR, interval RR and the pressure index used to calculate heart rate changes. The heart rate sensor is installed on the steering wheel of the driver's seat and the seat armrest of the passenger seat.

[0122] When the present application is actually used, if the vehicle is traveling on a bumpy road, R d HRV Score Monotonically increasing, indicating that the passenger is nervous, it is necessary to increase the suspension damping and limit the speed to suppress vibration transmission; if the vehicle is traveling on a smooth road, R d HRV tends to be stable Score Slight fluctuations within a certain range indicate that the passenger's emotions are stable, and this application maintains the default parameters.

[0123] In some embodiments, the first characteristic comfort score of skin surface conductivity is calculated by the following formula: SCR count is the number of peak values of skin surface conductivity per minute. SCR is converted to count Compress to [-1,1], and then map to [0,1] through linear transformation.

[0124] It should be noted that SCR count The change of skin surface conductivity reflects the emotional fluctuation or physiological arousal level. When the driver feels nervous, anxious or excited, the sympathetic nerve activity is enhanced, resulting in increased sweat gland secretion, increased skin conductance level and transient fluctuation (SCR). count Then it gets bigger. If SCR count =0, then Indicates that the driver or passenger is emotionally stable. If SCR count =5, GSR Score =1-0.76=0.24, indicating that the driver or passenger's emotions fluctuate. If SCR count >>5, then GSR Score →0, indicating that the driver or passenger is experiencing significant stress or discomfort. The skin conductance level (SCL) and instantaneous fluctuation (SCR) can be output by the galvanic skin response sensor, which is installed at the seatback contact point after the driver and passenger are seated.

[0125] When actually using this application, GSR ScoreWhen the speed is less than 0.4, the system will trigger active deceleration and limit the steering angle rate to avoid sudden turns or frequent lane changes. Score When it is greater than 0.8, the default parameters are maintained, focusing on traffic efficiency.

[0126] For example, when the vehicle is in an emergency obstacle avoidance scenario, the SCR count If the acceleration surges, the vehicle acceleration is immediately reduced so that the longitudinal acceleration is less than or equal to 1.5 km / s. 3 In addition, the seat support is enhanced to stabilize the occupant's posture. When the vehicle is in a long-term driving scenario, if the SCR count Continuously low, such as GSR Score ≈0.9, the vehicle speed is allowed to be increased to optimize traffic efficiency.

[0127] In some embodiments, the first characteristic comfort score of the three-dimensional head acceleration is calculated using the following formula: That is, the three-dimensional acceleration of the head is mapped to [0,1] through a negative exponential function, a head_RMS is the three-dimensional acceleration root mean square, and the calculation formula is as follows: a head_x is the acceleration of the head in the x direction, a head_y is the acceleration of the head in the y direction, a head_z is the acceleration of the head in the z direction, and the x, y, and z directions are perpendicular to each other.

[0128] It should be noted that a can be measured by a six-axis inertial measurement unit (IMU) installed inside the seat headrest head_x 、a head_y 、a head_z , Motion Score The smaller it is, the more severe the head shaking is and the less comfortable the driver or passenger is. head_RMS Significantly increased, making Motion Score Rapid descent, where the driver or passengers experience discomfort caused by bumps in the road, such as a head_RMS >0.4g, e -0.5×0.4 ≈0.82, indicating that the driver or passenger feels uncomfortable when the vehicle passes through bumpy roads or makes an emergency lane change; if a head_RMS The value tends to be smaller, Motion Score Maintain stability, that is, the driver or passenger performs at their best, such as a head_RMS =0.1g, Motion Score ≈0.95, indicating that the driver or passenger of the vehicle performed optimally.

[0129] For example, the vehicle passes through a speed bump, a head_RMSInstantly increased to 0.5g, Motion Score =e -0.5×0.5 ≈0.78, it is necessary to improve the suspension stiffness and reduce vibration transmission; when the vehicle is cruising smoothly on the highway (HWA / ICA / LCC) or piloting (NCA / NOA), the shaking of the driver or passenger's head is very small, such as a head_RMS =0.05g, Motion Score =e -0.5×0.05 ≈0.98, then maintain the default parameters.

[0130] In some embodiments, the first characteristic comfort score of the facial expression is calculated by the following formula:

[0131] AU4 intensity is the probability of frowning, is the number of blinks, that is, the first characteristic comfort score of facial expression is mapped to [0,1]. The lower the value, the stronger the discomfort detected by vision.

[0132] It should be noted that the frowning probability is used to characterize the intensity of corrugator muscle activity in facial action coding, reflecting the frequency and amplitude of the driver or passenger's frowning. The higher the value, the stronger the tension or discomfort. The number of blinks can be the number of blinks per minute when the driver or passenger frowns. Through the statistics of image tracking algorithm, abnormal blinking frequency, such as too high or too low, may indicate that the driver or passenger is tired, distracted or anxious, thereby reflecting the level of facial discomfort of the driver or passenger. 30 is the empirical threshold, which can be used to calibrate the quantity. That is, the driver or passenger does not frown and blinks at a normal frequency, then Vision Score =1-0.5×(0+0.5)=0.75; if If the driver or passenger frowns and blinks frequently, Vision Score =1-0.5×(0.8+1)=0.1.

[0133] When actually using this application, Vision Score <0.3, triggers active seat support, enhances air conditioning, and limits the vehicle's longitudinal acceleration to less than or equal to 1.5 km / s 3 ;Vision Score >0.7, maintain default parameters, focusing on traffic efficiency first.

[0134] For example, when the vehicle is in a long-distance driving scenario, the driver or passenger fatigue is monitored through the cockpit DMS system (driver fatigue monitoring system) or OMS system (occupant monitoring system). Alternatively, the vehicle's onboard camera can be used for detection. The onboard camera is installed on the vehicle's A-pillar, driver's rearview mirror, and the base of the middle and rear row ceiling screen. By observing the driver's and passenger's facial micro-expressions and eye movement tracking, facial action codes are output. Determine if the driver or passenger is tired, remind the driver to take a rest, and automatically turn on the reclining mode of the passenger seat and remind the passenger to take a rest; AU4 appears when the vehicle has emergency braking (AEB), automatic emergency steering (ESA), or sudden emergency braking and steering. intensity A sharp increase in a short period of time will immediately reduce the acceleration limit of the following vehicles, reducing the impact and pressure felt by the driver or passengers.

[0135] In this application, pressure distribution sensors can also be installed on the driver's and passenger seat cushions and backrests to detect the pressure center offset COP and the sitting stability index, and the steering wheel hands-off protection system HOD and the vehicle seat occupancy system can be used for detection.

[0136] In some embodiments, the comfort score control mapping relationship is as follows: if the comfort score is greater than 0 and less than or equal to a first value, the control parameters are: reducing the suspension damping coefficient by a first percentage, the longitudinal jerk is less than or equal to the first jerk, the steering angular rate is less than or equal to the first angular rate, and increasing the air conditioning air supply intensity by a first gear change; if the comfort score is greater than the first value and less than or equal to a second value, the control parameters are: maintaining the suspension damping coefficient unchanged, the longitudinal jerk is less than or equal to the second jerk, the steering angular rate is less than or equal to the second angular rate, and maintaining the air conditioning air supply intensity unchanged, with the second jerk being greater than the first jerk and the second angular rate being greater than the first angular rate; if the comfort score is greater than the second value, the control parameters are: increasing the suspension damping coefficient by a second percentage, the longitudinal jerk is less than or equal to a third jerk, the steering angular rate is less than or equal to a third angular rate, and reducing the air conditioning air supply intensity by a second gear change, with the third jerk being greater than the second jerk and the third angular rate being greater than the second angular rate. Maintaining the suspension damping coefficient unchanged can be replaced by maintaining the suspension damping coefficient at the baseline value.

[0137] For example, the comfort score control mapping relationship is: if the comfort score is greater than 0 and less than or equal to 0.3, the control parameters are to reduce the suspension damping coefficient by 30%, and the longitudinal acceleration is less than or equal to 1.5m / s 3, steering angle rate is less than or equal to ≤8° / s, air conditioning air supply intensity is increased by 2 levels; if the comfort score is greater than 0.3 and less than or equal to 0.6, the control parameters are to keep the suspension damping coefficient unchanged and the longitudinal acceleration less than or equal to 2.0m / s 3 , steering angle rate is less than or equal to 12° / s, and the air conditioning air supply intensity remains unchanged; if the comfort score is greater than 0.6, the control parameters are to increase the suspension damping coefficient by 20%, and the longitudinal acceleration is less than or equal to 3.0m / s 3 , the steering angle rate is less than or equal to 15° / s, and the air conditioning air supply intensity is reduced by 1 level.

[0138] In some embodiments, the reinforcement learning model also includes a multi-objective reward function, and the method also includes: using the multi-objective reward function to replace the multi-objective cost function, the multi-objective reward function is the difference between the weighted sum of the first reward function, the second reward function, and the third reward function and the penalty function, wherein the first reward function reflects the safety of the vehicle through the collision risk and lane keeping condition of the vehicle, the second reward function characterizes the comfort of the vehicle through the lateral comfort, longitudinal comfort and road bumps of the vehicle, and the third reward function characterizes the traffic efficiency of the vehicle through the speed maintenance and energy efficiency of the vehicle.

[0139] In some embodiments, the calculation formula of the multi-objective reward function is as follows: t =λ1R safety +λ2R comfort +λ3R efficiency -R penalty , R t is the multi-objective reward function at time t, R safety is the first reward function, λ1 is the weight of the first reward function, R comfort is the second reward function, λ2 is the weight of the first reward function, R efficiency is the third reward function, λ3 is the weight of the third reward function, R penalty is the penalty function.

[0140] In some embodiments, the method further includes: updating the multi-objective reward function according to the usage scenario and / or user preference data.

[0141] It should be noted that, with regard to the update of the multi-objective reward function, the weight update method of the first reward function can be the same as the weight update method of the first cost function, the weight update method of the second reward function can be the same as the weight update method of the second cost function, and the weight update method of the third reward function can be the same as the weight update method of the third cost function.

[0142] In some embodiments, the first reward function is the sum of a collision risk penalty function and a lane keeping reward function. The collision risk penalty function is determined by the distance between the vehicle and the vehicle in front or the time distance between the vehicle and the vehicle in front, and is used to characterize the collision risk of the vehicle. The lane keeping reward function is determined by the angle between the front of the vehicle and the center line of the lane line or the angle between the front of the vehicle and the center line of the passable area planned by the system when there is no lane line, and is used to characterize the lane keeping situation.

[0143] In some embodiments, the calculation formula of the first reward function is as follows: safety =R collision +R lane , R safety is the first reward function, R collision is the collision risk penalty function, R lane is the lane keeping reward function, d front The distance between the vehicle and the preceding vehicle or the time between the vehicle and the preceding vehicle; It is the angle between the front of the vehicle and the center line of the lane line, or the angle between the front of the vehicle and the center line of the passable area planned by the system when there is no lane line.

[0144] In some embodiments, the second reward function is the sum of a lateral comfort reward function, a longitudinal comfort reward function and a road bump and vibration suppression function, wherein the lateral comfort reward function is determined by the vehicle lateral acceleration and the vehicle roll angle, and is used to characterize the lateral comfort of the vehicle; the longitudinal comfort reward function is determined by the vehicle longitudinal acceleration and the vehicle longitudinal acceleration, and is used to characterize the longitudinal comfort of the vehicle; the road bump and vibration suppression function is determined by the vehicle vibration comfort score, and is used to characterize the road bump condition.

[0145] In some embodiments, the calculation formula of the second reward function is as follows: comfort =R lat +R long +R vib , R comfort is the second reward function, R lat is the lateral comfort reward function, R long is the longitudinal comfort reward function, R vib To suppress the road bump vibration function,

[0146] R lat =-β1·|α y |-β2|θ roll |, β1 is the weight of the vehicle's lateral acceleration, α y is the vehicle longitudinal acceleration, R long =-α1·|jerk|-α2·|ax |, α1 is the weight of the vehicle's longitudinal acceleration, jerk is the vehicle's longitudinal acceleration, α2 is the weight of the vehicle's longitudinal acceleration, a x is the vehicle longitudinal acceleration, R vib =-γ1·C vib , γ1 is the weight of vehicle vibration comfort score, C vib Rate the vehicle's vibration comfort.

[0147] In some embodiments, the method further includes: updating the weight of the vehicle longitudinal acceleration according to the amplitude of the driver's or passenger's head shaking.

[0148] In some implementations, the weight of the vehicle longitudinal acceleration is updated using the following formula:

[0149] is the weight of the vehicle longitudinal acceleration after update, α1 is the weight of the vehicle longitudinal acceleration before update, H shake is the head shaking amplitude.

[0150] In this case, if the driver or passenger's head shakes significantly, indicating that the driver or passenger is feeling uncomfortable, such as nodding frequently, the vehicle's longitudinal acceleration tolerance threshold is lowered.

[0151] In some embodiments, the third reward function is the sum of a vehicle speed maintenance function and an energy efficiency function. The vehicle speed maintenance function is determined by the vehicle speed and the expected vehicle speed and is used to characterize the vehicle speed maintenance situation. The energy efficiency function is determined by the brake cylinder pressure and the accelerator pedal opening and is used to characterize the energy efficiency.

[0152] In some embodiments, the calculation formula of the third reward function is as follows: efficiency =R speed +R energy , R speed is the vehicle speed maintenance function, R speed =η·

[0153] (vv desired ) 2 ,η is the weight of vehicle speed maintenance, v is vehicle speed, v desired is the expected speed,

[0154] R energy =-μ·(|P brake |+|T throttle |), μ is the weight of energy efficiency, P brake is the brake cylinder pressure, T throttle is the accelerator pedal opening.

[0155] In some embodiments, the penalty term function is the sum of a frequent lane change penalty function and a control quantity mutation penalty function, wherein the frequent lane change penalty function is determined by discrete lane change decisions, and the control quantity mutation penalty function is determined by continuous values of steering angle increments and continuous values of brake cylinder pressure increments.

[0156] In some embodiments, the penalty function is calculated as follows:

[0157] R penalty =R lane_change +R smooth , R lane_change is the frequent lane change penalty function, m lane (k) is the discrete lane change decision at time k, R smooth is the control quantity mutation penalty function, R smooth =-0.3·(|Δδ steer 2 |+|Δp brake 2 |), Δδ steer is the continuous value of the rotation angle increment, Δp brake It is the continuous value of brake cylinder pressure increment.

[0158] In some embodiments, the method further includes: if the passenger of the vehicle is identified as a target passenger, obtaining a passenger reward function of the target passenger, wherein the passenger reward function is used to characterize the comfort change of the target passenger and the energy consumption change of the vehicle; and using the passenger reward function to replace the multi-objective reward function.

[0159] In some embodiments, the passenger reward function is calculated as follows: c =

[0160] ΔC comfort -k energy ×ΔE energy =(C comfort (t2)-C comfort (t1))-k energy ×ΔE energy , ΔC comfort is the change in comfort score at time t2 relative to time t1, C comfort (t2) is the target comfort score at time t2, C comfort (t1) is the target comfort score at time t1, k energy To balance the weight of comfort and energy consumption, ΔE energy k is the change in energy consumption from time t1 to time t2. energy It can be 0.1, k energyLess than 1, by reducing the weight of energy consumption changes, priority is given to ensuring comfort optimization, avoiding energy saving at the expense of passenger experience, such as excessive speed reduction leading to extended travel time. Time t2 can be understood as the current moment, and time t1 can be understood as the previous moment. ΔC comfort A positive value indicates improved comfort, such as reduced bumps and suppressed lateral swings. ΔC comfort A negative value indicates a decrease in comfort, such as discomfort caused by sudden acceleration and sharp steering, which can be used to optimize the control strategy of the drive system to improve passenger comfort; ΔE energy A positive value indicates increased energy consumption, such as high motor power output and frequent braking, ΔE energy A negative value indicates reduced energy consumption, such as improved energy recovery efficiency and smooth driving, which is used to avoid excessive energy consumption and improve driving economy.

[0161] In some embodiments, controlling the vehicle to perform automatic driving according to the target control parameters includes: inputting the target control parameters into a pre-established chassis control model to obtain chassis control parameters, the chassis control model is established based on a dynamic matrix control algorithm, and the chassis control parameters include longitudinal acceleration, lateral acceleration, vertical acceleration, vehicle body roll angle and vehicle body pitch angle; according to the chassis control parameters, controlling the vehicle's suspension system, braking system and steering system.

[0162] In some embodiments, the input vectors of the chassis control model include suspension damping force, front wheel angle and brake cylinder pressure, and the output vectors include vehicle longitudinal acceleration, vehicle lateral acceleration, vehicle vertical acceleration, vehicle roll angle and vehicle pitch angle, that is, the input vector is

[0163] U=[F damp ,δ steer ,P brake ] T , F damp is the suspension damping force, continuously adjustable, ranging from 0 to 5000N; δ steer is the front wheel angle, with a range of 0 to 30°; P brake is the brake cylinder pressure, ranging from 0 to 20000 bar, and the output vector is Y=[a x ,a y ,a z ,θ roll ,θ pitch ] T , a x is the vehicle longitudinal acceleration, reflecting the vehicle acceleration and braking status, a y is the lateral acceleration of the vehicle, reflecting the steering force or lateral yaw state of the vehicle, a z is the vertical acceleration of the vehicle, reflecting the bump and pitch state of the vehicle, θ rollis the body roll angle, reflecting the rotation angle of the body around the longitudinal axis, θ pitch It is the pitch angle of the vehicle body, reflecting the rotation angle of the vehicle body around the horizontal axis.

[0164] In this application, on the one hand, the output vector provides real-time dynamic response data of the vehicle as a feedback signal to ensure that the actual state of the vehicle is consistent with the desired target (comfort, safety). roll >5°, the vehicle rolls excessively, increase the suspension damping or limit the steering angular velocity, or in other words, when the vehicle brakes suddenly or in an emergency, a x Sudden changes trigger slow braking or power smoothing strategies to suppress the increase in vehicle longitudinal acceleration. On the other hand, the output vector is related to the comfort of the driver or passenger and is used to control the roll, pitch and acceleration balance of the vehicle. roll When it is too large, such as in a sharp turn, the driver or passengers may feel lateral pressure. pitch When a sudden change occurs, it will cause the driver or passenger to have a "nodding" effect. These angle fluctuations can be suppressed by adjusting the suspension damping and brake pressure. For example, when the scene is a congested following scene, a x and a y The rate of change of represents the comfort of the driver or passenger. It can limit the longitudinal acceleration of the vehicle to a low level, thereby ensuring that the acceleration, deceleration and lateral swing of the vehicle in congested following will not cause discomfort to the passengers in the cabin. On the other hand, if the value of the output vector exceeds the safety threshold, such as a y >0.3g or θ roll >8°, trigger warning or emergency intervention control strategy.

[0165] It should be noted that if the sensor data of the vehicle is abnormal, such as the failure of the IMU (inertial measurement unit), it switches to the backup mode and uses other sensors such as cameras and lidar to indirectly estimate the state of the output vector Y.

[0166] For example, taking a curve scene as an example, the control target is to reduce the vehicle roll angle and the vehicle lateral acceleration, which is used to enhance F damp and limit δ steer , to reduce the lateral shaking of the driver or passengers during autonomous driving and improve the comfort experience; taking the bumpy road scene as an example, the control goal is to suppress the decrease of the vehicle's roll angle and the decrease of the vehicle's vertical acceleration, which is used for the active suspension to adjust the damping in advance, thereby achieving head acceleration a head_RMS Reduce; Taking the automatic emergency braking scenario as an example, the control goal is to suppress the decrease of the vehicle's longitudinal acceleration, which is used for staged slow braking and control of ΔP brake Gradually increase to achieve vehicle speed control jerk<1.5m / s 3 , ensure that the frequency of passengers nodding is reduced.

[0167] In some embodiments, the chassis control model further includes a chassis cost function for minimizing the deviation between the output vector and the expected output vector. The formula of the chassis cost function is as follows: Y is the output vector, Y ref The expected output vector, ΔU is the input vector change, U is the input vector, Q, R and S are weight matrices, which represent the balanced output variable deviation, input vector change error and input vector amplitude respectively. The output variable deviation can be understood as the tracking error, the input vector change error can be understood as the control increment error, and the input vector amplitude can be understood as the control increment energy consumption.

[0168] In some embodiments, the constraints for solving the chassis cost function are as follows: min vU≤U max 、|a y |≤0.3g, |θ roll |v5°、|ΔU(t)|≤ΔU max .U min is the minimum physical limit, U max is the maximum physical limit, a y is the vehicle lateral acceleration, θ roll is the vehicle roll angle, ΔU(t) is the input vector change at time t, ΔU max Enter the upper limit of the vector change allowed for a single step.

[0169] In some embodiments, the input vector change is determined by the following formula: is the input change of the k+1 iteration of the i-th subsystem, J i is the local cost function of the i-th subsystem, such as suspension comfort optimization, λ is the Lagrange multiplier, A i is the dynamic matrix of the ith subsystem, ΔU i is the input variation of the ith subsystem, Y coup The coupling effect of other subsystems on the current output, such as the roll caused by steering action.

[0170] In this way, the input vector is coupled and constrained, and the mathematical and physical constraints of the chassis control instructions of the suspension system, braking system and steering system are coordinated to ensure that multiple subsystems work together during the optimization process and avoid conflicts caused by local optimization.

[0171] In some embodiments, the output vector of the chassis control model is determined as follows: Y = Y0 + A·ΔU, where Y0 is the output vector when the input vector remains unchanged (ΔU = 0), ΔU is the change in the input vector, and A is a dynamic matrix used to characterize the linear influence of ΔU on Y.

[0172] In some embodiments, the dynamic matrix is formulated as follows:

[0173] Among them, G damp is the step response coefficient matrix of the suspension damping force, G steer is the step response coefficient matrix of the front wheel steering angle, G brake is the step response coefficient matrix of the brake cylinder pressure, C roll is the coupling coefficient of suspension damping force to vehicle roll acceleration, C pitch is the coupling coefficient of the suspension damping force change to the pitch acceleration, C lat is the coupling coefficient of the front wheel steering angle change to the vehicle lateral acceleration, C long is the coupling coefficient of the brake cylinder pressure change to the vehicle longitudinal acceleration.

[0174] In some embodiments, the step response coefficient matrix of the suspension damping force is as follows:

[0175] g damp (N) is the step response coefficient of the Nth suspension damping force.

[0176] In some embodiments, the step response coefficient of the suspension damping force is determined by the following formula: a z (t) is the suspension damping force transmission vehicle vertical acceleration function at time t, g damp (i) is the step response coefficient of the i-th suspension damping force, ΔF damp (ti+1) is the increase in the suspension damping force at time t-i+1.

[0177] In some implementations, the step response coefficient matrix of the front wheel steering angle is as follows:

[0178] g steer (N) is the step response coefficient of the Nth front wheel steering angle.

[0179] In some embodiments, the step response coefficient of the front wheel steering angle is determined by the following formula: a y (t) is the lateral acceleration function transmitted by the front wheel angle at time t, g steer (j) is the step response coefficient of the j-th front wheel angle, Δδ steer(t-j+1) is the front wheel steering angle increment at time t-j+1.

[0180] In some implementations, the step response coefficient matrix of the brake cylinder pressure is as follows:

[0181] g brake (N) is the step response coefficient of the Nth brake cylinder pressure.

[0182] In some implementations, the step response coefficient of the brake cylinder pressure is determined by the following formula:

[0183] a x (t) is the longitudinal acceleration function transmitted by the brake cylinder pressure at time t, g brake (k) is the step response coefficient of the kth brake cylinder pressure, ΔP brake (t-k+1) is the brake cylinder pressure increment at time t-k+1.

[0184] In some embodiments, the input vector change is:

[0185] ΔU=[Δu(t),Δu(t+1),…,Δu(t+H c -1)] T , Δu(t) is the input change at time t.

[0186] In some embodiments, the suspension damping force is determined by the following formula: To predict vibration acceleration, α is the feedforward gain coefficient.

[0187] In this way, it has predictive control, adaptive capabilities, hybrid control architecture and engineering practicality. It can predict vehicle vibration in advance (for example, 50ms) and actively adjust the damping coefficient, that is, predict road excitation in advance, actively offset vibration energy input, adapt to the vehicle in different load changes and suspension aging, optimize control resource allocation through frequency domain segmentation, and take into account both low-frequency comfort and high-frequency stability.

[0188] In some embodiments, the method further includes obtaining operating data of the vehicle driver in comfort mode, and training a generative adversarial network model based on the operating data to ensure that the decision output approximates the smoothness of human driving. The adversarial network model comprises a generator and a discriminator. The adversarial network model may be a GAN model. When the decision output approximates the smoothness of human driving, humanized driving is achieved.

[0189] In some embodiments, the method further comprises: updating the control strategy online by Q-learning, wherein the formula is: Q(s,a)←Q(s,a)+α[R+γmax′α Q(s′, á)-Q(s, a)], Q(s, a) is the current Q value, s represents the current state, such as comfort score, driving mode, road type, a represents the action performed, such as adjusting suspension damping, adjusting brake cylinder pressure, α is the learning rate, R is the immediate reward, R = ΔC comfort -0.1·ΔE break ,γ is the discount factor, max′ α Q(s′,á) is the maximum Q value of the next state.

[0190] It should be noted that Q-learning is a reinforcement learning algorithm based on a value function. It aims to learn the strategy of selecting the optimal action in a specific state (State) by interacting with the environment through trial and error, so as to maximize the long-term cumulative reward.

[0191] In some embodiments, the state space also includes comfort score change, usage scenario, lateral deviation, heart rate variability, and skin surface conductivity, and the action space also includes driving mode switching. Usage scenario can be understood as a driving scenario, including congested, high-speed, and bumpy driving scenarios, and the driving mode switching can be switching from comfort mode to energy-saving mode.

[0192] In this way, optimizing the passenger reward function can improve comfort and suppress ineffective energy consumption.

[0193] In some embodiments, the method further includes: if the passenger of the vehicle is identified as a target passenger, obtaining historical data of the target passenger, initializing a Q value table based on the historical data, or adjusting a learning rate to accelerate convergence of the control strategy.

[0194] In this way, the historical Q-value table can be directly loaded based on the passenger's historical data, such as motion sickness sensitivity or preferred damping mode, to avoid repeated learning. For example, the high damping strategy is enabled by default for elderly passengers.

[0195] In some embodiments, the NSGA-II algorithm is used to search for the Pareto optimal solution of the multi-objective cost function, and the Q-learning algorithm is used to explore the balance point of comfort, safety, and energy consumption, wherein the Q-learning algorithm predicts γmax′ through long-term reward. α Q(s′,á) selects the globally optimal policy.

[0196] For example, if the usage scenario is a bumpy scenario, high suspension damping action is preferred, sacrificing some energy consumption to improve comfort; if the usage scenario is a smooth cruising scenario, moderate acceleration is allowed to optimize energy efficiency; if the usage scenario is a target conflict driving scenario such as an emergency obstacle avoidance scenario, comfort needs to be sacrificed.

[0197] Taking dynamic adaptability improvement as an example, when the usage scenario is complex, such as rainy and snowy weather or sudden obstacles, the Q-learning algorithm adjusts the control parameters in real time to avoid the rigidity of rule-based control. When it detects that the passenger nods frequently or the head shakes more severely, it immediately reduces the acceleration limit to suppress discomfort. Taking long-term strategy learning as an example, through the cumulative reward mechanism, it learns the optimal strategy in different driving cycles, such as morning and evening rush hours, long-distance travel, etc. When the usage scenario is congested, it gradually learns to slow down the brakes in advance to reduce sudden stops, rather than relying on a fixed following distance. Taking hardware compatibility assurance as an example, in response to the computing power limitations of the on-board embedded ECU, the Q-learning algorithm ensures that the single iteration time is less than 5ms through sparse matrix optimization and lightweight state coding, such as discretized comfort scoring, to meet the real-time requirements of autonomous driving for vehicles.

[0198] In this application, on the one hand, a reinforcement learning model is used to achieve real-time trade-offs between comfort and safety in complex scenarios, which is superior to traditional rules or MPC methods and improves dynamic adaptability; on the other hand, personalized adaptation can be performed when planning and controlling vehicle driving based on the real-time attribute characteristics of the vehicle's driver and passengers, and a customized driving style can be provided by combining user preferences with biofeedback; on the other hand, multi-system collaborative optimization control and deep coupling of the reinforcement learning model with the chassis control model achieve "planning-execution" integrated optimization, avoiding poor driving experience caused by data transmission and fragmentation.

[0199] Figure 2 A frame diagram of an automatic driving control device in an embodiment of the present application is shown. Figure 2 According to a second aspect of the present application, an automatic driving control device 100 is provided, the device comprising:

[0200] A first acquisition unit 101 acquires vehicle usage scenario data and / or user preference data;

[0201] a first updating unit 102 for updating a multi-objective cost function of a pre-established reinforcement learning model based on the usage scenario and / or user preference data, wherein the reinforcement learning model is a framework for interaction between the vehicle and the environment, and its model structure is described by a multi-tuple, the multi-tuple including at least the multi-objective cost function, wherein the multi-objective cost function is used to balance the safety, comfort, and economy of the vehicle during autonomous driving;

[0202] a second acquisition unit 103 for acquiring environmental parameters and preset control parameters of the vehicle, and inputting the environmental parameters and the preset control parameters into the updated reinforcement learning model to output target control parameters, wherein the preset control parameters include a suspension damping coefficient, longitudinal jerk, and steering angular rate, and the target control parameters include a suspension damping force, a front wheel steering angle, and a brake cylinder pressure;

[0203] The first control unit 104 controls the vehicle to perform automatic driving according to the target control parameters.

[0204] In some embodiments, the multi-objective cost function is obtained by weighted summation of a first cost function, a second cost function, and a third cost function, the first cost function is used to characterize the safety of the vehicle, the second cost function is used to characterize the comfort of the vehicle, and the third cost function is used to characterize the economy of the vehicle. The first updating unit is configured as: a first obtaining unit, based on a scenario adjustment mapping relationship, obtains a weight adjustment set according to the usage scenario data, and / or based on a preference weight mapping relationship, obtains a target weight set according to the user preference data, wherein the scenario adjustment mapping relationship includes: multiple scenarios and an adjustment set corresponding to each scenario, the adjustment set includes adjusting the weights of the first cost function, the second cost function, and the third cost function, the preference weight mapping relationship includes: multiple preferences and a weight set corresponding to each preference, the weight set includes the weights of the first cost function, the second cost function, and the third cost function; a second updating unit, based on the weight adjustment set, adjusts the weights of the first cost function, the second cost function, and the third cost function, and / or updates the weights of the first cost function, the second cost function, and the third cost function to the target weight set.

[0205] In some embodiments, the device further includes a third acquisition unit, which is used to acquire the preset control parameters, and the third acquisition unit is configured as: a fourth acquisition unit, which acquires multiple passenger physiological characteristics of the vehicle and a current feature weight data set, wherein the multiple passenger physiological characteristics include heart rate variability, skin surface conductivity, three-dimensional head acceleration, and facial expressions, and the current feature data set includes: current weights of multiple passenger physiological characteristics; a second acquisition unit, which obtains a first feature comfort score for each passenger physiological characteristic according to the passenger physiological characteristic, and obtains a second feature comfort score by performing the first feature comfort score based on the weight of the passenger physiological characteristic; a third acquisition unit, which sums each of the second feature comfort scores to obtain a target comfort score; a fourth acquisition unit, which obtains the preset control parameters according to the target comfort score based on a pre-established score control mapping relationship, and the comfort score control mapping relationship includes: multiple comfort scores and control parameters corresponding to each comfort score.

[0206] In some embodiments, the device also includes a fifth obtaining unit, which is used to obtain the current feature weight data set, and the fifth obtaining unit is configured as: a fifth obtaining unit, which obtains the current passenger information of the vehicle; a sixth obtaining unit, which obtains the current feature weight data set based on the current passenger information based on the passenger weight mapping relationship, and the passenger weight mapping relationship includes: multiple passenger information and a weight data set corresponding to each passenger information, and the weight data set includes: weights of multiple passenger physiological characteristics.

[0207] In some embodiments, the reinforcement learning model also includes a multi-objective reward function, and the device also includes: a first replacement unit, which uses the multi-objective reward function to replace the multi-objective cost function, and the multi-objective reward function is the difference between the weighted sum of the first reward function, the second reward function, and the third reward function and the penalty function, wherein the first reward function reflects the safety of the vehicle through the collision risk and lane keeping condition of the vehicle, the second reward function characterizes the comfort of the vehicle through the lateral comfort, longitudinal comfort and road bumps of the vehicle, and the third reward function characterizes the traffic efficiency of the vehicle through the speed maintenance and energy efficiency of the vehicle.

[0208] In some embodiments, the device further includes: a sixth acquisition unit, which acquires a passenger reward function of the target passenger if the passenger of the vehicle is identified as a target passenger, wherein the passenger reward function is used to characterize the comfort change of the target passenger and the energy consumption change of the vehicle; and a second replacement unit, which replaces the multi-objective reward function with the passenger reward function.

[0209] In some embodiments, the first control unit is configured as: a seventh obtaining unit, inputting the target control parameters into a pre-established chassis control model to obtain chassis control parameters, wherein the chassis control model is established based on a dynamic matrix control algorithm, and the chassis control parameters include longitudinal acceleration, lateral acceleration, vertical acceleration, vehicle body roll angle and vehicle body pitch angle; a second control unit, controlling the suspension system, braking system and steering system of the vehicle according to the chassis control parameters.

[0210] Based on the same inventive concept, as a third aspect, the present application also provides a computer-readable storage medium on which a computer program is stored. The computer program includes executable instructions. When the executable instructions are executed by a processor, the method described in any embodiment of the first aspect of the present application is implemented.

[0211] In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary implementations of the present application described in the above "Exemplary Method" section of this specification.

[0212] refer to Figure 3 As shown, a program product 200 for implementing the above method according to an embodiment of the present application is described. The program product 200 may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0213] As another aspect, the present application also provides an electronic device capable of implementing the above method.

[0214] Refer to the following Figure 4 hereinafter, an electronic device 300 according to this embodiment of the present application is described. Figure 4 The electronic device 300 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0215] like Figure 4 As shown, the electronic device 300 is a general-purpose computing device, including: a processing unit 310: used to execute program code; a storage unit 320: storing program code, including a volatile storage unit (such as RAM 321, a cache storage unit 322) and a read-only storage unit (ROM 323), and a program / utility 324 containing a program module 325; a bus 330: connecting different system components and supporting multiple bus structures; an I / O interface 350: used to communicate with external devices (such as a keyboard, a Bluetooth device, etc.); a network adapter 360: connected through the bus 330, supporting communication with a local area network, a wide area network and a public network.

Claims

1. An automatic driving control method, characterized in that: The method comprises: Obtaining vehicle usage scenario data and / or user preference data; updating a multi-objective cost function of a pre-established reinforcement learning model based on the usage scenario and / or user preference data, wherein the reinforcement learning model is a framework for interaction between the vehicle and the environment, and its model structure is described by a multi-tuple, the multi-tuple including at least the multi-objective cost function, wherein the multi-objective cost function is used to balance the safety, comfort, and economy of the vehicle during autonomous driving; Obtaining environmental parameters and preset control parameters of the vehicle, and inputting the environmental parameters and the preset control parameters into the updated reinforcement learning model to output target control parameters, wherein the preset control parameters include a suspension damping coefficient, longitudinal jerk, and steering angular rate, and the target control parameters include a suspension damping force, a front wheel steering angle, and a brake cylinder pressure; The vehicle is controlled to perform automatic driving according to the target control parameters.

2. The automatic driving control method according to claim 1, characterized in that: The multi-objective cost function is obtained by weighted summation of a first cost function, a second cost function, and a third cost function, wherein the first cost function is used to characterize the safety of the vehicle, the second cost function is used to characterize the comfort of the vehicle, and the third cost function is used to characterize the economy of the vehicle. The multi-objective cost function of the pre-established reinforcement learning model is updated according to the usage scenario and / or user preference data, including: Based on the scenario adjustment mapping relationship, a weight adjustment set is obtained according to the usage scenario data, and / or based on the preference weight mapping relationship, a target weight set is obtained according to the user preference data, wherein the scenario adjustment mapping relationship includes: multiple scenarios and an adjustment set corresponding to each scenario, the adjustment set includes adjusting weights of the first cost function, the second cost function, and the third cost function, and the preference weight mapping relationship includes: multiple preferences and a weight set corresponding to each preference, the weight set includes weights of the first cost function, the second cost function, and the third cost function; The weights of the first cost function, the second cost function and the third cost function are adjusted according to the weight adjustment set, and / or the weights of the first cost function, the second cost function and the third cost function are updated to the target weight set.

3. The automatic driving control method according to claim 1, wherein: The preset control parameters are obtained by the following steps: Acquiring multiple physiological characteristics of passengers in the vehicle and a current characteristic weight dataset, wherein the multiple physiological characteristics of passengers include heart rate variability, skin surface conductivity, three-dimensional head acceleration, and facial expression, and the current characteristic dataset includes current weights of the multiple physiological characteristics of passengers; For each of the passenger physiological characteristics, obtain a first characteristic comfort score according to the passenger physiological characteristic, and obtain a second characteristic comfort score by performing the calculation on the first characteristic comfort score based on the weight of the passenger physiological characteristic; Summing each of the second characteristic comfort scores to obtain a target comfort score; Based on a pre-established score control mapping relationship, the preset control parameter is obtained according to the target comfort score, and the comfort score control mapping relationship includes: multiple comfort scores and control parameters corresponding to each comfort score.

4. The automatic driving control method according to claim 1, wherein: The current feature weight dataset is obtained by the following steps: Obtaining current passenger information of the vehicle; Based on the passenger weight mapping relationship, a current feature weight data set is obtained according to the current passenger information. The passenger weight mapping relationship includes: multiple passenger information and a weight data set corresponding to each passenger information. The weight data set includes: weights of multiple passenger physiological characteristics.

5. The automatic driving control method according to claim 1, characterized in that: The reinforcement learning model also includes a multi-objective reward function, and the method further includes: The multi-objective reward function is used to replace the multi-objective cost function. The multi-objective reward function is the difference between the weighted sum of the first reward function, the second reward function, and the third reward function and the penalty function, wherein the first reward function reflects the safety of the vehicle through the collision risk and lane keeping condition of the vehicle, the second reward function characterizes the comfort of the vehicle through the lateral comfort, longitudinal comfort and road bumps of the vehicle, and the third reward function characterizes the traffic efficiency of the vehicle through the speed maintenance and energy efficiency of the vehicle.

6. The automatic driving control method according to claim 5, characterized in that: The method further comprises: If the passenger of the vehicle is identified as a target passenger, obtaining a passenger reward function of the target passenger, wherein the passenger reward function is used to characterize a change in comfort of the target passenger and a change in energy consumption of the vehicle; The passenger reward function is used to replace the multi-objective reward function.

7. The automatic driving control method according to claim 1, characterized in that: The controlling the vehicle to perform automatic driving according to the target control parameter includes: Inputting the target control parameters into a pre-established chassis control model to obtain chassis control parameters, wherein the chassis control model is established based on a dynamic matrix control algorithm, and the chassis control parameters include longitudinal acceleration, lateral acceleration, vertical acceleration, vehicle body roll angle, and vehicle body pitch angle; The suspension system, braking system and steering system of the vehicle are controlled according to the chassis control parameters.

8. An automatic driving control device, characterized in that: The device comprises: A first acquiring unit, acquiring vehicle usage scenario data and / or user preference data; a first updating unit configured to update a multi-objective cost function of a pre-established reinforcement learning model based on the usage scenario and / or user preference data, wherein the reinforcement learning model is a framework for interaction between the vehicle and the environment, and its model structure is described by a multi-tuple, the multi-tuple including at least the multi-objective cost function, wherein the multi-objective cost function is used to balance safety, comfort, and economy of the vehicle during autonomous driving; a second acquisition unit, configured to acquire environmental parameters and preset control parameters of the vehicle, and input the environmental parameters and the preset control parameters into the updated reinforcement learning model to output target control parameters, wherein the preset control parameters include a suspension damping coefficient, a longitudinal jerk, and a steering angular rate, and the target control parameters include a suspension damping force, a front wheel steering angle, and a brake cylinder pressure; A first control unit controls the vehicle to perform automatic driving according to the target control parameters.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program includes executable instructions, and when the executable instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: one or more processors; A memory for storing executable instructions of the processor, wherein when the executable instructions are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.