Vehicle control method, system, device and equipment based on gestures and medium

By combining driver gestures and vehicle working conditions data, and using deep recursive Q network to generate target vehicle control instructions, the existing vehicle control technology is solved inadequate convenience, safety and responsiveness, and precise vehicle control in harsh environments is achieved.

CN120440060APending Publication Date: 2025-08-08CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510768848.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing vehicle control technology has shortcomings in terms of convenience, safety and accurate response to drivers' intentions to control vehicles. Traditional manual control relies on driver skills. Mobile smart car control is limited by network signals, traditional gesture control signals are blurred, and autonomous driving performance is limited in harsh environments.

Method used

By obtaining the driver's gesture data and vehicle working condition data, using the deep recursive Q network to generate target vehicle control instructions, combined with the reward function optimization model training, the optimal vehicle control strategy is generated, and the driver's intention, vehicle condition and road conditions are comprehensive decisions.

Benefits of technology

It realizes accurate response to driver intentions in harsh environments, improves driving safety and reliability, and improves the convenience and safety of vehicle control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120440060A_ABST
    Figure CN120440060A_ABST
Patent Text Reader

Abstract

The invention relates to a gesture-based vehicle control method, system and device, equipment and a medium. The gesture-based vehicle control method comprises the following steps: acquiring current vehicle control intention information expressed by a driver through a vehicle control gesture in a driving process; working condition data of the vehicle in the driving process are obtained, and the working condition data comprise current driving state data, current environment state data and a historically generated vehicle control instruction; generating a current target vehicle control instruction according to the vehicle control intention information and the working condition data; and controlling the vehicle according to the target vehicle control instruction. By implementing the gesture-based vehicle control method, convenience and safety can be considered, and the real vehicle control intention of the user can be accurately responded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle technology, and in particular to a gesture-based vehicle control method, system, device, equipment, and medium. Background Art

[0002] In the field of modern vehicle control technology, vehicle control technology is constantly innovating, showing a trend of diversification and intelligence. The current mainstream vehicle control technologies mainly include traditional manual control technology, mobile phone intelligent control technology, traditional gesture control technology, and autonomous driving technology.

[0003] Traditional manual control technology is the most basic and common method of controlling a vehicle. It relies on the driver to directly operate various vehicle control components, such as the steering wheel, accelerator, and brakes, to achieve control. Traditional gesture control technology controls the vehicle by recognizing the driver's gestures, eliminating the need for the driver to directly touch the vehicle's controls. Mobile phone intelligent control technology uses smartphones as control terminals and enables remote control of the vehicle through wireless communication technology. Autonomous driving technology achieves autonomous driving by integrating advanced equipment, complex algorithms, and control systems.

[0004] However, these technologies all face a series of challenges in practical applications, which limit their further performance improvement and widespread application. Traditional manual control technology requires the driver to have proficient driving skills and high reaction ability, mobile phone intelligent vehicle control technology is limited by network signal quality, traditional gesture control control signals are ambiguous, and autonomous driving technology is limited by weather conditions and the complexity of road conditions. In general, existing control methods have problems such as operational convenience, reliability, safety and / or difficulty in accurately responding to the driver's actual control intentions. Therefore, there is an urgent need for a vehicle control method that is both convenient and safe and can accurately respond to the user's actual control intentions. Summary of the Invention

[0005] Based on this, the present application provides a gesture-based vehicle control method, system, device, equipment and medium that can take into account convenience, safety and accurately respond to the user's actual vehicle control intentions.

[0006] In the first aspect, the present application provides a gesture-based vehicle control method, which includes: obtaining the current vehicle control intention information expressed by the driver through vehicle control gestures during driving; obtaining the operating condition data of the vehicle during driving, wherein the operating condition data includes the current driving status data, the current environmental status data and the historically generated vehicle control instructions; generating the current target vehicle control instructions based on the vehicle control intention information and the operating condition data; and controlling the vehicle according to the target vehicle control instructions.

[0007] In combination with the first aspect, in a first possible implementation manner of the first aspect, the aforementioned step of obtaining the vehicle control intention information expressed by the driver through gestures during driving includes: obtaining gesture data, wherein the gesture data is obtained by performing gesture recognition on the driver's gesture image, the gesture data includes gesture category and gesture dynamic features, and the gesture dynamic features include at least one of gesture speed features, gesture amplitude features and gesture direction features; and determining the gesture data as vehicle control intention information to represent the driver's vehicle control intention.

[0008] In combination with the first aspect, in a second possible implementation manner of the first aspect, the aforementioned step of generating target vehicle control instructions based on vehicle control intention information and operating condition data includes: inputting the operating condition data and vehicle control intention information into a deep recursive Q network; performing feature extraction and feature fusion on the operating condition data and vehicle control intention information through the network, and processing the fused features through the loop layer in the network to obtain hidden information; calculating the cumulative reward value of each vehicle control instruction corresponding to the gesture executed under the current operating condition based on the hidden information, and taking the vehicle control instruction with the largest cumulative reward value as the target vehicle control instruction under the current operating condition.

[0009] In combination with the first aspect, in a third possible implementation of the first aspect, before generating a target vehicle control instruction based on the vehicle control intention information and the operating condition data, the method also includes: calculating a reward value after the vehicle control instruction of the training sample is executed based on a reward function, wherein the reward function is used to evaluate the behavior of the vehicle executing the corresponding vehicle control instruction according to the user gesture; using the reward value to calculate the cumulative benefit, and the deviation between the cumulative benefit and the target benefit; updating the network parameters of the deep recursive Q network according to the deviation, wherein the network parameters include at least one of the weight coefficient of the reward function, the discount factor, and the learning rate.

[0010] In combination with the third possible implementation method of the first aspect, in the fourth possible implementation method of the first aspect, the reward function includes at least one of a power response reward, a path planning accuracy reward, a control feedback reward and an environmental adaptation reward; the power response reward is used to evaluate whether the actual power output of the vehicle is consistent with the driver's expected power output after the vehicle executes the corresponding vehicle control command according to the user gesture; the path planning accuracy reward is used to evaluate whether the actual driving path of the vehicle is consistent with the driver's expected driving path after the vehicle executes the corresponding vehicle control command according to the user gesture; the control feedback reward is used to evaluate whether the vehicle can adapt to environmental changes and drive stably after the vehicle executes the corresponding vehicle control command according to the user gesture; the environmental adaptation reward is used to evaluate whether the vehicle can adapt to harsh environments and drive without slipping after the vehicle executes the corresponding vehicle control command according to the user gesture.

[0011] In combination with the fourth possible implementation method of the first aspect, in the fifth possible implementation method of the first aspect, in the reward function: the power response reward includes the weighted sum of the driving torque deviation and the acceleration deviation; the path planning accuracy reward includes the weighted sum of the path lateral deviation and the steering angle deviation; the control feedback reward includes the weighted sum of the vehicle speed deviation, the lateral acceleration and the longitudinal acceleration; the environmental adaptation reward includes the weighted sum of the vehicle driving safety score and the vehicle slip penalty score.

[0012] In the second aspect, the present application provides a vehicle control system, which includes: a vehicle control device for implementing the first aspect or any one of the embodiments of the first aspect; a data collection and acquisition device for collecting gesture images and vehicle operating condition data; and a gesture recognition device for obtaining the current vehicle control intention information expressed by the driver through the vehicle control gesture during driving by performing image processing on the gesture image.

[0013] On the third aspect, the present application provides a vehicle control device, which includes: an acquisition unit for acquiring the current vehicle control intention information expressed by the driver through vehicle control gestures during driving; acquiring the operating condition data of the vehicle during driving, wherein the operating condition data includes the current driving status data, the current environmental status data and the historically generated vehicle control instructions; a generation unit for generating the current target vehicle control instructions based on the vehicle control intention information and the operating condition data; and a control unit for controlling the vehicle according to the target vehicle control instructions.

[0014] In a fourth aspect, the present application also provides a vehicle control device, which includes a processor and a memory, and the processor and the memory are connected via a bus; the processor is used to execute multiple instructions; the memory is used to store multiple instructions, and the instructions are suitable for being loaded by the processor and executed as the gesture-based vehicle control method of the first aspect or any one of the embodiments of the first aspect.

[0015] In a fifth aspect, the present application also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are suitable for being loaded and executed by a processor as a gesture-based vehicle control method such as the first aspect or any one of the embodiments of the first aspect.

[0016] In summary, the present application provides a gesture-based vehicle control method, system, device, equipment and medium. In the present application, by integrating the driver's vehicle control intention information with the vehicle operating condition data, a target vehicle control instruction that conforms to the user's driving habits and adapts to the current operating conditions can be generated. This method not only takes into account the driver's current vehicle control intention, but also incorporates his historical control instructions to ensure that the generation process of the target vehicle control instruction fully reflects the driver's driving intention and takes into account the current environmental conditions and driving conditions. Therefore, in the actual driving process, when the vehicle executes this target vehicle control instruction, it can not only accurately capture the driver's true vehicle control intention and retain the convenience of gesture control, but also significantly improve the safety and reliability of the driving process, thereby enhancing the driver's driving experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flow chart of a gesture-based vehicle control method in one embodiment;

[0018] Figure 2 is a flow chart of a gesture-based vehicle control method in another embodiment;

[0019] Figure 3 is a schematic structural diagram of a vehicle control device in one embodiment;

[0020] Figure 4 Schematic diagram of the structure of a vehicle control device in one embodiment. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0022] It should be noted that the diagrams provided in the present embodiment are only schematic illustrations of the basic concept of the present invention. The diagrams only show the components related to the present invention and are not drawn according to the number, shape and size of the components during actual implementation. The type, quantity and ratio of each component during actual implementation can be changed at will, and the component layout form may also be more complex. The structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the content disclosed in the specification for people familiar with this technology to understand and read. They are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they have no technical significance. Any modification of the structure, change of the proportional relationship or adjustment of the size should still fall within the scope of the technical content disclosed by the present invention without affecting the efficacy and purpose of the present invention. At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" quoted in this specification are only for the convenience of description and are not used to limit the scope of the implementation of the present invention. Changes or adjustments in their relative relationships should also be considered as the scope of the implementation of the present invention without substantially changing the technical content.

[0023] In the field of modern vehicle control technology, vehicle control technologies are constantly innovating, showing a trend towards diversification and intelligence. Current mainstream vehicle control technologies include traditional manual control, mobile phone intelligent control, traditional gesture control, and autonomous driving. These technologies face a series of challenges in practical application, limiting their further performance improvement and widespread adoption.

[0024] Traditional manual control technology is the most basic and time-honored method of vehicle control. The driver directly controls the vehicle's direction, speed, and start and stop through mechanical devices such as the steering wheel, brake pedal, and accelerator pedal. This method is not very convenient, requiring the driver to coordinate the operation of multiple control devices. It is highly dependent on the driver's operating skills and reaction ability, and requires the driver to make accurate judgments and timely operations under various road conditions and environments. However, in special scenarios, such as parking in a small space and driving at low speed through complex road conditions, traditional manual control methods put tremendous operating pressure on the driver, making it easy to make operational errors, which in turn affects driving safety.

[0025] Smart car control technology via mobile phone allows car owners to control the vehicle through their mobile phone. From anywhere with a mobile phone signal, they can send commands to the vehicle using their linked phone, enabling functions such as remote starting, shutting down, unlocking and unlocking the doors, and remote summoning. However, smart car control technology via mobile phone also has significant limitations. First, its operational convenience is limited by network signals, requiring a stable network connection between the phone and the vehicle. In areas with poor signal, commands may be delayed or unable to be transmitted, resulting in operational failure. Over-reliance on network signals also increases safety risks. Once the network signal fails, the vehicle may not be able to respond to commands in a timely manner, leading to safety accidents. Second, the functions of smart car control via mobile phone are relatively limited. It cannot achieve real-time and precise control of the vehicle's driving status, accurately respond to the driver's actual control intentions, and meet the user's needs for comprehensive and efficient vehicle management.

[0026] Traditional gesture control technology usually only considers a single piece of information, such as only the current gesture information, or even only the gesture category in the gesture information, but lacks consideration of other information. Other information includes driving status data, current environmental status data, historical vehicle control instructions, and at least one of the dynamic characteristics of gestures. Therefore, the control signal generated by traditional gesture control technology is ambiguous, making it difficult for the vehicle to accurately respond to the driver's actual control intention, and safety and reliability are difficult to guarantee, which makes driving safety accidents prone to occur.

[0027] Autonomous driving technology uses sensors (such as cameras and radar) to perceive the vehicle's surroundings and, through complex algorithms, makes decisions and controls to enable autonomous driving. However, existing autonomous driving technology still has many shortcomings in terms of environmental adaptability, reliability, and safety. In severe weather (such as heavy rain, snow, and dense fog), sensor performance can be affected, leading to inaccurate perception and, in turn, compromising the decision-making and control capabilities of autonomous driving technology. Furthermore, autonomous driving technology has limited ability to handle complex road conditions (such as road construction and emergency scenes), potentially leading to decision errors and increasing the risk of traffic accidents.

[0028] In general, existing control methods have problems with operational convenience, reliability, safety, and / or difficulty in accurately responding to the driver's actual control intentions. Therefore, there is an urgent need for a vehicle control method that combines convenience, safety, and the ability to accurately respond to the user's actual control intentions.

[0029] In this regard, the present application proposes a gesture-based vehicle control method, which improves on the traditional manual control technology. Since the current driving state and environment of the vehicle are usually constantly changing during driving, the present application combines vehicle control gestures and driving conditions to automatically control the vehicle, fully considering factors such as driver's intentions, vehicle conditions, and road conditions, so that the vehicle can make a comprehensive decision on the best vehicle control instructions based on the driver's current vehicle control gestures, historical vehicle control instructions, and the vehicle's current vehicle condition and road conditions, rather than relying solely on gestures to control the vehicle. By executing the target vehicle control instructions generated by the vehicle control method during driving, it can improve driving safety and reliability while accurately responding to the driver's actual vehicle control intentions.

[0030] For example, when a driver gestures to direct the vehicle forward, if the vehicle is in severe weather conditions such as fog with extremely low visibility, accelerating based solely on gestures can lead to a high risk of an accident. However, the vehicle control method provided in this application can effectively prevent blind acceleration of the vehicle even in adverse environmental conditions, thereby significantly improving driving safety and reliability.

[0031] It should be noted that the gesture-based vehicle control method provided in this application can be applied to a vehicle control system, a vehicle control device of a vehicle control system, or a vehicle control device. The vehicle control system, the vehicle control device of a vehicle control system, or the vehicle control device can be a server, computer, terminal device, controller, processor, etc. that can implement the vehicle control method of this application. In addition, the vehicle control system, the vehicle control device of a vehicle control system, or the vehicle control device can exchange data with other servers, terminal devices, controllers, processors, etc. and execute the vehicle control method proposed in this application.

[0032] In order to better understand the vehicle control method based on gesture proposed in this application, as Figure 1 As shown, this application proposes an embodiment of this method. Next, this application will combine Figure 1 The flowchart shown in the figure uses the vehicle control device as the execution subject to illustrate the gesture-based vehicle control method proposed in this application. Specifically:

[0033] 110: Obtaining the current vehicle control intention information expressed by the driver through vehicle control gestures during driving;

[0034] 120: Acquire the vehicle's operating condition data during driving, wherein the operating condition data includes current driving state data, current environmental state data, and historical vehicle control instructions;

[0035] 130: Generate the current target vehicle control instruction based on the vehicle control intention information and the working condition data;

[0036] 140: Control the vehicle according to the target vehicle control command.

[0037] In step 110, the vehicle control intention information is used to represent the driver's vehicle control intention. The vehicle control device can directly obtain the vehicle control intention information processed by other devices, or obtain the vehicle control intention information by performing gesture recognition on a gesture image including the driver's vehicle control gesture.

[0038] Control gestures can be single or combined. For example, single hand gestures can include clenching a fist, waving forward, waving left, or waving right. Combined hand gestures can include waving forward twice or clenching fists twice in a row, meeting more complex control needs.

[0039] In one practicable embodiment, a vehicle control device can obtain gesture data and identify the gesture data as vehicle control intention information. Gesture data is obtained by performing gesture recognition on the driver's gesture image. This gesture data not only covers gesture categories, used to distinguish different types of gestures, but also includes gesture dynamic characteristics. Gesture dynamic characteristics can be further subdivided into multiple sub-items, including but not limited to gesture speed characteristics, i.e., the speed of gesture movement; gesture amplitude characteristics, i.e., the range of gesture movement; and gesture direction characteristics, i.e., the specific direction of gesture movement.

[0040] Gesture categories are used to indicate the type of vehicle movement desired by the driver. Gesture categories correspond to vehicle control gestures. For example, a forward waving motion (palm extended forward with arm pushing forward) indicates the driver desires the vehicle to move forward, and the forward waving gesture is categorized as vehicle forward. Conversely, a backward waving motion (palm extended backward with arm pulling backward) indicates the driver desires the vehicle to move backward, and the backward waving gesture is categorized as vehicle backward. A leftward or rightward waving motion (palm extended sideways with arm swinging left or right) indicates the driver desires the vehicle to turn, and the leftward waving gesture is categorized as vehicle left turn, and the rightward waving gesture is categorized as vehicle right turn. A fist clenching motion (all fingers clenched into a fist) indicates the driver desires the vehicle to stop, and the fist clenching gesture is categorized as vehicle stop. Furthermore, this application provides some composite gestures. For example, two consecutive forward waving motions are categorized as acceleration forward, and two consecutive fist clenching motions are categorized as emergency braking. By defining these gesture categories, drivers' more complex and diverse control needs can be met. The gesture categories clearly communicate the type of control commands the driver wants the vehicle to perform.

[0041] Gesture dynamic features are used to indicate the intensity of vehicle movement desired by the driver. Gesture dynamic features include at least one of gesture speed, gesture amplitude, and gesture direction. For example, when the gesture category is vehicle forward movement, the gesture speed feature is used to determine the target acceleration of the vehicle's movement, and the gesture amplitude feature is used to determine the target speed of the vehicle's movement. The faster the waving speed and the larger the amplitude, the higher the forward speed and acceleration in the instruction. When the gesture category is vehicle turning right or vehicle turning left, the gesture speed feature is used to determine the target angle change rate of the vehicle's movement, and the gesture direction feature is used to determine the target steering angle of the vehicle's movement. The larger the arm swing amplitude, the larger the steering angle, and the greater the angle change rate.

[0042] At least one of these characteristics is taken into account, using gesture categories and / or gesture dynamics to more accurately represent the driver's actual control intent. Based on this comprehensive and detailed gesture data, the vehicle control device analyzes and processes it to ultimately determine the user's control intent, enabling precise vehicle control.

[0043] In step 120, the driving state data may be data such as vehicle speed and acceleration used to represent the driving state of the vehicle, and the environmental state data may include road data and weather data used to represent the driving state of the vehicle. The road data may include road surface type, road object distribution, curve curvature, road slope, and other data used to represent the road conditions around the vehicle. The weather data may be data such as sunny, rainy, and cloudy used to represent the weather conditions around the vehicle. If there are no historical vehicle control instructions, the historical vehicle control instructions are empty. If there are historical vehicle control instructions, the historical vehicle control instructions may be the vehicle control instructions generated last or the vehicle control instructions generated N times before, where N is a positive integer greater than 1. The vehicle control instructions generated last are generated by the vehicle control device based on the last driving state data, the last environmental state data, and the vehicle control instructions generated the day before last.

[0044] In step 130, the vehicle control device generates corresponding target control instructions according to the optimal control strategy. These control instructions may include drive torque T and steering angle θ. After receiving these instructions via the CAN bus, the vehicle control device or other control device parses their content and converts them into specific control signals, which are then sent to the vehicle's corresponding actuators, such as the motor controller, steering motor, and braking system. This allows the vehicle's driving state to be precisely controlled, thus enabling gesture-based vehicle control.

[0045] The optimal control strategy can be generated by a reinforcement learning model that can process time series information, such as Deep Recurrent Q-Network (DRQN), Distributed DQN (R2D2), Asynchronous Advantage Actor-Critic (A3C), and each of the aforementioned models includes a loop layer for processing time series information. The loop layer can be, for example, a Long Short-Term Memory Network (LSTM) or a Gated Recurrent Unit (GRU), etc. The aforementioned time series information includes historically generated control instructions in this application. The aforementioned reinforcement learning model transmits historical information through hidden states to enhance the model's modeling ability for time series data. By incorporating the time dimension into the state representation, the model can make continuous decisions based on historical experience.

[0046] In the reinforcement learning model, the aforementioned vehicle control intention information and operating condition data together constitute the state space State of the reinforcement learning model, while the control instructions are used to represent the vehicle action space Action.

[0047] To guide the model to learn the optimal control strategy, it is crucial to design a reasonable reward mechanism. When the vehicle's driving state is stable under the current gesture operation, meets the driver's expectations, and can adapt to the surrounding environment (for example, smoothly turning around a curve, maintaining a stable driving on a bumpy road, etc.), a positive reward should be given. Conversely, if the vehicle becomes unstable (such as speeding, oversteering, etc.) or fails to adapt to environmental changes, a negative reward should be given. For example, on a slippery road in the rain, if the vehicle can smoothly decelerate and avoid skidding through gesture operation, a high positive reward will be obtained; if the gesture operation causes the vehicle to lose control, a large negative reward will be obtained.

[0048] The reinforcement learning model is continuously trained based on a reward mechanism. During each training session, the model selects a control command based on the current operating conditions based on the principle of maximizing rewards. After the vehicle executes the control command, it provides reward feedback based on the actual driving performance. The model adjusts its parameters based on this reward feedback, increasing the probability of selecting the optimal control command under the same or similar operating conditions. Through extensive training, the model gradually learns the optimal control strategy for different operating conditions, thereby improving the vehicle control device's adaptability to various gestures and operating conditions.

[0049] Preferably, the DRQN model is used to generate the optimal control strategy. Compared with A3C and R2D2, the DRQN model is suitable for rapid development, especially in some observable environments and resource-constrained tasks. Next, this application uses the DRQN model as an example to illustrate the process of instruction generation:

[0050] The aforementioned step of generating target vehicle control instructions based on vehicle control intention information and operating condition data includes: inputting the operating condition data and vehicle control intention information into a deep recursive Q network; performing feature extraction and feature fusion on the operating condition data and vehicle control intention information through the network, and processing the fused features through the loop layer in the network to obtain hidden information; calculating the cumulative reward value of each vehicle control instruction corresponding to the execution gesture under the current operating condition based on the hidden information, and taking the vehicle control instruction with the largest cumulative reward value as the target vehicle control instruction under the current operating condition.

[0051] Among them, the deep recursive Q network can include an input layer, at least one convolutional layer, at least one fully connected layer, a loop layer and an output layer. The working condition data and vehicle control intention information are input into the input layer of the deep recursive Q network, and the working condition data and vehicle control intention information are subjected to feature extraction and feature fusion through the convolutional layer and fully connected layer of the network. The fused features are processed through the loop layer in the network to obtain hidden information, and then the hidden information is input into the fully connected layer. The cumulative reward value of each vehicle control instruction corresponding to the gesture executed under the current working condition is calculated through the fully connected layer, and the vehicle control instruction with the largest cumulative reward value is determined as the target vehicle control instruction.

[0052] Next, this application still takes the DRQN model as an example to illustrate the model training process: before generating the target vehicle control instruction based on the vehicle control intention information and working condition data, the reward value after the vehicle control instruction of the training sample is executed is calculated based on the reward function, wherein the reward function is used to evaluate the behavior of the vehicle executing the corresponding vehicle control instruction according to the user gesture; the reward value is used to calculate the cumulative benefit, as well as the deviation between the cumulative benefit and the target benefit; the network parameters of the deep recursive Q network are updated according to the deviation, wherein the network parameters include at least one of the weight coefficient, discount factor, and learning rate of the reward function.

[0053] Among them, the reward function is the core mechanism in reinforcement learning (including algorithms such as DRQN), which guides the agent to learn the optimal strategy through quantified feedback signals.

[0054] In one possible implementation, the reward function R total Including Power Response Reward R power , path planning accuracy reward R path , control feedback reward R control and environmental adaptation reward R envAt least one of the following: the power response reward is used to evaluate whether the actual power output of the vehicle is consistent with the driver's expected power output after the vehicle executes the corresponding vehicle control command according to the user gesture; the path planning accuracy reward is used to evaluate whether the actual driving path of the vehicle is consistent with the driver's expected path after the vehicle executes the corresponding vehicle control command according to the user gesture; the control feedback reward is used to evaluate whether the vehicle can adapt to environmental changes and drive stably after the vehicle executes the corresponding vehicle control command according to the user gesture; the environmental adaptation reward is used to evaluate whether the vehicle can adapt to harsh environments and drive without slipping after the vehicle executes the corresponding vehicle control command according to the user gesture.

[0055] In another possible implementation, the reward function R total Including Power Response Reward R power , path planning accuracy reward R path , control feedback reward R control and environmental adaptation reward R env , R total =R power +R path +R control +R env For example, the dynamic response reward includes the weighted sum of the driving torque deviation and the acceleration deviation, the path planning accuracy reward includes the weighted sum of the path lateral deviation and the steering angle deviation, the control feedback reward includes the weighted sum of the vehicle speed deviation, the lateral acceleration and the longitudinal acceleration, and the environmental adaptation reward includes the weighted sum of the vehicle driving safety score and the vehicle skidding penalty score. The mathematical expressions of each reward function include:

[0056] Power Response Reward: R power =-ω1·|TT target |-ω2·|aa target |; Among them, R power is the power response reward, T is the current driving torque, T target is the target driving torque, TT target is the driving torque deviation, a is the current acceleration, a target is the target acceleration, ω1 and ω2 are weight coefficients for balancing the contribution of driving torque and acceleration, aa target is the acceleration deviation;

[0057] Path planning accuracy reward: R path =-ω3·d lateral -ω2·|θ-θ target |; Among them, R path is the path planning accuracy reward, d lateral is the lateral deviation between the actual driving path and the expected driving path, θ is the current steering angle, and θ targetis the target steering angle, ω3 and ω4 are weight coefficients for balancing the contribution of lateral deviation and steering angle, θ-θ target Indicates steering angle deviation;

[0058] Control feedback reward: R control =-ω5·|vv target |-ω6·|a lateral |-ω7·|a vertical |; Among them, R control is the control feedback reward, v is the current vehicle speed, v target is the target speed, a lateral is the lateral acceleration (used to measure steering stability), a vertical is the longitudinal acceleration (used to measure the smoothness of bumpy roads), ω5, ω6 and ω7 are weight coefficients used to balance the contribution of vehicle speed, lateral acceleration and vertical acceleration, and vv target is the vehicle speed deviation;

[0059] Acclimation Bonus: R env =ω8·safety_score-ω9·slip_penalty, where R env is the environmental adaptation reward, safety_score is the safety score (such as whether the vehicle is stable, whether it avoids skidding, etc.), slip_penalty is the slip penalty (such as the negative reward when the tire slips), ω8 and ω9 are weight coefficients used to balance the contribution of safety and slip penalty.

[0060] It should be noted that the previous mathematical formulas are a representation of the reward function. In fact, in order to ensure that the sub-items with different dimensions in the reward function can be weighted and summed, the sub-items can also be normalized to convert different physical quantities into comparable dimensionless values.

[0061] For example, in the path planning accuracy reward R path In the lateral and |θ-θ target |After normalization, weighted summation is performed. The sub-items of different dimensions can be added after normalization.

[0062] Moreover, the normalization operation can also improve the problem of excessive or too small rewards for certain sub-items in the reward function, because the numerical ranges of multiple sub-items of the reward function may vary greatly. If added directly, the rewards of some sub-items will dominate the rewards of other sub-items, causing the model to ignore the optimization goals of other sub-items. Therefore, after calculating the reward values of each sub-item of the reward function, the reward values of each sub-item can be normalized, and then the normalized values can be weighted summed.

[0063] The training sample set of the reinforcement learning model contains samples of multiple consecutive time steps, each of which can be represented as a four-tuple (s t , a t , r, s t+1 ), where s t Including the vehicle control intention information and working condition data at time step t, and the vehicle control instructions at time step t-1; a t Including the vehicle control instructions of t time steps; r represents the reward value, s t+1 This includes the vehicle control intention information and operating condition data at time step t+1, as well as the vehicle control instructions at time step t. The training method involves constructing an experience replay pool, extracting a subset of samples from the pool, inputting them into the target model to obtain the cumulative reward (i.e., Q-value) and target reward (i.e., target Q-value), updating the network parameters of the deep recursive Q network based on the deviation, and repeating the training operation until the model meets the convergence conditions, resulting in a trained reinforcement learning model.

[0064] It should be noted that the Q value is obtained by online network processing of the deep recursive Q network, and the target Q value is obtained by target network processing of the deep recursive Q network. The network parameters of the online network of the deep recursive Q network are updated in real time according to the deviation, and the network parameters of the target network are updated by regularly copying the trained network parameters of the online Q network to the target network.

[0065] In one implementation, the value update function used by DQRN includes: Among them, s t is the current state, a t is the current action, r t+1 For reward, t+1 is the next state, γ is the discount factor, α is the learning rate, Q(s t ,a t ) is to execute the current action a t The cumulative revenue after Q(s t+1 ,a), represents the target return, Indicates the target value mentioned above.

[0066] In another embodiment, the present application also provides a vehicle control system, which includes: a vehicle control device for generating target vehicle control instructions through the aforementioned embodiment or implementable method; a data collection and acquisition device for collecting gesture images and vehicle operating condition data; a gesture recognition device for obtaining the current vehicle control intention information expressed by the driver through vehicle control gestures during driving by performing image processing on the gesture images.

[0067] In order to better understand the vehicle control method applied to the vehicle control system, the present application also provides Figure 2 Flowchart, next, combined with Figure 2 To explain:

[0068] Figure 2 The data acquisition and retrieval device is used to collect data and gesture definitions, for example, including collecting the driver's gesture information (such as gesture images, etc.), collecting current operating condition data in real time, and collecting the user's driving condition history data (including slope, steering angle, weather, road surface or road type, etc.); and the data acquisition and retrieval device continuously determines whether the data is complete and continues to collect if it is incomplete;

[0069] Figure 2 The gesture recognition device is used to extract key features of gestures and recognize the driver's gestures based on the gesture recognition algorithm and the defined gestures, thereby generating gesture data. The gesture data includes gesture categories and gesture dynamic features. For example, multimodal fusion technology can be used to extract features from gesture data. The image features of the gestures are obtained through image processing and processed through a deep neural network architecture to extract the spatial position and motion trajectory features of the gestures. The deep features are combined with the image features to obtain gesture data.

[0070] Figure 2 The signal processing device is used to define the state space State (such as control instructions, speed, acceleration, working conditions, etc.), define the action space Action (such as driving torque, steering angle, etc.), design the reward function Reward (such as power response, path planning accuracy, control strategy feedback, etc.), and generate vehicle control instructions (i.e. target control instructions) based on the defined State, Action and Reward based on a deep reinforcement learning algorithm (such as DRQN). Figure 2 The vehicle control device receives the target control instruction generated by the signal processing device and executes the target control instruction.

[0071] It should be understood that although Figure 1 and Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 and Figure 2At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these sub-steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0072] join Figure 3 The present application also provides a vehicle control device. The embodiment of the present application can divide the functional modules of the device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. Specifically, the vehicle control device includes: an acquisition unit 310, which is used to obtain the current vehicle control intention information expressed by the driver through the vehicle control gesture during driving; the acquisition unit 310 is also used to obtain the working condition data of the vehicle during driving, wherein the working condition data includes the current driving state data, the current environmental state data and the historically generated vehicle control instructions; the generation unit 320 is used to generate the current target vehicle control instructions based on the vehicle control intention information and the working condition data; the control unit 330 is used to control the vehicle according to the target vehicle control instructions.

[0073] In one practicable manner, the aforementioned acquisition unit 310 is specifically used to: acquire gesture data, wherein the gesture data is obtained by performing gesture recognition on the driver's gesture image, the gesture data includes gesture category and gesture dynamic features, and the gesture dynamic features include at least one of gesture speed features, gesture amplitude features, and gesture direction features; and determine the gesture data as vehicle control intention information to represent the driver's vehicle control intention.

[0074] In one practicable manner, the aforementioned generation unit 320 is specifically used to: input the operating condition data and the vehicle control intention information into a deep recursive Q network; perform feature extraction and feature fusion on the operating condition data and the vehicle control intention information through the network, and process the fused features through the loop layer in the network to obtain hidden information; calculate the cumulative reward value of each vehicle control instruction corresponding to the gesture executed under the current operating condition based on the hidden information, and use the vehicle control instruction with the largest cumulative reward value as the target vehicle control instruction under the current operating condition.

[0075] In one practicable embodiment, the vehicle control device further includes a training unit 330, which is used to: calculate a reward value after the vehicle control command of the training sample is executed based on a reward function, wherein the reward function is used to evaluate the behavior of the vehicle executing the corresponding vehicle control command according to the user gesture; use the reward value to calculate the cumulative benefit, and the deviation between the cumulative benefit and the target benefit; update the network parameters of the deep recursive Q network according to the deviation, wherein the network parameters include at least one of the weight coefficient of the reward function, the discount factor, and the learning rate.

[0076] In one feasible manner, the reward function includes at least one of a power response reward, a path planning accuracy reward, a control feedback reward and an environmental adaptation reward; the power response reward is used to evaluate whether the actual power output of the vehicle is consistent with the driver's expected power output after the vehicle executes the corresponding vehicle control command according to the user gesture; the path planning accuracy reward is used to evaluate whether the actual driving path of the vehicle is consistent with the driver's expected driving path after the vehicle executes the corresponding vehicle control command according to the user gesture; the control feedback reward is used to evaluate whether the vehicle can adapt to environmental changes and drive stably after the vehicle executes the corresponding vehicle control command according to the user gesture; the environmental adaptation reward is used to evaluate whether the vehicle can adapt to harsh environments and drive without slipping after the vehicle executes the corresponding vehicle control command according to the user gesture.

[0077] In one feasible manner, in the reward function: the power response reward includes the weighted sum of the driving torque deviation and the acceleration deviation; the path planning accuracy reward includes the weighted sum of the path lateral deviation and the steering angle deviation; the control feedback reward includes the weighted sum of the vehicle speed deviation, the lateral acceleration and the longitudinal acceleration; the environmental adaptation reward includes the weighted sum of the vehicle driving safety score and the vehicle slip penalty score.

[0078] join Figure 4The present application also provides a vehicle control device, which may include: a processor 410 and a memory 420. The processor 410 and the memory 420 are connected via a bus 430. The processor 410 is configured to execute multiple instructions; the memory is configured to store multiple instructions, which are suitable for being loaded by the processor and executed as described in the above embodiment of the remote testing method based on the human-computer interaction interface. The processor may be an electronic control unit (ECU), a central processing unit (CPU), a general-purpose processor, a coprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of an 5SP and a microprocessor, and so on. In this embodiment, the processor may be a single-chip microcomputer. Various control functions can be implemented by programming the single-chip microcomputer. The processor has the advantages of strong computing power and fast processing speed. Specifically, the processor 410 is used to execute the function of the acquisition unit 310, which is used to obtain the current vehicle control intention information expressed by the driver through vehicle control gestures during driving; it is also used to obtain the operating condition data of the vehicle during driving, wherein the operating condition data includes the current driving status data, the current environmental status data and the historically generated vehicle control instructions; the processor 410 is also used to execute the function of the generation unit 320, which is used to generate the current target vehicle control instructions based on the vehicle control intention information and the operating condition data; the processor 410 is also used to execute the function of the control unit 330, which is used to control the vehicle according to the target vehicle control instructions.

[0079] In one practicable manner, the aforementioned processor 410 is specifically used to: obtain gesture data, wherein the gesture data is obtained by performing gesture recognition on the driver's gesture image, the gesture data includes gesture category and gesture dynamic features, and the gesture dynamic features include at least one of gesture speed features, gesture amplitude features and gesture direction features; determine the gesture data as vehicle control intention information to represent the driver's vehicle control intention.

[0080] In one practicable manner, the aforementioned processor 410 is specifically used to: input the operating condition data and vehicle control intention information into a deep recursive Q network; perform feature extraction and feature fusion on the operating condition data and vehicle control intention information through the network, and process the fused features through the loop layer in the network to obtain hidden information; calculate the cumulative reward value of each vehicle control instruction corresponding to the gesture executed under the current operating condition based on the hidden information, and use the vehicle control instruction with the largest cumulative reward value as the target vehicle control instruction under the current operating condition.

[0081] In one practicable embodiment, the processor 410 is further configured to execute the functions of the training unit 330: calculating a reward value after the vehicle control command of the training sample is executed based on a reward function, wherein the reward function is used to evaluate the behavior of the vehicle in executing the corresponding vehicle control command according to the user gesture; calculating the cumulative benefit and the deviation between the cumulative benefit and the target benefit using the reward value; and updating the network parameters of the deep recursive Q network according to the deviation, wherein the network parameters include at least one of the weight coefficient of the reward function, the discount factor, and the learning rate.

[0082] In one feasible manner, the reward function includes at least one of a power response reward, a path planning accuracy reward, a control feedback reward and an environmental adaptation reward; the power response reward is used to evaluate whether the actual power output of the vehicle is consistent with the driver's expected power output after the vehicle executes the corresponding vehicle control command according to the user gesture; the path planning accuracy reward is used to evaluate whether the actual driving path of the vehicle is consistent with the driver's expected driving path after the vehicle executes the corresponding vehicle control command according to the user gesture; the control feedback reward is used to evaluate whether the vehicle can adapt to environmental changes and drive stably after the vehicle executes the corresponding vehicle control command according to the user gesture; the environmental adaptation reward is used to evaluate whether the vehicle can adapt to harsh environments and drive without slipping after the vehicle executes the corresponding vehicle control command according to the user gesture.

[0083] In one feasible manner, in the reward function: the power response reward includes the weighted sum of the driving torque deviation and the acceleration deviation; the path planning accuracy reward includes the weighted sum of the path lateral deviation and the steering angle deviation; the control feedback reward includes the weighted sum of the vehicle speed deviation, the lateral acceleration and the longitudinal acceleration; the environmental adaptation reward includes the weighted sum of the vehicle driving safety score and the vehicle slip penalty score.

[0084] This application also provides a computer-readable storage medium storing a plurality of instructions suitable for being loaded by a processor and executing the method of any of the aforementioned embodiments. The processor is configured to execute the plurality of instructions; and the memory is configured to store the plurality of instructions, which are loaded by the processor and executed by the vehicle control method of the aforementioned embodiments.

[0085] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0086] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A vehicle control method based on gestures, characterized in that: include: Obtain the driver's current vehicle control intention information expressed through vehicle control gestures during driving; Acquiring operating condition data of the vehicle during driving, wherein the operating condition data includes current driving state data, current environmental state data, and historically generated vehicle control instructions; Generate the current target vehicle control instruction according to the vehicle control intention information and the working condition data; The vehicle is controlled according to the target vehicle control instruction.

2. The method according to claim 1, characterized in that The step of obtaining the current vehicle control intention information expressed by the driver through the vehicle control gesture during driving includes: Acquiring gesture data, wherein the gesture data is obtained by performing gesture recognition on a gesture image of the driver, the gesture data including a gesture category and a gesture dynamic feature, the gesture dynamic feature including at least one of a gesture speed feature, a gesture amplitude feature, and a gesture direction feature; The gesture data is determined as vehicle control intention information to represent the driver's vehicle control intention.

3. The method according to claim 1, characterized in that The step of generating the current target vehicle control instruction according to the vehicle control intention information and the operating condition data includes: Inputting the operating condition data and vehicle control intention information into a deep recursive Q network; The network is used to extract and fuse the working condition data and the vehicle control intention information, and the fused features are processed through a loop layer in the network to obtain hidden information; The cumulative reward value of each vehicle control instruction corresponding to the execution gesture under the current working condition is calculated according to the hidden information, and the vehicle control instruction with the largest cumulative reward value is used as the target vehicle control instruction under the current working condition.

4. The method according to claim 1, wherein Before generating the current target vehicle control instruction according to the vehicle control intention information and the operating condition data, the method further includes: Calculating a reward value after the vehicle control command of the training sample is executed based on a reward function, wherein the reward function is used to evaluate the vehicle's behavior of executing the corresponding vehicle control command according to the user gesture; Calculating the cumulative benefit and the deviation between the cumulative benefit and the target benefit using the reward value; Network parameters of a deep recursive Q network are updated according to the deviation, wherein the network parameters include at least one of a weight coefficient of a reward function, a discount factor, and a learning rate.

5. The method according to claim 4, characterized in that The reward function includes at least one of a power response reward, a path planning accuracy reward, a control feedback reward, and an environment adaptation reward; The power response reward is used to evaluate whether the vehicle's actual power output is consistent with the driver's expected power output after the vehicle executes the corresponding vehicle control command according to the user's gesture; The path planning accuracy reward is used to evaluate whether the vehicle's actual driving path is consistent with the driver's intended driving path after the vehicle executes the corresponding vehicle control command according to the user's gesture; Control feedback rewards are used to evaluate whether the vehicle can adapt to environmental changes and maintain stable driving after executing corresponding control commands according to user gestures; The environmental adaptation reward is used to evaluate whether the vehicle can adapt to harsh environments and drive without slipping after executing the corresponding vehicle control commands according to user gestures.

6. The method according to claim 5, characterized in that In the reward function: The power response reward includes the weighted sum of the driving torque deviation and the acceleration deviation; The path planning accuracy reward includes the weighted sum of the path lateral deviation and the steering angle deviation; The control feedback reward includes the weighted sum of vehicle speed deviation, lateral acceleration and longitudinal acceleration; The environmental adaptation reward includes the weighted sum of the vehicle driving safety score and the vehicle skidding penalty score.

7. A vehicle control system, characterized in that: The vehicle control system includes: A vehicle control device for executing the method according to any one of claims 1 to 6; Data collection and acquisition device, used to collect gesture images and vehicle operating condition data; The gesture recognition device is used to obtain the current vehicle control intention information expressed by the driver through the vehicle control gesture during driving by performing image processing on the gesture image.

8. A vehicle control device, characterized in that: The vehicle control device comprises: an acquisition unit, configured to acquire the current vehicle control intention information expressed by the driver through vehicle control gestures during driving; and acquire operating condition data of the vehicle during driving, wherein the operating condition data includes current driving state data, current environmental state data, and historically generated vehicle control instructions; A generating unit, configured to generate a target vehicle control instruction according to the vehicle control intention information and the operating condition data; A control unit is used to control the vehicle according to the target vehicle control instruction.

9. A vehicle control device, characterized in that: The vehicle control device includes a processor and a memory, which are connected via a bus; the processor is used to execute multiple instructions; the memory is used to store the multiple instructions, and the instructions are suitable for being loaded by the processor and executing the gesture-based vehicle control method described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executed by the gesture-based vehicle control method according to any one of claims 1 to 6.