A method, system and device for autonomous decision-making of intelligent connected vehicles based on reinforcement learning

By adopting reinforcement learning and DQN algorithms in the intelligent connected vehicle autonomous decision-making system, combined with the information of on-board sensors and connected devices, the rationality and accuracy of vehicle autonomous decision-making in urban road scenarios are solved, and more efficient understanding and decision-making of traffic conditions are achieved.

CN119296054BActive Publication Date: 2025-05-16SHENZHEN URBAN TRANSPORT PLANNING CENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411816443.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-16
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the rationality and accuracy of autonomous decision-making of vehicles, especially in urban road scenarios, where it is difficult for vehicles to fully understand the overall traffic conditions and environmental changes.

Method used

An intelligent connected vehicle autonomous decision-making system based on reinforcement learning is adopted. The system receives information from on-board sensors and connected devices through the data acquisition module, uses the DQN algorithm and value function to make decisions and corrections, and ultimately realizes autonomous decision-making of vehicles.

Benefits of technology

By leveraging the global road conditions information provided by IoT devices, vehicles can have a more comprehensive understanding of traffic road conditions and environmental changes, thereby improving the rationality and accuracy of independent decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119296054B_ABST
    Figure CN119296054B_ABST
Patent Text Reader

Abstract

A method, system and device for autonomous decision-making of an intelligent networked vehicle based on reinforcement learning, relates to the field of automatic driving, and specifically to a method, system and device for autonomous decision-making of an intelligent networked vehicle based on reinforcement learning. In order to solve the problem of how to improve the rationality and accuracy of vehicle autonomous decision-making in the prior art, the present invention provides the following solution: the system includes data acquisition, autonomous driving decision-making and control modules. The data acquisition module processes sensor data and message information and sends it to the autonomous driving decision-making module. The autonomous driving decision-making module uses the DQN algorithm to convert data into preliminary decisions, and uses the value function to correct the final decision. The control module controls the vehicle driving state according to the final decision to achieve efficient autonomous driving. The present invention has good application prospects in vehicle intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving, and in particular to a method, system and device for autonomous decision-making of an intelligent connected vehicle based on reinforcement learning. Background Art

[0002] The present invention belongs to the field of the combination of artificial intelligence and network connection technology, and specifically introduces a method, system and related devices for an intelligent networked vehicle based on reinforcement learning in an urban road scenario to make autonomous decisions by receiving surrounding road condition information provided by an Internet of Things device.

[0003] In recent years, with the maturity of artificial intelligence technology, self-driving cars have become a development trend. Driving behavior decision-making is an important functional module of smart cars, and it is also the focus and difficulty of autonomous driving technology research. Driving behavior decision-making is a safe and reasonable driving operation made based on the environmental information obtained by the system and the current vehicle status. The quality of its decision is an important indicator to measure the level of vehicle intelligence.

[0004] However, in this field, the key technical issues mainly focus on how to ensure the rationality and accuracy of vehicle autonomous decision-making. Solving these problems requires in-depth research on data processing, algorithm optimization and system design to achieve a more reliable and efficient autonomous driving experience. Summary of the invention

[0005] In order to solve the problem of how to improve the rationality and accuracy of vehicle autonomous decision-making in the prior art, the present invention provides the following solutions:

[0006] An autonomous decision-making system for an intelligent networked vehicle based on reinforcement learning, the system comprising: a data acquisition module, an autonomous driving decision-making module and a control module;

[0007] The data acquisition module is used to pre-process the collected data of the vehicle-mounted sensor and the message information of the network-connected device to obtain the pre-processed collected data and the pre-processed message information, and send them to the autonomous driving decision module;

[0008] The autonomous driving decision module is used to convert the pre-processed collected data into a preliminary driving decision through the DQN algorithm, and is also used to construct the pre-processed message information into a value function, and use the value function to correct the preliminary driving decision, thereby obtaining a final vehicle autonomous decision, and at the same time sending the final vehicle autonomous decision to the control module;

[0009] The control module is used to control the vehicle driving state according to the final vehicle autonomous decision.

[0010] Furthermore, the specific steps of preprocessing the collected data of the vehicle-mounted sensor are:

[0011] The data acquisition module obtains the collected data of the vehicle-mounted sensor through the RTSP video stream address;

[0012] Decoding the collected data of the vehicle-mounted sensor into a single-frame image in a unified RGB format;

[0013] The single frame image is subjected to color space conversion and image filtering and denoising processing.

[0014] Furthermore, the specific steps of pre-processing the message information of the networked device are:

[0015] The data acquisition module receives message information from networked devices through V2X technology;

[0016] The data acquisition module reads the key fields of the message information of the networked device and saves it in the local environment.

[0017] Furthermore, the DQN algorithm is specifically as follows:

[0018] S1. Random initialization state, autonomous decision-making system has The probability of selecting an action through Q-Network , The probability of performing random actions;

[0019] S2. After the action is completed, the state , vehicle information , reward value and the next state Save to experience pool;

[0020] S3. Repeat S1 and S2 until the experience value is full, and use the experience pool to update the target Q-Network;

[0021] S4. After the target Q-Network converges, the preprocessed collected data is used as input to obtain the preliminary driving decision.

[0022] Furthermore, the value function is:

[0023] ;

[0024] in is the reward function, is the penalty function;

[0025] The value functions include: speed value function, lane change value function and parking value function;

[0026] The speed value function is:

[0027] ;

[0028] in, is the speed bonus factor, The lower speed limit of the road. The upper speed limit of the road. is the current speed of the vehicle;

[0029] The lane change value function is:

[0030] ;

[0031] in, is the lane congestion factor;

[0032] The parking value function is:

[0033] ;

[0034] in, is the road condition factor.

[0035] An autonomous decision-making method for an intelligent networked vehicle based on reinforcement learning, the method comprising: a data collection step, an autonomous driving decision-making step and a control step;

[0036] The data collection step is used to pre-process the collected data of the vehicle-mounted sensor and the message information of the networked device to obtain the pre-processed collected data and the pre-processed message information;

[0037] The autonomous driving decision step is used to convert the pre-processed collected data into a preliminary driving decision through the DQN algorithm, and is also used to construct the pre-processed message information into a value function, and use the value function to correct the preliminary driving decision, thereby obtaining a final vehicle autonomous decision;

[0038] The control step is used to control the vehicle driving state according to the final vehicle autonomous decision.

[0039] The present invention also provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the above-mentioned monitoring method.

[0040] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned monitoring method is implemented.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] Compared with the traditional method that only relies on the information within the limited distance range of the vehicle-mounted sensor to build a value function to make decisions, the present invention uses the Internet of Things devices to provide global road condition information and synchronize it to the vehicle decision-making system, so that the vehicle can have a more comprehensive understanding of the overall traffic conditions and environmental changes, thereby improving the rationality and accuracy of the vehicle's autonomous decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a structural diagram of the autonomous decision-making system of the present invention;

[0044] Figure 2 This is the flow chart of the DQM algorithm. DETAILED DESCRIPTION

[0045] Embodiment 1, combination Figure 1 and Figure 2 This embodiment is described. An autonomous decision-making system for an intelligent networked vehicle based on reinforcement learning, the system comprising: a data acquisition module, an autonomous driving decision-making module and a control module;

[0046] The data acquisition module is used to pre-process the collected data of the vehicle-mounted sensor and the message information of the network-connected device to obtain the pre-processed collected data and the pre-processed message information, and send them to the autonomous driving decision module;

[0047] The autonomous driving decision module is used to convert the pre-processed collected data into a preliminary driving decision through the DQN algorithm, and is also used to construct the pre-processed message information into a value function, and use the value function to correct the preliminary driving decision, thereby obtaining a final vehicle autonomous decision, and at the same time sending the final vehicle autonomous decision to the control module;

[0048] The control module is used to control the vehicle driving state according to the final vehicle autonomous decision.

[0049] Specifically, through Figure 1 It can be learned that the system first receives the vehicle sensor data and the message information of the networked device through the data acquisition module, and pre-processes the input data respectively. The autonomous driving decision module relies on the Hailo chip of the vehicle computing platform to input the pre-processed vehicle sensor data into the reinforcement learning DQN algorithm to infer the preliminary driving decision. Based on the message information of the pre-processed networked device, the value function is constructed to correct the preliminary driving decision in real time to form the final driving decision. Improve the rationality and real-time performance of the vehicle's autonomous decision-making. Send the driving decision to the control module to control the vehicle's actions.

[0050] The specific steps for preprocessing the collected data of the vehicle-mounted sensor are:

[0051] The data acquisition module obtains the collected data of the vehicle-mounted sensor through the RTSP video stream address;

[0052] Decoding the collected data of the vehicle-mounted sensor into a single-frame image in a unified RGB format;

[0053] The single frame image is subjected to color space conversion and image filtering and denoising processing.

[0054] The specific steps for preprocessing the message information of the networked device are:

[0055] The data acquisition module receives message information from networked devices through V2X technology;

[0056] The data acquisition module reads the key fields of the message information of the networked device and saves it in the local environment.

[0057] Specifically, the data acquisition module can also receive the radar point cloud data and perform preprocessing. The implementation steps are as follows:

[0058] Use height threshold and RANSAC plane fitting to remove ground points to highlight the target objects above the ground;

[0059] Remove outliers and noise points through statistical filtering (radius filtering, conditional filtering) to retain valid data;

[0060] Use registration algorithms (ICP, NDT) to align point cloud data from different frames to generate a continuous, seamless point cloud map;

[0061] Divide the point cloud into voxel grids and replace all points within the grid with the centroid of each grid to reduce the amount of data and improve processing efficiency;

[0062] Randomly select part of the point cloud data to reduce data density and reduce the computational burden.

[0063] The DQN algorithm is specifically:

[0064] S1. Random initialization state, autonomous decision-making system has The probability of selecting an action through Q-Network , The probability of performing random actions;

[0065] S2. After the action is completed, the state , vehicle information , reward value and the next state Save to experience pool;

[0066] S3. Repeat S1 and S2 until the experience value is full, and use the experience pool to update the target Q-Network;

[0067] S4. After the target Q-Network converges, the preprocessed collected data is used as input to obtain the preliminary driving decision.

[0068] Specific, combined Figure 2 It can be seen that the system first randomly initializes the state and selects an action through Q-Network according to a certain probability, or performs a random action with another probability. After each action, the system saves the state, vehicle information, reward value, and next state to the experience pool, and repeats this process until the experience pool is full.

[0069] When updating the target Q-Network, the system uses the data in the experience pool to calculate the loss function, where the loss function includes the difference between the neural network output value and the target value. The loss function is

[0070] ;

[0071] in, y For the target Q-Network, the reward value is and the next state The result of reasoning.

[0072] At the same time, preliminary driving decisions include control operations such as acceleration, deceleration, lane change and parking.

[0073] The value function is:

[0074] ;

[0075] in is the reward function, is the penalty function;

[0076] The value functions include: speed value function, lane change value function and parking value function;

[0077] The speed value function is:

[0078] ;

[0079] in, is the speed bonus factor, The lower speed limit of the road. The upper speed limit of the road. is the current speed of the vehicle;

[0080] Specifically, According to the positive correlation and real-time adjustment of road vehicle density monitored by networked equipment, The value range is , vehicle speed adjustment reference for initial driving decision , positive value means acceleration, negative value means deceleration.

[0081] The lane change value function is:

[0082] ;

[0083] in, is the lane congestion factor;

[0084] Specifically, According to the vehicle density adjustment of the lane where the vehicle is currently located monitored by the connected equipment, The value range is ,according to The value of covers the initial lane-changing behavior of the vehicle. When it is positive, the decision to change lanes is made. Change lane to the left. If the value is positive, the lane will change to the right; if the value is negative, the initial decision result will be maintained.

[0085] The parking value function is:

[0086] ;

[0087] in, is the road condition factor.

[0088] Specifically, The connected equipment monitors whether there are any vehicles driving illegally or in accidents in the lane where the vehicle is currently located, and makes real-time adjustments. The value range is ,when A negative value overrides the initial decision and ultimately decides that the vehicle should stop. If it is greater than zero, it is also necessary to The value is used to slow down by a percentage.

[0089] The final vehicle autonomous decision is based on three value functions , as well as A composite decision is obtained, with the highest priority being When the decision vehicle stops, all other vehicle control decisions are canceled. The final vehicle autonomous decision is organized into a structured message and sent to the control module.

[0090] Here is an example to illustrate the structured message format:

[0091] {

[0092] "id": "xxxx",

[0093] "type": "1",

[0094] "uuid": "0kz5jlizotnlov2amct7fd4jl3s4kqza",

[0095] "time": "20210313155853",

[0096] "acttion1":"speed:XXX",

[0097] "acttion2":"change:XXX",

[0098] "acttion3":"stop:XXX"

[0099] }

[0100] Field Description:

[0101]

[0102] The autonomous driving decision-making module based on the Hailo chip integrates the final vehicle autonomous decision information such as time, device code and unique identifier into a structured message and sends it to the vehicle control module to execute the decision-making behavior.

[0103] An autonomous decision-making method for an intelligent networked vehicle based on reinforcement learning, the method comprising: a data collection step, an autonomous driving decision-making step and a control step;

[0104] The data collection step is used to pre-process the collected data of the vehicle-mounted sensor and the message information of the networked device to obtain the pre-processed collected data and the pre-processed message information;

[0105] The autonomous driving decision step is used to convert the pre-processed collected data into a preliminary driving decision through the DQN algorithm, and is also used to construct the pre-processed message information into a value function, and use the value function to correct the preliminary driving decision, thereby obtaining a final vehicle autonomous decision;

[0106] The control step is used to control the vehicle driving state according to the final vehicle autonomous decision.

[0107] Embodiment 2: The computer device of the present invention may be a device including a processor and a memory, such as a single chip microcomputer including a central processing unit, etc. Moreover, the processor is used to implement the steps of the above detection method when executing a computer program stored in the memory.

[0108] Embodiment 3: Computer readable storage medium embodiment.

[0109] The computer-readable storage medium of the present invention can be any form of storage medium that can be read by a processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned detection method can be implemented.

[0110] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments may be envisioned within the scope of the invention thus described. In addition, it should be noted that the language used in this specification is selected primarily for readability and teaching purposes, rather than for explaining or defining the subject matter of the present invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the present invention is illustrative, not restrictive, with respect to the scope of the present invention, which is defined by the appended claims.

Claims

1. An autonomous decision-making system for intelligent connected vehicles based on reinforcement learning, characterized in that: The system includes: a data acquisition module, an autonomous driving decision module and a control module; The data acquisition module is used to pre-process the collected data of the vehicle-mounted sensor and the message information of the network-connected device to obtain the pre-processed collected data and the pre-processed message information, and send them to the autonomous driving decision module; The autonomous driving decision module is used to convert the pre-processed collected data into a preliminary driving decision through the DQN algorithm, and is also used to construct the pre-processed message information into a value function, and use the value function to correct the preliminary driving decision, thereby obtaining a final vehicle autonomous decision, and sending the final vehicle autonomous decision to the control module; The control module is used to control the vehicle driving state according to the final vehicle autonomous decision; The value function is: ; in is the reward function, is the penalty function; The value functions include: speed value function, lane change value function and parking value function; The speed value function is: ; in, is the speed bonus factor, The lower speed limit of the road. The upper speed limit of the road. is the current speed of the vehicle; According to the positive correlation and real-time adjustment of road vehicle density monitored by networked equipment, The value range is , vehicle speed adjustment reference for initial driving decision , positive value is acceleration, negative value is deceleration; The lane change value function is: ; in, is the lane congestion factor; According to the vehicle density adjustment of the lane where the vehicle is currently located monitored by the connected equipment, The value range is ,according to The value of covers the initial lane-changing behavior of the vehicle. When it is positive, the decision to change lanes is made. Change lane to the left. If the value is positive, the lane will change to the right, and if the value is negative, the initial decision result will be maintained; The parking value function is: ; in, is the road condition factor; The connected equipment monitors whether there are any vehicles driving illegally or in accidents in the lane where the vehicle is currently located, and makes real-time adjustments. The value range is ,when When the value of is negative, it overrides the result of the preliminary decision and finally decides that the vehicle should stop; if If it is greater than zero, it is also necessary to The value is used to slow down by percentage; The final vehicle autonomous decision is based on three value functions , as well as A composite decision is obtained, with the highest priority being , when the decision vehicle stops, all other vehicle control decisions are canceled.

2. The intelligent connected vehicle autonomous decision-making system based on reinforcement learning according to claim 1, characterized in that: The specific steps for preprocessing the collected data of the vehicle-mounted sensor are: The data acquisition module obtains the collected data of the vehicle-mounted sensor through the RTSP video stream address; Decoding the collected data of the vehicle-mounted sensor into a single-frame image in a unified RGB format; The single frame image is subjected to color space conversion and image filtering and denoising processing.

3. The intelligent connected vehicle autonomous decision-making system based on reinforcement learning according to claim 1, characterized in that: The specific steps for preprocessing the message information of the networked device are: The data acquisition module receives message information from networked devices through V2X technology; The data acquisition module reads the key fields of the message information of the networked device and saves it in the local environment.

4. The intelligent connected vehicle autonomous decision-making system based on reinforcement learning according to claim 1, characterized in that: The DQN algorithm is specifically as follows: S1. Random initialization state, autonomous decision-making system has The probability of selecting an action through Q-Network , The probability of performing random actions; S2. After the action is completed, the state , vehicle information , reward value and the next state Save to experience pool; S3. Repeat S1 and S2 until the experience value is full, and use the experience pool to update the target Q-Network; S4. After the target Q-Network converges, the preprocessed collected data is used as input to obtain the preliminary driving decision.

5. A method for autonomous decision-making of intelligent connected vehicles based on reinforcement learning, characterized in that: The method comprises: a data collection step, an autonomous driving decision step and a control step; The data collection step is used to pre-process the collected data of the vehicle-mounted sensor and the message information of the networked device to obtain the pre-processed collected data and the pre-processed message information; The autonomous driving decision step is used to convert the pre-processed collected data into a preliminary driving decision through the DQN algorithm, and is also used to construct the pre-processed message information into a value function, and use the value function to correct the preliminary driving decision, thereby obtaining a final vehicle autonomous decision; The control step is used to control the vehicle driving state according to the final vehicle autonomous decision; The value function is: ; in is the reward function, is the penalty function; The value functions include: speed value function, lane change value function and parking value function; The speed value function is: ; in, is the speed bonus factor, The lower speed limit of the road. The upper speed limit of the road. is the current speed of the vehicle; According to the positive correlation and real-time adjustment of road vehicle density monitored by networked equipment, The value range is Vehicle speed adjustment reference for preliminary driving decisions , positive value is acceleration, negative value is deceleration; The lane change value function is: ; in, is the lane congestion factor; According to the vehicle density adjustment of the lane where the vehicle is currently located monitored by the connected equipment, The value range is according to The value of covers the initial lane-changing behavior of the vehicle. When it is positive, the decision to change lanes is made. Change lane to the left. If the value is positive, the lane will change to the right, and if the value is negative, the initial decision result will be maintained; The parking value function is: ; in, is the road condition factor; The connected equipment monitors whether there are any vehicles driving illegally or in accidents in the lane where the vehicle is currently located, and makes real-time adjustments. The value range is ,when When the value of is negative, it overrides the result of the preliminary decision and finally decides that the vehicle should stop; if If it is greater than zero, it is also necessary to The value is used to slow down by percentage; The final vehicle autonomous decision is based on three value functions , as well as A composite decision is obtained, with the highest priority being , when the decision vehicle stops, all other vehicle control decisions are canceled.

6. A computer device comprising a memory and a processor, characterized in that A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the autonomous decision-making method according to claim 5.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the autonomous decision-making method of claim 5 is implemented.

Citation Information

Patent Citations

  • Urban traffic-based scheduling method for automatically driving special vehicles

    CN110956837A

  • Vehicle-mounted automatic driving algorithm optimization method, system and device based on road side unit

    CN117622221A

  • Decision planning method for automatic driving vehicle in urban traffic scene based on reinforcement learning

    CN118396034A