Vehicle lane changing decision method and device, electronic equipment and storage medium

By combining ensemble learning and reinforcement learning with various environmental and vehicle status data, accurate lane-changing decision commands are generated, solving the problem of misjudgment in autonomous driving under the influence of single-modal data and improving the safety and comfort of autonomous driving.

CN119796250BActive Publication Date: 2025-12-30CHONGQING CHANGAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510068262.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-12-30
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Existing lane-changing decision-making methods for autonomous vehicles mainly rely on single-modal input data, which are easily affected by factors such as weather and lighting, leading to decreased system performance or misjudgments.

Method used

An ensemble learning approach is adopted, which inputs multiple weak learners with various types of environmental data (such as images and point cloud data) to generate a set of environmental and state recognition results. Combined with vehicle state data, the results are input into a reinforcement learning model to generate lane-changing decision instructions.

Benefits of technology

It improves the accuracy of environmental recognition and the reliability of lane-changing decisions, reduces misjudgments caused by single-modal data, and enhances the safety and comfort of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119796250B_ABST
    Figure CN119796250B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of vehicle lane change decision method, device, electronic equipment and storage medium, at least two types of environment data collected for the surrounding environment of target vehicle are acquired, and, the vehicle state data of the target vehicle is acquired;For each type of environment data, the environment data is input into at least two corresponding weak learners, and the corresponding environment recognition result set of the type of environment data is obtained;The vehicle state data is input into at least two corresponding weak learners, and the corresponding state recognition result set of the target vehicle is obtained;Input data is obtained from each of the environment recognition result set and the state recognition result set, and the input data is input into reinforcement learning model, and the corresponding longitudinal acceleration and steering angle change rate are obtained;According to the longitudinal acceleration and the steering angle change rate, the lane change decision instruction corresponding to the target vehicle is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and specifically to a vehicle lane-changing decision-making method, device, electronic device, and storage medium. Background Technology

[0002] Autonomous vehicles (also known as self-driving automobiles) are intelligent vehicles that achieve driverless operation through computer systems. They rely on the collaborative efforts of artificial intelligence, computer vision, radar, monitoring devices, and global positioning systems to enable computers to operate motor vehicles automatically and safely without any active human intervention.

[0003] Lane change decision-making is a crucial function in autonomous driving systems. An effective lane change decision-making system can not only improve vehicle driving safety but also enhance driving smoothness and comfort. Currently, most lane change decision-making methods rely on single-modal input data (e.g., camera images, LiDAR data, GPS information, etc.) for lane change decisions.

[0004] However, this method of making lane-changing decisions based on a single modality of input data is easily affected by factors such as weather and lighting, which can lead to a decrease in system performance or misjudgments. Summary of the Invention

[0005] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this application provides a vehicle lane-changing decision-making method, device, electronic device and storage medium.

[0006] Firstly, this application provides a vehicle lane-changing decision-making method, including:

[0007] Acquire at least two types of environmental data collected from the surrounding environment of the target vehicle, and acquire vehicle status data of the target vehicle;

[0008] For each type of environmental data, the environmental data is input into at least two corresponding weak learners to obtain a set of environmental recognition results corresponding to the type of environmental data;

[0009] The vehicle state data is input into at least two corresponding weak learners to obtain a set of state recognition results corresponding to the target vehicle.

[0010] Input data is obtained from each of the environmental recognition result sets and the state recognition result sets, and the input data is input into the reinforcement learning model to obtain the corresponding longitudinal acceleration and steering angle change rate;

[0011] Based on the longitudinal acceleration and the rate of change of steering angle, a lane-changing decision command corresponding to the target vehicle is generated.

[0012] In one possible implementation, obtaining input data from each of the environment identification result sets and the state identification result sets includes:

[0013] Determine at least two data acquisition times;

[0014] For each set of environmental identification results, environmental identification data collected at each data acquisition time is obtained from the set of environmental identification results, and at least two sets of environmental identification data are combined in chronological order to obtain environmental time-series identification data.

[0015] From the set of state recognition results, obtain the state recognition data collected at each data acquisition time, and combine at least two of the state recognition data in chronological order to obtain state time sequence recognition data;

[0016] At least two of the environmental timing identification data and the state timing identification data are used as the input data.

[0017] In one possible implementation, determining at least two data acquisition moments includes:

[0018] Obtain the target time, preset time step, and vehicle lane change time;

[0019] The cutoff time is determined based on the target time and the vehicle lane-changing time.

[0020] Starting from the target time, an intermediate time is determined at each preset time step until the cutoff time is reached;

[0021] The target time and all the intermediate times are taken as the data acquisition times.

[0022] In one possible implementation, generating the lane-changing decision command corresponding to the target vehicle based on the longitudinal acceleration and the rate of change of the steering angle includes:

[0023] Obtain the current vehicle speed and current steering angle of the target vehicle;

[0024] Determine the start and end times at least two of the data acquisition times;

[0025] The time difference between the start time and the end time is determined as the target duration;

[0026] The target vehicle speed is obtained by setting and calculating the current vehicle speed, the longitudinal acceleration, and the target duration.

[0027] The target steering angle is obtained by setting and calculating the current steering angle, the rate of change of the steering angle, and the target duration.

[0028] Based on the longitudinal acceleration, the target vehicle speed, the rate of change of steering angle, and the target steering angle, a lane-changing decision instruction corresponding to the target vehicle is generated.

[0029] In one possible implementation, generating the lane-changing decision command corresponding to the target vehicle based on the longitudinal acceleration, the target vehicle speed, the rate of change of the steering angle, and the target steering angle includes:

[0030] Obtain the preset correction conditions;

[0031] Based on the correction conditions, the longitudinal acceleration, the target vehicle speed, the rate of change of the steering angle, and the target steering angle are corrected to obtain the corresponding corrected longitudinal acceleration, corrected vehicle speed, corrected rate of change of the steering angle, and corrected steering angle.

[0032] Based on the corrected longitudinal acceleration, the corrected vehicle speed, the corrected steering angle change rate, and the corrected steering angle, a lane-changing decision command corresponding to the target vehicle is generated.

[0033] In one possible implementation, the environmental data includes image data and point cloud data, and the acquisition of at least two types of environmental data collected regarding the environment surrounding the target vehicle includes:

[0034] Image data is acquired by multiple image acquisition devices installed on the target vehicle, and point cloud data is acquired by multiple radars installed on the target vehicle.

[0035] In one possible implementation, the method further includes:

[0036] During the training of the reinforcement learning model, lane change prediction data corresponding to the output of the reinforcement learning model is obtained. The lane change prediction data includes the comfort level of the target vehicle after changing lanes according to the output, the lane distance between the target vehicle and the center line of the lane to be changed, the safe distance between the target vehicle and the vehicle in front, and the vehicle collision situation of the target vehicle.

[0037] The reward value of the lane change prediction data is determined according to the reward function corresponding to the reinforcement learning model. The reward value includes a comfort reward value corresponding to the comfort level, a lane distance reward value corresponding to the lane distance, a safety distance reward value corresponding to the safety distance, and a collision reward value corresponding to the vehicle collision situation.

[0038] The model parameters of the reinforcement learning model are adjusted according to the reward value until the reward value reaches the target reward value, at which point the model training ends.

[0039] Secondly, this application provides a vehicle lane-changing decision-making device, comprising:

[0040] The data acquisition module is used to acquire at least two types of environmental data collected from the surrounding environment of the target vehicle, and to acquire the vehicle status data of the target vehicle.

[0041] The first input module is used to input the environmental data into at least two corresponding weak learners for each type of environmental data, so as to obtain a set of environmental recognition results corresponding to the type of environmental data.

[0042] The second input module is used to input the vehicle state data into at least two corresponding weak learners to obtain a set of state recognition results corresponding to the target vehicle.

[0043] The third input module is used to obtain input data from each of the environment recognition result sets and the state recognition result sets, and input the input data into the reinforcement learning model to obtain the corresponding longitudinal acceleration and steering angle change rate;

[0044] The instruction generation module is used to generate lane change decision instructions corresponding to the target vehicle based on the longitudinal acceleration and the steering angle change rate.

[0045] In one possible implementation, the third input module is specifically used for:

[0046] Determine at least two data acquisition times;

[0047] For each set of environmental identification results, environmental identification data collected at each data acquisition time is obtained from the set of environmental identification results, and at least two sets of environmental identification data are combined in chronological order to obtain environmental time-series identification data.

[0048] From the set of state recognition results, obtain the state recognition data collected at each data acquisition time, and combine at least two of the state recognition data in chronological order to obtain state time sequence recognition data;

[0049] At least two of the environmental timing identification data and the state timing identification data are used as the input data.

[0050] In one possible implementation, the third input module is further configured to:

[0051] Obtain the target time, preset time step, and vehicle lane change time;

[0052] The cutoff time is determined based on the target time and the vehicle lane-changing time.

[0053] Starting from the target time, an intermediate time is determined at each preset time step until the cutoff time is reached;

[0054] The target time and all the intermediate times are taken as the data acquisition times.

[0055] In one possible implementation, the instruction generation module is specifically used for:

[0056] Obtain the current vehicle speed and current steering angle of the target vehicle;

[0057] Determine the start and end times at least two of the data acquisition times;

[0058] The time difference between the start time and the end time is determined as the target duration;

[0059] The target vehicle speed is obtained by setting and calculating the current vehicle speed, the longitudinal acceleration, and the target duration.

[0060] The target steering angle is obtained by setting and calculating the current steering angle, the rate of change of the steering angle, and the target duration.

[0061] Based on the longitudinal acceleration, the target vehicle speed, the rate of change of steering angle, and the target steering angle, a lane-changing decision instruction corresponding to the target vehicle is generated.

[0062] In one possible implementation, the instruction generation module is further configured to:

[0063] Obtain the preset correction conditions;

[0064] Based on the correction conditions, the longitudinal acceleration, the target vehicle speed, the rate of change of the steering angle, and the target steering angle are corrected to obtain the corresponding corrected longitudinal acceleration, corrected vehicle speed, corrected rate of change of the steering angle, and corrected steering angle.

[0065] Based on the corrected longitudinal acceleration, the corrected vehicle speed, the corrected steering angle change rate, and the corrected steering angle, a lane-changing decision command corresponding to the target vehicle is generated.

[0066] In one possible implementation, the environmental data includes image data and point cloud data, and the data acquisition module is specifically used for:

[0067] Image data is acquired by multiple image acquisition devices installed on the target vehicle, and point cloud data is acquired by multiple radars installed on the target vehicle.

[0068] In one possible implementation, the apparatus further includes a model training module for:

[0069] During the training of the reinforcement learning model, lane change prediction data corresponding to the output of the reinforcement learning model is obtained. The lane change prediction data includes the comfort level of the target vehicle after changing lanes according to the output, the lane distance between the target vehicle and the center line of the lane to be changed, the safe distance between the target vehicle and the vehicle in front, and the vehicle collision situation of the target vehicle.

[0070] The reward value of the lane change prediction data is determined according to the reward function corresponding to the reinforcement learning model. The reward value includes a comfort reward value corresponding to the comfort level, a lane distance reward value corresponding to the lane distance, a safety distance reward value corresponding to the safety distance, and a collision reward value corresponding to the vehicle collision situation.

[0071] The model parameters of the reinforcement learning model are adjusted according to the reward value until the reward value reaches the target reward value, at which point the model training ends.

[0072] Thirdly, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0073] Memory, used to store computer programs;

[0074] When a processor executes a program stored in memory, it implements any of the steps described in the first aspect.

[0075] Fourthly, a computer-readable storage medium is provided, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the methods described in the first aspect.

[0076] The technical solutions provided in this application have the following advantages compared with the prior art:

[0077] The vehicle lane-changing decision-making method, apparatus, electronic device, and storage medium provided in this application first acquire at least two types of environmental data collected from the surrounding environment of the target vehicle, and acquire vehicle state data of the target vehicle. Then, for each type of environmental data, the environmental data is input into at least two corresponding weak learners to obtain an environmental recognition result set corresponding to the type of environmental data, and the vehicle state data is input into at least two corresponding weak learners to obtain a state recognition result set corresponding to the target vehicle. Next, input data is obtained from each environmental recognition result set and state recognition result set, and the input data is input into a reinforcement learning model to obtain the corresponding longitudinal acceleration and steering angle change rate. Finally, based on the longitudinal acceleration and steering angle change rate, a lane-changing decision command corresponding to the target vehicle is generated. In this application, firstly, by using at least two types of environmental data to identify the vehicle's surrounding environment, the accuracy of environmental identification can be improved. Secondly, by feeding vehicle state data and various types of environmental data into various weak learners respectively, outputting identification results, and then using the identification results of each weak learner as input to the reinforcement learning model, further extracting features from the input results, the vehicle situation and the surrounding environment situation can be better identified, thereby making better lane-changing decisions. Attached Figure Description

[0078] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0079] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0080] Figure 1 A flowchart of a vehicle lane-changing decision method provided in an embodiment of this application;

[0081] Figure 2 A flowchart illustrating another vehicle lane-changing decision-making method provided in this application embodiment;

[0082] Figure 3 A flowchart illustrating another vehicle lane-changing decision-making method provided in this application embodiment;

[0083] Figure 4 A flowchart illustrating another vehicle lane-changing decision-making method provided in this application embodiment;

[0084] Figure 5 A basic framework diagram of reinforcement learning provided for embodiments of this application;

[0085] Figure 6 This is an overall flowchart of a vehicle lane-changing decision-making method provided in an embodiment of this application;

[0086] Figure 7 This is a schematic diagram of the structure of a vehicle lane-changing decision device provided in an embodiment of this application;

[0087] Figure 8 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0088] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0089] To facilitate understanding of the embodiments of this application, the following will provide further explanation and description with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application.

[0090] See Figure 1 This is a flowchart illustrating an embodiment of a vehicle lane-changing decision-making method provided in this application. Figure 1 As shown, the process may include the following steps:

[0091] Step 101: Obtain at least two types of environmental data collected from the surrounding environment of the target vehicle, and obtain vehicle status data of the target vehicle.

[0092] The aforementioned environmental data refers to data collected regarding the environment surrounding the target vehicle, including image data and point cloud data.

[0093] Specifically, acquiring at least two types of environmental data surrounding the target vehicle may include the following steps: acquiring image data using multiple image acquisition devices installed on the target vehicle, and acquiring point cloud data using multiple radars installed on the target vehicle.

[0094] In the application, image acquisition devices (such as cameras, webcams, etc.) and radar are installed at various locations on the target vehicle. The image acquisition devices can collect image data (such as pictures, videos, etc.) of the environment around the vehicle, and the radar can collect radar signals of the environment around the vehicle, which are finally presented as point cloud data.

[0095] Since image data lacks depth information and point cloud data is always sparse, using only one type of data as input for subsequent algorithms results in low accuracy. This solution integrates image-type data and point cloud-type data, which can offset the shortcomings of a single type of data and thus improve the accuracy of subsequent algorithm recognition.

[0096] Vehicle status data refers to numerical information about the current vehicle status, such as total CAN (Controller Area Network) information, speed, wheel speed, angular velocity, and distance to the vehicle in front. In applications, this numerical information is collected by sensors installed at various locations on the vehicle body.

[0097] Step 102: For each type of environmental data, input the environmental data into at least two corresponding weak learners to obtain a set of environmental recognition results corresponding to the type of environmental data.

[0098] The weak learners corresponding to the image data mentioned above refer to image-based object recognition algorithms, such as CNN (Convolutional Neural Networks), SVM (Support Vector Machine), decision trees, and K-nearest neighbors. The weak learners corresponding to the point cloud data mentioned above refer to point cloud-based object recognition algorithms, such as VoxelNet, PointNet, PointRCNN, and octrees.

[0099] In this embodiment, an ensemble learning approach is adopted, using some weak learners to perform preliminary feature extraction on the collected environmental data. Specifically, image data is input into multiple image-based target recognition algorithms, and each image-based target recognition algorithm outputs a corresponding environmental recognition result, thus obtaining a set of multiple environmental recognition results. Similarly, point cloud data is input into multiple point cloud-based target recognition algorithms, and each point cloud-based target recognition algorithm outputs a corresponding environmental recognition result, thus obtaining a set of multiple environmental recognition results.

[0100] It should be noted that, in the application, the algorithms corresponding to each type of environmental data have the following three requirements: 1) Different algorithms have the same or similar output labels; 2) The algorithm is applicable to the data to be labeled; 3) Each algorithm has a different network architecture or recognition differences. For example, an algorithm may be more accurate in recognizing a certain category, but may not perform well in other categories or scenarios. Other algorithms should have the ability to compensate for such "weaknesses" to a certain extent. In this way, the requirement of ensemble learning for each algorithm to be "good but different" is met.

[0101] Step 103: Input the vehicle state data into at least two corresponding weak learners to obtain the state recognition result set corresponding to the target vehicle.

[0102] In this embodiment of the application, the above-mentioned ensemble learning idea is also adopted. The vehicle state data is input into at least two corresponding weak learners, and each weak learner outputs a corresponding state recognition result, thereby obtaining a set of multiple state recognition results.

[0103] Step 104: Obtain input data from each of the environment recognition result sets and the state recognition result sets, and input the input data into the reinforcement learning model to obtain the corresponding longitudinal acceleration and steering angle change rate.

[0104] Step 105: Generate lane change decision instructions corresponding to the target vehicle based on the longitudinal acceleration and the steering angle change rate.

[0105] For ease of understanding, steps 104 and 105 will be explained uniformly below:

[0106] The reinforcement learning model described above is based on a model trained using reinforcement learning, a learning method that learns how to maximize rewards through interaction with an environment. Reinforcement learning involves an agent (whose core is a neural network), an environment, and a reward function. The agent performs actions in the environment and receives rewards according to the reward function. The agent's goal is to learn a policy that maximizes cumulative rewards in the environment. The core idea of ​​reinforcement learning is to gradually learn how to maximize rewards in different states through exploration and exploitation.

[0107] Input data refers to data obtained from the same time or time period in various recognition result sets.

[0108] Based on this, in this embodiment, the results identified by each weak learner are used as input to the reinforcement learning model, enabling the model to further extract features from the input data. This allows for better identification of vehicle and environmental conditions, and the output of more accurate decision information (i.e., longitudinal acceleration and steering angle change rate) based on the identified conditions. Finally, based on the longitudinal acceleration and steering angle change rate, a lane-changing decision command corresponding to the target vehicle is generated. This command controls the target vehicle to perform a lane-changing operation. Specifically, the output is converted into operational information and connected to the vehicle's operating system, allowing direct control of the vehicle's direction, throttle, and brakes to achieve assisted driving.

[0109] The technical solution provided in this application first acquires at least two types of environmental data collected from the surrounding environment of the target vehicle, and acquires vehicle state data of the target vehicle. Then, for each type of environmental data, the environmental data is input into at least two corresponding weak learners to obtain an environmental recognition result set corresponding to the type of environmental data, and the vehicle state data is input into at least two corresponding weak learners to obtain a state recognition result set corresponding to the target vehicle. Next, input data is obtained from each environmental recognition result set and state recognition result set, and the input data is input into a reinforcement learning model to obtain the corresponding longitudinal acceleration and steering angle change rate. Finally, based on the longitudinal acceleration and steering angle change rate, a lane-changing decision command corresponding to the target vehicle is generated. In this application, firstly, by using at least two types of environmental data to identify the vehicle's surrounding environment, the accuracy of environmental identification can be improved. Secondly, by feeding vehicle state data and various types of environmental data into various weak learners respectively, outputting identification results, and then using the identification results of each weak learner as input to the reinforcement learning model, further extracting features from the input results, the vehicle situation and the surrounding environment situation can be better identified, thereby making better lane-changing decisions.

[0110] See Figure 2 This is a flowchart illustrating an embodiment of another vehicle lane-changing decision-making method provided in this application. Figure 2 The process shown above Figure 1 Based on the illustrated process, this section describes how to obtain input data from each of the environment identification result sets and the state identification result sets. For example... Figure 2 As shown, the process may include the following steps:

[0111] Step 201: Determine at least two data acquisition times.

[0112] Data acquisition time refers to the moment when environmental data and vehicle status data are acquired. In applications, the time and frequency at which corresponding data are acquired through image acquisition devices, radar, and sensors are the same; that is, the corresponding data is acquired at the same time every preset time interval.

[0113] In this embodiment of the application, step 201 may specifically include the following steps: obtaining a target time, a preset time step, and a vehicle lane-changing time; determining a cutoff time based on the target time and the vehicle lane-changing time; determining an intermediate time every preset time step starting from the target time until the cutoff time is reached; and using the target time and all the intermediate times as data acquisition times.

[0114] The target time can be any time when data is collected; the preset time step refers to the time interval between data collection by the image acquisition device, radar and sensors; the vehicle lane changing time refers to the time it takes for a vehicle to change from the current lane to other lanes (e.g., two to three seconds).

[0115] In this implementation, firstly, the cutoff time is obtained by subtracting the vehicle lane-changing time from the target time. Then, an intermediate time is obtained by subtracting a preset time step from the target time. Another intermediate time is obtained by subtracting two preset time steps from the target time, and so on, until the cutoff time is reached. Thus, multiple data acquisition times are obtained.

[0116] Step 202: For each set of environmental identification results, obtain the environmental identification data collected at each data acquisition time from the set of environmental identification results, and combine at least two of the environmental identification data in chronological order to obtain environmental time sequence identification data.

[0117] Step 203: Obtain the state identification data collected at each data acquisition time from the state identification result set, and combine at least two of the state identification data in chronological order to obtain state time sequence identification data.

[0118] Step 204: Use at least two of the environmental timing identification data and the state timing identification data as the input data.

[0119] For ease of understanding, steps 202-204 will be explained uniformly below:

[0120] In this embodiment of the application, environmental identification data collected at each data acquisition moment is obtained from each environmental identification result set, and sorted in chronological order to obtain environmental temporal identification data; and state identification data collected at each data acquisition moment is obtained from the state identification result set, and sorted in chronological order to obtain state temporal identification data.

[0121] Specifically, the neural network output tensor of the nth weak learner at time t (i.e., the target time) is defined as C. t n This value reflects the recognition result after the learner extracts the features. Considering the temporal nature of the scene, a preset time step of τ is defined to incorporate historical states.

[0122] Therefore, the output of each learner at time t is as follows:

[0123]

[0124] Let there be a 2τ historical state, where the data acquisition time includes t, t-τ (the acquisition time before t), and t-2τ (the two acquisition times before t). Then the state st at time t is defined as follows:

[0125] s t =[T t ,T t-τ ,T t-2τ ]

[0126] The state st at time t in the above formula consists of two parts. Let T be the neural network output tensor of the nth weak learner at time t, used to describe the result recognized by that learner; t-τ ,T t-2τ It is a record of history, used to introduce the characteristics of temporal changes and extract historical experience.

[0127] pass Figure 2 The process shown can take into account the temporal nature of the lane-changing scenario by incorporating historical states, which helps vehicles better recognize changes in the surrounding environment and improves the accuracy of decision-making.

[0128] See Figure 3 This is a flowchart illustrating an embodiment of another vehicle lane-changing decision-making method provided in this application. Figure 3 The process shown above Figure 1 Based on the illustrated process, this section describes how to generate lane-changing decision instructions for the target vehicle according to the longitudinal acceleration and the rate of change of the steering angle. For example... Figure 3 As shown, the process may include the following steps:

[0129] Step 301: Obtain the current vehicle speed and current steering angle of the target vehicle;

[0130] Step 302: Determine the start time and end time in at least two of the data acquisition times;

[0131] Step 303: Determine the time difference between the start time and the end time as the target duration;

[0132] Step 304: Perform setting calculations on the current vehicle speed, the longitudinal acceleration, and the target duration to obtain the target vehicle speed;

[0133] Step 305: Perform setting calculations on the current steering angle, the steering angle change rate, and the target duration to obtain the target steering angle;

[0134] Step 306: Generate a lane-changing decision command corresponding to the target vehicle based on the longitudinal acceleration, the target vehicle speed, the rate of change of steering angle, and the target steering angle.

[0135] For ease of understanding, steps 301-306 will be explained uniformly below:

[0136] For lane-changing scenarios, the reinforcement learning model primarily learns how to adjust speed and steering angle to reach the destination. Therefore, the output of the action is a two-dimensional vehicle control signal: longitudinal vehicle acceleration and the rate of change of steering angle. During lane changes, longitudinal acceleration mainly controls the vehicle's throttle and brakes, while the rate of change of steering angle controls the steering wheel angle. The longitudinal vehicle acceleration av ranges from [-3m / s², 3m / s²]. Specifically, values ​​in the range [-3m / s², 0] represent deceleration, and values ​​in the range [0, 3m / s²] represent acceleration. Similarly, the steering angle mainly controls the magnitude of the lane change, equivalent to the steering wheel angle, and the rate of change of steering angle ay ranges from [-3° / s, 3° / s]. Specifically, values ​​in the range [-3° / s, 0] represent left turns, and values ​​in the range [0, 3° / s] represent right turns. The steering angle should be constrained by speed. The faster the speed, the smaller the steering angle should be when changing lanes. When the speed is slower, the steering angle can be larger.

[0137] In this embodiment of the application, the action space of reinforcement learning is designed as a continuous action space, and the DDPG algorithm is selected. Therefore, the action spaces av and ay are defined as follows: av∈[-3m / s2,3m / s2], ay∈[-3° / s,3° / s];

[0138] Therefore, the formula for velocity transformation is: V t+τ =V t +a v *τ;

[0139] Among them, V t+τ For the target vehicle speed, V t Given the current vehicle speed, a v Let τ be the longitudinal acceleration and τ be the target duration.

[0140] The formula for transforming the direction angle is: θ t+τ =θ t +a y *τ;

[0141] Where, θ t+τ Let θ be the target steering angle. t a is the current steering angle. y τ is the rate of change of steering angle, and τ is the target duration.

[0142] That is, in the embodiments of this application, firstly, the target duration is obtained by calculating the time difference between the start time and the end time. Then, the target vehicle speed is obtained by setting the current vehicle speed, longitudinal acceleration and target duration using the speed transformation formula. The target steering angle is obtained by setting the current steering angle, steering angle change rate and target duration using the steering angle transformation formula. Finally, a lane change decision instruction corresponding to the target vehicle is generated based on the longitudinal acceleration, target vehicle speed, steering angle change rate and target steering angle to guide the target vehicle to adjust the current vehicle speed to the target vehicle speed according to the longitudinal acceleration and adjust the current steering angle to the target steering angle according to the steering angle change rate.

[0143] In applications, directly generating lane-changing decision instructions based on the output of a reinforcement learning model can easily lead to illegal actions and create dangers. For example, outputting a large steering angle at a high vehicle speed can easily cause the vehicle to roll over; or, outputting acceleration can cause the vehicle to travel too fast, exceeding the speed of the vehicle in front in the target lane after the lane change, which can easily lead to a rear-end collision. Such output actions are dangerous and are judged as illegal. Therefore, in another embodiment of this application, generating the lane-changing decision instructions corresponding to the target vehicle based on the longitudinal acceleration, the target vehicle speed, the rate of change of the steering angle, and the target steering angle may further include the following steps:

[0144] Obtain preset correction conditions; correct the longitudinal acceleration, the target vehicle speed, the steering angle change rate, and the target steering angle according to the correction conditions to obtain the corresponding corrected longitudinal acceleration, corrected vehicle speed, corrected steering angle change rate, and corrected steering angle; generate the lane change decision command corresponding to the target vehicle based on the corrected longitudinal acceleration, corrected vehicle speed, corrected steering angle change rate, and corrected steering angle.

[0145] The above correction conditions can be set by the user according to actual needs, such as the steering angle range corresponding to the vehicle speed, the target vehicle speed being less than the speed of the vehicle in front after changing lanes, etc.

[0146] As described above, in this embodiment, the longitudinal acceleration, target vehicle speed, steering angle change rate, and target steering angle are corrected according to the correction conditions. Based on the corrected longitudinal acceleration, corrected vehicle speed, corrected steering angle change rate, and corrected steering angle, the final lane-changing decision command is generated. This avoids generating illegal actions and improves the safety of autonomous driving.

[0147] pass Figure 3The process shown, for continuous action space, can calculate the target vehicle speed and target steering angle at the end of the lane change action based on the target vehicle's current speed, current steering angle, longitudinal acceleration, steering angle change rate, and the time required for lane change (i.e., target time). Thus, based on the longitudinal acceleration, target vehicle speed, steering angle change rate, and target steering angle, the corresponding lane change decision command for the target vehicle is generated.

[0148] See Figure 4 This is a flowchart illustrating another embodiment of the vehicle lane-changing decision-making method provided in this application. Figure 4 As shown, the process may include the following steps:

[0149] Step 401: During the training of the reinforcement learning model, obtain lane change prediction data corresponding to the output result of the reinforcement learning model. The lane change prediction data includes the comfort level of the target vehicle after changing lanes according to the output result, the lane distance between the target vehicle and the center line of the lane to be changed, the safe distance between the target vehicle and the vehicle in front, and the vehicle collision situation of the target vehicle.

[0150] Step 402: Determine the reward value of the lane change prediction data according to the reward function corresponding to the reinforcement learning model, wherein the reward value includes a comfort reward value corresponding to the comfort level, a lane distance reward value corresponding to the lane distance, a safety distance reward value corresponding to the safety distance, and a collision reward value corresponding to the vehicle collision situation.

[0151] Step 403: Adjust the model parameters of the reinforcement learning model according to the reward value until the reward value reaches the target reward value, and the model training ends.

[0152] For ease of understanding, steps 401-403 will be explained uniformly below:

[0153] In the embodiments of this application, such as Figure 5As shown, a reinforcement learning model includes an agent, an environment, and a reward function. The agent performs actions in the environment and receives rewards according to the reward function. The agent's goal is to learn a policy that maximizes the cumulative reward in the environment. The core idea of ​​reinforcement learning is to learn how to maximize rewards in different states through exploration and utilization. The reward function consists of four parts: first, lane-changing reward / penalty, which means the closer the center of the target vehicle is to the centerline of the target lane (i.e., the lane line after the lane change), the better; second, distance reward / penalty, maintaining a safe distance from the vehicle in front throughout the process; third, comfort reward / penalty, where comfort typically depends on the rate of change of longitudinal acceleration and steering angle, and should be maintained as smoothly as possible during lane changes; and fourth, collision penalty, where the vehicle will be severely penalized in the event of a collision.

[0154] 1) Lane change bonus formula:

[0155] r l =-abs[L t -Y]

[0156] Where abs represents taking the absolute value, L t This represents the lateral coordinate position of the target vehicle's center point (the center point of the diagonal of a rectangle considered as the entire vehicle body) at time t. Y represents the lateral coordinate of the target lane centerline. The farther the vehicle is from the target lane centerline, the smaller this value and the more severe the penalty; the closer the vehicle is to the target lane centerline, the larger this value.

[0157] 2) Distance reward / penalty formula:

[0158]

[0159] Where r d For distance-based rewards, D is the distance to the vehicle in front. A reward is given if the distance is greater than the safe distance, and a penalty is imposed if the distance is less than the safe distance. The safe distance is generally related to the speed of the vehicle in front, and a maximum safe distance is defined as follows:

[0160]

[0161] Among them, V t t' is the current vehicle speed, t' is the person's reaction time, u is the speed of the vehicle in front, and d is the speed of the vehicle in front. s The basic safety threshold is generally set at 5 meters. It can be seen that the safe distance is directly proportional to speed and inversely proportional to the speed of the vehicle in front. The key point is that when a vehicle changing lanes crosses the center line of the lane, the vehicle in front of it will jump to become the vehicle in front of the target lane.

[0162] 3) Comfort reward and penalty:

[0163] Comfort level sIt depends on the rate of change of longitudinal acceleration and steering angle.

[0164] 4) Collision Rewards and Penalties:

[0165] Where, r c The collision penalty is a binary reward, for example, when no collision occurs, r... c When a collision occurs, r is 0. c It is -100.

[0166] Therefore, the entire reward and punishment function formula is as follows:

[0167] r t =w l r l +w d r d +w s r s +w c r c

[0168] Among them, w l ,w d ,w s ,w c , where w is the penalty coefficient, used to adjust the proportion of each reward / penalty in the reward function. In application, w c r c The reward function is two or three orders of magnitude smaller than other reward and penalty items, thus ensuring safety.

[0169] pass Figure 4 The process shown can train a reinforcement learning model based on a reward function, enabling the model to gradually learn how to maximize rewards under different states, thereby improving the model's processing power and the accuracy of its output.

[0170] See Figure 6 This is an overall flowchart of a vehicle lane-changing decision-making method provided in an embodiment of this application. Figure 6 As shown, the process may include the following steps:

[0171] Upon receiving a lane-changing initiation command (either from a driver requesting a lane change, a navigation-guided route requesting a lane change, or an autonomous driving path planning system), the vehicle lane-changing decision-making method of this application is activated to begin making intelligent lane-changing action decisions.

[0172] Specifically, firstly, cameras, radar, and sensors on the vehicle acquire camera images, radar point cloud data, and signal values ​​representing the vehicle's own state to capture images of the surrounding environment. Then, these data are fed into various weak learners, which output corresponding target detection and recognition results. Temporal features are extracted from the learners' results and used as input to a reinforcement learning agent, allowing the agent to further extract features from the input. The agent interacts with the environment to train its model, thereby better recognizing environmental conditions and outputting corresponding actions. Finally, corrective actions from a professional driver correct the agent's output, resulting in the final lane-changing decision. Furthermore, the corrected actions can be input into a simulation environment and calculated using a reward function to further optimize the reinforcement learning agent.

[0173] In this application, firstly, by using at least two types of environmental data to identify the vehicle's surrounding environment, the accuracy of environmental identification can be improved. Secondly, by feeding vehicle state data and various types of environmental data into various weak learners respectively, outputting identification results, and then using the identification results of each weak learner as input to the reinforcement learning model, further extracting features from the input results, the vehicle situation and the surrounding environment situation can be better identified, thereby making better lane-changing decisions.

[0174] Based on the same technical concept, embodiments of this application also provide a vehicle lane-changing decision device, such as... Figure 7 As shown, the device includes:

[0175] The data acquisition module 71 is used to acquire at least two types of environmental data collected from the surrounding environment of the target vehicle, and to acquire the vehicle status data of the target vehicle.

[0176] The first input module 72 is used to input the environmental data into at least two corresponding weak learners for each type of environmental data to obtain a set of environmental recognition results corresponding to the type of environmental data;

[0177] The second input module 73 is used to input the vehicle state data into at least two corresponding weak learners to obtain a set of state recognition results corresponding to the target vehicle.

[0178] The third input module 74 is used to obtain input data from each of the environment recognition result sets and the state recognition result sets, and input the input data into the reinforcement learning model to obtain the corresponding longitudinal acceleration and steering angle change rate;

[0179] The instruction generation module 75 is used to generate a lane-changing decision instruction corresponding to the target vehicle based on the longitudinal acceleration and the steering angle change rate.

[0180] In one possible implementation, the third input module is specifically used for:

[0181] Determine at least two data acquisition times;

[0182] For each set of environmental identification results, environmental identification data collected at each data acquisition time is obtained from the set of environmental identification results, and at least two sets of environmental identification data are combined in chronological order to obtain environmental time-series identification data.

[0183] From the set of state recognition results, obtain the state recognition data collected at each data acquisition time, and combine at least two of the state recognition data in chronological order to obtain state time sequence recognition data;

[0184] At least two of the environmental timing identification data and the state timing identification data are used as the input data.

[0185] In one possible implementation, the third input module is further configured to:

[0186] Obtain the target time, preset time step, and vehicle lane change time;

[0187] The cutoff time is determined based on the target time and the vehicle lane-changing time.

[0188] Starting from the target time, an intermediate time is determined at each preset time step until the cutoff time is reached;

[0189] The target time and all the intermediate times are taken as the data acquisition times.

[0190] In one possible implementation, the instruction generation module is specifically used for:

[0191] Obtain the current vehicle speed and current steering angle of the target vehicle;

[0192] Determine the start and end times at least two of the data acquisition times;

[0193] The time difference between the start time and the end time is determined as the target duration;

[0194] The target vehicle speed is obtained by setting and calculating the current vehicle speed, the longitudinal acceleration, and the target duration.

[0195] The target steering angle is obtained by setting and calculating the current steering angle, the rate of change of the steering angle, and the target duration.

[0196] Based on the longitudinal acceleration, the target vehicle speed, the rate of change of steering angle, and the target steering angle, a lane-changing decision instruction corresponding to the target vehicle is generated.

[0197] In one possible implementation, the instruction generation module is further configured to:

[0198] Obtain the preset correction conditions;

[0199] Based on the correction conditions, the longitudinal acceleration, the target vehicle speed, the rate of change of the steering angle, and the target steering angle are corrected to obtain the corresponding corrected longitudinal acceleration, corrected vehicle speed, corrected rate of change of the steering angle, and corrected steering angle.

[0200] Based on the corrected longitudinal acceleration, the corrected vehicle speed, the corrected steering angle change rate, and the corrected steering angle, a lane-changing decision command corresponding to the target vehicle is generated.

[0201] In one possible implementation, the environmental data includes image data and point cloud data, and the data acquisition module is specifically used for:

[0202] Image data is acquired by multiple image acquisition devices installed on the target vehicle, and point cloud data is acquired by multiple radars installed on the target vehicle.

[0203] In one possible implementation, the apparatus further includes a model training module for:

[0204] During the training of the reinforcement learning model, lane change prediction data corresponding to the output of the reinforcement learning model is obtained. The lane change prediction data includes the comfort level of the target vehicle after changing lanes according to the output, the lane distance between the target vehicle and the center line of the lane to be changed, the safe distance between the target vehicle and the vehicle in front, and the vehicle collision situation of the target vehicle.

[0205] The reward value of the lane change prediction data is determined according to the reward function corresponding to the reinforcement learning model. The reward value includes a comfort reward value corresponding to the comfort level, a lane distance reward value corresponding to the lane distance, a safety distance reward value corresponding to the safety distance, and a collision reward value corresponding to the vehicle collision situation.

[0206] The model parameters of the reinforcement learning model are adjusted according to the reward value until the reward value reaches the target reward value, at which point the model training ends.

[0207] Based on the same technical concept, embodiments of this application also provide an electronic device, such as... Figure 8 As shown, it includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0208] Memory 113 is used to store computer programs;

[0209] When processor 111 executes a program stored in memory 113, it performs the following steps:

[0210] Acquire at least two types of environmental data collected from the surrounding environment of the target vehicle, and acquire vehicle status data of the target vehicle;

[0211] For each type of environmental data, the environmental data is input into at least two corresponding weak learners to obtain a set of environmental recognition results corresponding to the type of environmental data;

[0212] The vehicle state data is input into at least two corresponding weak learners to obtain a set of state recognition results corresponding to the target vehicle.

[0213] Input data is obtained from each of the environmental recognition result sets and the state recognition result sets, and the input data is input into the reinforcement learning model to obtain the corresponding longitudinal acceleration and steering angle change rate;

[0214] Based on the longitudinal acceleration and the rate of change of steering angle, a lane-changing decision command corresponding to the target vehicle is generated.

[0215] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0216] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0217] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0218] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0219] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described vehicle lane-changing decision methods.

[0220] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the vehicle lane-changing decision methods described above.

[0221] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0222] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A vehicle lane change decision method, characterized by, The method comprises: acquiring at least two types of environment data collected for the target vehicle's surrounding environment, and acquiring vehicle state data of the target vehicle; for each type of environment data, inputting the environment data into at least two corresponding weak learners to obtain a set of environment recognition results corresponding to the type of environment data; inputting the vehicle state data into at least two corresponding weak learners to obtain a set of state recognition results corresponding to the target vehicle; acquiring input data from each of the set of environment recognition results and the set of state recognition results, and inputting the input data into a reinforcement learning model to obtain corresponding longitudinal acceleration and steering angle change rate; generating a lane change decision instruction corresponding to the target vehicle according to the longitudinal acceleration and the steering angle change rate; wherein the generation of the lane change decision instruction corresponding to the target vehicle according to the longitudinal acceleration and the steering angle change rate comprises: determining at least two data collection time points; acquiring a current vehicle speed and a current steering angle corresponding to the target vehicle; determining a start time point and an end time point among the at least two data collection time points; determining a target time length as a time difference between the start time point and the end time point; performing a set operation on the current vehicle speed, the longitudinal acceleration, and the target time length to obtain a target vehicle speed; performing a set operation on the current steering angle, the steering angle change rate, and the target time length to obtain a target steering angle; generating a lane change decision instruction corresponding to the target vehicle according to the longitudinal acceleration, the target vehicle speed, the steering angle change rate, and the target steering angle.

2. The method of claim 1, wherein, The acquisition of input data from each of the set of environment recognition results and the set of state recognition results comprises: for each set of environment recognition results, acquiring environment recognition data collected at each of the data collection time points from the set of environment recognition results, and combining at least two of the environment recognition data in chronological order to obtain environment time series recognition data; acquiring state recognition data collected at each of the data collection time points from the set of state recognition results, and combining at least two of the state recognition data in chronological order to obtain state time series recognition data; combining at least two of the environment time series recognition data and the state time series recognition data as the input data.

3. The method of claim 2, wherein, The determination of at least two data collection time points comprises: acquiring a target time point, a preset time step, and a vehicle lane change time length; determining a cutoff time point according to the target time point and the vehicle lane change time length; determining an intermediate time point every preset time step from the target time point until reaching the cutoff time point; combining the target time point and all intermediate time points as data collection time points.

4. The method of claim 1, wherein, The generation of the lane change decision instruction corresponding to the target vehicle according to the longitudinal acceleration, the target vehicle speed, the steering angle change rate, and the target steering angle comprises: acquiring a preset correction condition; According to the modified conditions, the longitudinal acceleration, the target vehicle speed, the steering angle change rate and the target steering angle are modified to obtain corresponding modified longitudinal acceleration, modified vehicle speed, modified steering angle change rate and modified steering angle; According to the modified longitudinal acceleration, the modified vehicle speed, the modified steering angle change rate and the modified steering angle, the target vehicle corresponding lane changing decision instruction is generated.

5. The method of claim 1, wherein, The environmental data includes image data and point cloud data, and the at least two types of environmental data collected for the target vehicle surrounding environment are obtained, including: The image data is collected by a plurality of image acquisition devices arranged on the target vehicle, and the point cloud data is collected by a plurality of radars arranged on the target vehicle.

6. The method of claim 1, wherein, The method further comprises: In the process of training the reinforcement learning model, the output result of the reinforcement learning model is obtained, and the lane changing prediction data corresponding to the output result is obtained, wherein the lane changing prediction data includes the comfort degree of the target vehicle after lane changing, the lane distance between the target vehicle and the lane line center line of the lane, the safety distance between the target vehicle and the front vehicle, and the vehicle collision situation of the target vehicle; According to the reward function corresponding to the reinforcement learning model, the reward value of the lane changing prediction data is determined, wherein the reward value of the lane changing prediction data includes the comfort degree reward value corresponding to the comfort degree, the lane distance reward value corresponding to the lane distance, the safety distance reward value corresponding to the safety distance, and the collision reward value corresponding to the vehicle collision situation; According to the reward value of the lane changing prediction data, the model parameters of the reinforcement learning model are adjusted until the reward value of the lane changing prediction data reaches the target reward value, and the model training is ended.

7. A vehicle lane change decision device characterized by comprising: The device comprises: A data acquisition module is configured to acquire at least two types of environmental data collected for a target vehicle surrounding environment, and acquire vehicle state data of the target vehicle; A first input module is configured to input the environmental data into at least two corresponding weak learners for each type of environmental data, to obtain a set of environmental recognition results corresponding to the type of environmental data; A second input module is configured to input the vehicle state data into at least two corresponding weak learners to obtain a set of state recognition results corresponding to the target vehicle; A third input module is configured to acquire input data from each of the set of environmental recognition results and the set of state recognition results, and input the input data into a reinforcement learning model to obtain corresponding longitudinal acceleration and steering angle change rate; An instruction generation module is configured to generate a lane changing decision instruction corresponding to the target vehicle according to the longitudinal acceleration and the steering angle change rate; The instruction generation module is specifically configured to: Determine at least two data acquisition time points; Acquire the current vehicle speed and the current steering angle corresponding to the target vehicle; Determine a start time and an end time among the at least two data acquisition time points; Determine the time difference between the start time and the end time as a target time length; Performing a setting operation on the current vehicle speed, the longitudinal acceleration and the target time length to obtain a target vehicle speed; Performing a setting operation on the current steering angle, the steering angle change rate and the target time length to obtain a target steering angle; Generating a lane change decision instruction corresponding to the target vehicle according to the longitudinal acceleration, the target vehicle speed, the steering angle change rate and the target steering angle.

8. An electronic device, comprising: The vehicle lane change decision method comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory are in communication with each other through the communication bus. The memory is used for storing a computer program. The processor is used for executing the program stored in the memory to realize the vehicle lane change decision method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a vehicle lane change decision method program, and the vehicle lane change decision method program is executed by the processor to realize the steps of the vehicle lane change decision method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Vehicle self-learning lane changing decision-making system and method considering driving behavior characteristics

    CN113291308A

  • Automatic driving vehicle lane changing decision control method based on hierarchical reinforcement learning

    CN114013443A