A social evaluation method in driving interaction behavior based on inverse reinforcement learning

By applying the social evaluation method of reverse reinforcement learning in autonomous vehicles, the traffic accident problem caused by insufficient personification of autonomous vehicles is solved, effective social evaluation and optimization of driving interaction behavior is achieved, and traffic safety and social interaction capabilities of autonomous vehicles are improved.

CN118095063BActive Publication Date: 2025-06-06TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410116481.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-06-06
Estimated Expiration
2044-01-26

AI Technical Summary

Technical Problem

In the traffic environment of mixed human-machine traffic, the insufficient anthropomorphization of autonomous vehicles has led to frequent traffic accidents. The main reason is that the misunderstanding of human driving intentions and unclear expression of their own intentions, which leads to abnormal driving behaviors, stimulates the road anger of human drivers, and increases safety hazards.

Method used

The social evaluation method in driving interactive behavior based on inverse reinforcement learning is adopted. By constructing a game theory interactive behavior model, designing behavior variables such as acceleration and steering angle, combining individual and group benefits functions, the inverse reinforcement learning algorithm is used to optimize weight parameters, design social parameters, and evaluate the sociality of interactive behavior.

Benefits of technology

Effectively identify and evaluate drivers' social driving behaviors, help autonomous vehicles better understand the intentions of human drivers, reduce abnormal driving behaviors, improve traffic safety, and enhance the social interaction capabilities of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118095063B_ABST
    Figure CN118095063B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for evaluating sociality in driving interaction behavior based on inverse reinforcement learning, comprising the following steps: constructing an interactive behavior model based on game theory, including designing game participants, designing game behaviors, designing game benefit functions, and solving the model to obtain interactive trajectories; using an inverse reinforcement learning algorithm to identify social parameters in the interactive behavior model, and then evaluating the sociality in the interactive process. Compared with the prior art, the present invention has the advantages of being able to take into account sociality and multi-level chain action dependency, and being able to analyze complex strong interactive scenarios with large deviations between human-machine mixed driving and rule-abiding driving behaviors and high risk from a technical level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving vehicles, and in particular to a social evaluation method in driving interaction behavior based on inverse reinforcement learning. Background Art

[0002] In recent years, with the rapid development of new technologies such as information communication, Internet, big data, and artificial intelligence, autonomous vehicles (AV) are gradually replacing human-driven vehicles (HV) and becoming a new generation of vehicles for intelligent mobile space and application terminals. However, due to issues such as safety performance, public acceptance, and the large number of non-autonomous vehicles, there is still a long way to go before a 100% autonomous driving environment. In the next few decades, autonomous vehicles will share the road with human drivers.

[0003] In a mixed traffic environment with humans and machines, the lack of anthropomorphism of autonomous vehicles is the main cause of traffic accidents. At present, in the process of interaction between AV and HV, AV often performs "abnormal" driving behaviors from the perspective of human drivers due to misunderstanding of human driving intentions or unclear expression of their own intentions. This type of "abnormal" driving behavior will not only stimulate road rage among human drivers, but also make it difficult for human drivers to predict the behavior of autonomous vehicles, thus causing safety hazards. The test report released by Waymo on October 30, 2020 counted 18 traffic accidents that occurred in a total of 6.1 million miles of autonomous driving mileage from 2019 to September 2020, of which 14 accidents were caused by the driving intentions of autonomous vehicles being misunderstood by human-driven cars, resulting in rear-end collisions.

[0004] The lack of anthropomorphism in self-driving cars is mainly due to the lack of sociality. In driving interaction scenarios, human drivers will not only consider their own benefits when making decisions, but also take into account the benefits of the group. The individual's preference for the distribution of benefits directly determines the final interaction results, so the process is naturally social. Therefore, evaluating the driver's "social" driving behavior is an intuitive method to help AV correctly interpret HV driving intentions. However, the action dependency relationship between interactive objects in the interactive scenario makes the traditional behavior parameter identification method for a single individual no longer applicable. Therefore, the present invention proposes a social identification method based on inverse reinforcement learning for the social parameters in the driving interaction behavior model under multi-object action dependency, and explicitly evaluates the sociality of the two interacting parties in the interactive scenario. Summary of the invention

[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a social evaluation method in driving interaction behavior based on inverse reinforcement learning.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] The present invention provides a method for evaluating sociality in driving interaction behavior based on inverse reinforcement learning, comprising the following steps:

[0008] Step S1, constructing an interactive behavior model based on game theory, including designing game participants, where the objects are the two parties of interaction: the main vehicle and the interactive object;

[0009] Step S2, designing a game behavior, wherein the behavior includes two variables: acceleration a and steering angle h;

[0010] Step S3, designing a game benefit function, the function includes a first individual benefit, a second individual benefit and a group benefit, wherein the three benefits are linearly summed to obtain an overall benefit function, which is expressed by the following formula:

[0011] U=ω I1 ×R I1 +ω I2 ×R I2 +ω G ×R G ;

[0012] Among them, U is the overall profit function, ω I1 ,ω I2 ,ω G The first individual income R I1 、The second individual income R I2 and group payoff R g The weight parameter of

[0013] Step S4, solving the model to obtain the trajectories of the main vehicle and the interactive object respectively;

[0014] Step S5, assigning an initial value to the weight parameter;

[0015] Step S6: Generate an optimal interaction trajectory according to the current overall benefit function; during the interaction process, both parties will take the optimal action of the next frame in the optimal interaction trajectory until the interaction ends;

[0016] Step S7, extracting the feature μ(π) of the optimal interaction trajectory in the interaction process through the inverse reinforcement learning algorithm, and comparing it with the trajectory feature μ(E) in the empirical data;

[0017] Step S8: Solve the optimization problem using the maximum margin algorithm To update the weight parameters, repeatedly perform steps S6 to S7 until convergence;

[0018] Step S9: designing social parameters according to the weight parameters, and implementing evaluation of interactive behaviors through the social parameters.

[0019] Furthermore, the main vehicle includes a steering vehicle and a straight-moving vehicle, and the interactive objects correspond to the main vehicle as the straight-moving vehicle and the steering vehicle.

[0020] Furthermore, the range of the acceleration a is (-3m / s 2 ,3m / s 2 ); the range of the steering angle h is (-π / 6,π / 6).

[0021] Furthermore, the first individual benefit is expressed by the following formula:

[0022]

[0023] Among them, the first individual income R I1 is the driving process benefit, τ i represents the projection length of the actual driving trajectory of vehicle i on the reference path, It represents the total length of the actual driving trajectory of vehicle i from the starting point of the interaction to the target point projected on the reference path, and N represents the total number of frames simulated from the starting point of the interaction to the target point.

[0024] Furthermore, the second individual benefit is expressed by the following formula:

[0025]

[0026] Among them, the first individual income R i2 is the lane deviation gain, represents the projection point of the nth actual trajectory point of vehicle i on the reference path, It represents the nth actual trajectory point of vehicle i, and the lane offset value is the maximum distance between the actual trajectory point and the reference path.

[0027] Furthermore, the group benefit is expressed by the following formula:

[0028]

[0029] Among them, the group benefit R G To eliminate conflict benefits, n min represents the time frame when the interaction trajectories of vehicle i and vehicle j are closest, Indicates that vehicle i is in the nth min The position of the time frame, Indicates that vehicle j is in the nth min The position of the time frame.

[0030] Furthermore, solving the model in step S4 includes the following steps:

[0031] Step S4-1: The main vehicle plans the initial optimal trajectory based on the initial position of the interactive object with the goal of maximizing the benefit of the interactive object.

[0032] Step S4-2: Based on the initial optimal trajectory, the optimal trajectory of the main vehicle is planned with the goal of maximizing the utility of the main vehicle.

[0033] Step S4-3: Based on the optimal trajectory of the main vehicle, the optimal trajectory of the interactive object is updated with the goal of maximizing the utility of the interactive object.

[0034] Step S4-4: The optimal trajectory of the main vehicle and the optimal trajectory of the interacting object Recorded in If the newly obtained trajectory of the main vehicle and the trajectory of the interacting object Meeting the stationing conditions: The simulation process ends and is output on p record Otherwise, execute steps S4-2 and S4-3 to obtain the new trajectory of the main vehicle. and the trajectory of the interacting object

[0035] Furthermore, the social parameter is expressed by the following formula:

[0036]

[0037] in, is the social parameter, ω I1 ,ω I2 ,ω G The first individual income R I1 、The second individual income R I2 and group payoff R G The weight parameter of The closer it is to π / 2, the more cooperative the driver is. The closer it is to -π / 2, the more competitive the driver is.

[0038] In a second aspect, the present invention provides a computer system comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any one of the methods described above.

[0039] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the above-mentioned methods when executed by a processor.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. The driving interaction behavior model constructed by the present invention based on game theory can take into account both sociality and multi-level chain action dependencies, and can analyze complex and strong interaction scenarios with high risk, such as mixed human-machine driving, large deviation between social driving behavior and rule-based driving behavior, from a technical level.

[0042] 2. The social evaluation method based on inverse reinforcement learning proposed in the present invention can identify the heterogeneous sociality of drivers. By evaluating the sociality in the driving interaction process, it can be used to design, train and optimize autonomous driving algorithms with "social" driving behavior.

[0043] 3. By evaluating the social nature of driving behavior, the present invention provides a technical basis for promoting the evaluation of the "sociality" level of autonomous driving decision-making and planning algorithms, helps improve the social interaction capabilities of AVs, and is conducive to building an HV-AV mixed traffic system. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of the steps of the social evaluation method in driving interaction behavior based on inverse reinforcement learning of the present invention;

[0045] Figure 2 This is a framework diagram of the social evaluation method in driving interaction behavior based on inverse reinforcement learning of the present invention;

[0046] Figure 3 It is an unprotected left turn-straight-ahead interaction scene diagram of the present invention;

[0047] Figure 4 A social definition diagram for the present invention;

[0048] Figure 5 is a flow chart of the inverse reinforcement learning algorithm of the present invention;

[0049] Figure 6 The spatial coverage map of the traffic participant trajectories of the present invention;

[0050] Figure 7 This is a diagram showing the effectiveness verification results of the inverse reinforcement learning of the present invention;

[0051] Figure 8 A distribution diagram of the weight coefficients of the benefit function of the inverse reinforcement learning of the present invention;

[0052] Fig. 9This is a diagram of the social identification results of the inverse reinforcement learning of the present invention. DETAILED DESCRIPTION

[0053] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0054] Example

[0055] This embodiment provides a method for social evaluation in driving interaction behavior based on inverse reinforcement learning, which is used to perform social evaluation on interaction behaviors in an unprotected left turn interaction scenario at the intersection of Jianhe Road and Xianxia Road in Shanghai.

[0056] The data collection is to record the movement of traffic participants through video recording, combined with the high-precision video processing tool George to extract the movement parameters of traffic participants. The video records the traffic flow at the intersection from 4:00 to 5:30 in the afternoon, including the off-peak period from 4:00 to 4:30 and the peak period from 4:30 to 5:30. During the data collection process, the point where the left front wheel of the vehicle contacts the ground is used as the mark point, and the time interval is 0.12s. The pixel point where the mark point of each vehicle is located is manually marked in the image, and the software outputs the actual vehicle position coordinate point. The trajectory space coverage map is shown in the figure. Figure 6 As shown, it can be seen that the actual trajectories of various traffic participants in the intersection do not strictly follow the lanes.

[0057] Specifically, the method comprises the following steps: Figure 1 and Figure 2 As shown, this step can be divided into two parts:

[0058] The first part is to build an interactive behavior model based on game theory, which includes the following steps:

[0059] Step (1.1) Design the game participants:

[0060] In such Figure 3 In the scenario, the participants involved have two interactive individuals, namely the left-turning car and the straight-moving car. In particular, both the left-turning car and the straight-moving car can be the main car, that is, when the virtual driving interaction scene is constructed with the left-turning car, the main car is the left-turning car, and its interactive object is the straight-moving car. Similarly, when the virtual driving interaction is constructed with the straight-moving car, the main car is the straight-moving car, and its interactive object is the left-turning car.

[0061] In this embodiment, the left-turning vehicle is the main vehicle and the straight-moving vehicle is the interaction object. Figure 6 In the actual scenario, the left turn-straight-through interaction event is extracted. The participants involved have two interaction individuals, namely the left-turning vehicle and the straight-through vehicle.

[0062] Step (1.2) Design game behavior:

[0063] For this game model, the action includes two variables, namely acceleration a and steering angle h, which are recorded as motion control quantity u = [a, h]. Different from the discrete actions such as overtaking and giving way set in most game models, the action variables in this model are all continuous variables. According to the actual data analysis, the acceleration limit is (-3m / s 2 ,3m / s 2 ), the steering angle is limited to (-π / 6,π / 6).

[0064] Step (1.3) Design the game payoff function:

[0065] In order to express the sociality in driving interaction behavior through the benefit function, the present invention designs two individual benefit functions R I1 , R I2 and a group benefit R G , the overall benefit function is the linear sum of the three benefits, as shown below:

[0066] U=ω I1 ×R I1 +ω I2 ×R I2 +ω G ×R G

[0067] where ω I1 ,ω I2 ,ω G are the parameters to be identified for the three benefits. In the existing game model, these parameters are mostly specified manually, which has the problem of being difficult to reproduce the real traffic conditions. Therefore, it is necessary to identify the parameters to achieve the purpose of approaching the real traffic scene.

[0068] Specifically, the first individual benefit: the driving process benefit term R I1 , considering that drivers usually want to reach the destination from the starting point in the shortest time, the speed of the driving process is directly related to the driver's benefit. The specific expression of the driving process benefit item is as follows:

[0069]

[0070] Among them, τ i represents the projection length of the actual driving trajectory of vehicle i on the reference path, It represents the total length of the actual driving trajectory of vehicle i from the starting point of the interaction to the target point projected on the reference path, and N represents the total number of frames simulated from the starting point of the interaction to the target point.

[0071] Second individual benefit: Lane deviation benefit term R I2 Lane deviation refers to the degree of deviation of the vehicle relative to the center line of the lane during driving. Lane deviation has an important impact on the driver's driving safety and comfort. The specific expression of the lane deviation benefit term is as follows:

[0072]

[0073] in, represents the projection point of the nth actual trajectory point of vehicle i on the reference path, It represents the nth actual trajectory point of vehicle i, and the lane offset here is the maximum distance between the actual trajectory point and the reference path.

[0074] Group benefits: Resolving conflicts G ,In driving scenarios, resolving conflicts and maintaining traffic flow are very important goals. Conflict resolution refers to dealing with potential conflicts or conflict situations with other traffic participants to ensure safe and smooth traffic. The specific expression of conflict resolution is as follows:

[0075]

[0076] Among them, n min represents the time frame when the interaction trajectories of vehicle i and vehicle j are closest, Indicates that vehicle i is in the nth min The position of the time frame, Indicates that vehicle j is in the nth min The position of the time frame.

[0077] Step (1.4) solves the game process, including the following:

[0078] First, the left-turning car will plan a trajectory based on the initial position of the straight-moving car, and the planning principle is to maximize the benefits of the interactive object. Then, the left-turning car will calculate the optimal trajectory under the current trajectory based on the straight-moving car trajectory planned in the virtual scene. The principle of calculating the optimal trajectory is also to maximize the benefits of the left-turning car. In the process of continuous iteration, when convergence to an equilibrium state, the left-turning car will take the optimal behavior under the equilibrium solution as the behavior of the next simulation second.

[0079] The specific solution includes the following steps:

[0080] (1.4.1): To maximize the benefit of the straight-moving vehicle, initialize the optimal trajectory for the straight-moving vehicle Optimal trajectory for a straight-moving vehicle The calculation method is:

[0081]

[0082]

[0083]

[0084]

[0085] Among them, U i (·) is the revenue function of the straight-moving vehicle; is the position of the straight-moving vehicle i at time t, i.e., the superscript t represents the time point, and the subscript i is the identity number of a specific straight-moving vehicle (i = 1, 2, 3 ... N); u is the motion control amount of the straight-moving vehicle; is the limit value of the vehicle motion control amount; r(·) is the function for calculating the distance from a specific position to the center line of the lane where it is located; w lane is the lane width; w veh is the vehicle's body width.

[0086] (1.4.2): Based on the initial optimal trajectory of the straight-moving vehicle, the trajectory of the left-turning vehicle is planned with the goal of maximizing the utility of the left-turning vehicle. The trajectory of the left-turning vehicle is calculated as follows:

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093] Among them, U 0 (·) is the revenue function of the left-turning vehicle; R I1 (·), R I2 (·) are the individual benefits of left-turning vehicles, R G (·) is the group benefit of left-turning vehicles; l veh is the vehicle's body length.

[0094] (1.4.3): Based on the optimal trajectory of the left-turning vehicle, the trajectory of the straight-moving vehicle is planned with the goal of maximizing the utility of the straight-moving vehicle. The trajectory of the straight-moving vehicle is calculated as follows:

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101] Among them, U i (·) is the revenue function of the straight-moving vehicle, R I1 (·), R I2 (·) are the individual benefits of the straight-moving vehicles, R G (·) is the group benefit of the straight-moving vehicle; l veh is the vehicle's body length.

[0102] (1.4.4): The optimal trajectory of the left-turning vehicle and the optimal trajectory of the straight-moving vehicle Recorded in If the newly acquired left-turning vehicle trajectory and the trajectory of the straight-moving vehicle Meeting the stationing conditions: The simulation process ends and is output on p record ; Otherwise, execute steps (1.4.2) and (1.4.3) to obtain the new trajectory of the left-turning vehicle and the trajectory of the straight-moving vehicle

[0103] After step (1.4.4) is completed, you can record The interactive trajectory results in the virtual scene are obtained.

[0104] The second part uses inverse reinforcement learning to identify social parameters in the interactive behavior model and then evaluate the sociality of the interactive process, such as Figure 5 As shown, the specific steps are as follows:

[0105] Step (2.1) assigning an initial value to the weight parameter;

[0106] Step (2.2) generates an optimal interaction trajectory according to the overall benefit function currently described; during the interaction process, both parties will take the optimal action of the next frame in the optimal interaction trajectory until the interaction ends;

[0107] Step (2.3) extracts the features μ(π) of the optimal interaction trajectory during the interaction process through the inverse reinforcement learning algorithm and compares it with the trajectory features μ(E) in the empirical data;

[0108] Specifically, the left-turning vehicle and the straight-moving vehicle, as the interacting subjects, observe the environmental state at time t and obtain the state quantity s t , through the game strategy π, the action a at time t is generated t =π(s t ), the environment generates the state of the next moment after receiving the action returned by the interactive subject. For each observed state quantity s t , the interacting subjects will make corresponding actions a t , so the trajectory δ can be defined as the initial state-action pair (s 1 ,a 1 ) to the terminal state—action pair (s N ,a N ) is a state-action sequence:

[0109] δ={(s 1 ,a 1 ),…,(s N ,a N )}

[0110] The state information required by each interactive subject in the decision-making process is divided into the state of the vehicle s according to different sources ego and the state s of the interacting object inter ,Right now:

[0111] s ego =[p x-ego ,p y-ego ,v x-eho ,v y-ego ,h ego ]

[0112] s inter =[p x-inter ,p y-inter ,v x-inter ,v y-inter ,h inter ]

[0113] Among them, p x ,p y ,v x ,v y , h represent the vehicle's x-position, vehicle's y-position, vehicle's x-speed, vehicle's y-speed, and vehicle's front angle, respectively.

[0114] The feature φ based on the interaction trajectory is extracted as μ(π) as follows:

[0115]

[0116] Step (2.4) uses the maximum margin algorithm to solve the optimization problem To update the weight parameters of the overall benefit function, repeatedly perform steps 2.2 to 2.4 until convergence;

[0117] Step (2.5) Design social parameters according to the weight parameters like Figure 4 shown.

[0118] Social Parameters It is expressed by the following formula:

[0119]

[0120] when The closer it is to π / 2, the more cooperative the driver is. The closer it is to -π / 2, the more competitive the driver is.

[0121] Inverse reinforcement learning is used to identify the social parameters of the interactive behavior model, and the interactive trajectory fragments of left-turning vehicles are extracted from the intersection of Jianhe Road and Xianxia Road. The Wasserstein Distance of the trajectory in the horizontal and vertical directions (referred to as Trajectory_x and Trajectory_y) is calculated to measure the trajectory accuracy. Figure 7 As shown, the specific indicator values ​​are shown in Table 1. Wasserstein Distance of simulation data and empirical data trajectories. Figure 7 (a) represents the trajectory distribution of left-turning vehicles in actual data. Figure 7 (b) shows the trajectory distribution of left-turning vehicles in the simulation data based on the game model.

[0122]

[0123] Table 1

[0124] The distribution of the weight coefficient of the reward function of inverse reinforcement learning in the left-turn rushing scenario is as follows: Figure 8 As shown, Figure 8 (a) shows the distribution of the weight coefficients of the benefit function of the left-turn vehicle, which are the weight coefficients of the three features of driving progress, lane deviation and conflict resolution. Figure 8 (b) shows the distribution of weight coefficients of the profit function of straight-moving vehicles.

[0125] The identification results are obtained through social parameters such as Fig. 9 As shown in the figure, the horizontal axis represents the social tendency of left-turning vehicles, and the vertical axis represents the social tendency of straight-moving vehicles. Each red scattered point represents the social tendency of both left-turning vehicles and straight-moving vehicles in an interaction event. From this, we can clearly see that nearly half of the left-turning vehicles show strong competitiveness, while the competitiveness of the vast majority of left-turning vehicles is weak, thus achieving the evaluation of interactive behavior.

[0126] In summary, the social evaluation method based on inverse reinforcement learning can explicitly express and approximate implicit social normative behaviors and identify the "social" driving interaction behaviors of human drivers.

[0127] This embodiment provides a computer system, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any one of the above methods.

[0128] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program / instructions are executed by a processor, the steps of any one of the above methods are implemented.

[0129] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A social evaluation method in driving interaction behavior based on inverse reinforcement learning, characterized in that: The following steps are involved: Step S1, constructing an interactive behavior model based on game theory, including designing game participants, where the objects are the two parties of interaction: the main vehicle and the interactive object; Step S2, designing a game behavior, wherein the behavior includes two variables: acceleration a and steering angle h; Step S3, designing a game benefit function, the function includes a first individual benefit, a second individual benefit and a group benefit, wherein the three benefits are linearly summed to obtain an overall benefit function, which is expressed by the following formula: U=ω I1 ×r I1 +oh I2 ×R I2 +oh G ×R G ; Among them, U is the overall profit function, ω I1 ,ω I2 ,ω G The first individual income R I1 、The second individual income R I2 and group payoff R G The weight parameter of Step S4, solving the model to obtain the trajectories of the main vehicle and the interactive object respectively; Step S5, assigning an initial value to the weight parameter; Step S6: Generate an optimal interaction trajectory according to the current overall benefit function; during the interaction process, both parties will take the optimal action of the next frame in the optimal interaction trajectory until the interaction ends; Step S7, extracting the feature μ(π) of the optimal interaction trajectory in the interaction process through the inverse reinforcement learning algorithm, and comparing it with the trajectory feature μ(E) in the empirical data; Step S8: Solve the optimization problem using the maximum margin algorithm To update the weight parameters, repeatedly perform steps S6 to S7 until convergence; Step S9: designing social parameters according to the weight parameters, and implementing evaluation of interactive behaviors through the social parameters.

2. According to claim 1, a social evaluation method in driving interaction behavior based on inverse reinforcement learning is characterized in that: The main vehicle includes a steering vehicle and a straight-moving vehicle, and the interactive objects correspond to the main vehicle as the straight-moving vehicle and the steering vehicle.

3. The social evaluation method in driving interaction behavior based on inverse reinforcement learning according to claim 1 is characterized in that: The range of the acceleration a is (-3m / s 2 ,3m / s 2 ); the range of the steering angle h is (-π / 6,π / 6).

4. The social evaluation method in driving interaction behavior based on inverse reinforcement learning according to claim 1 is characterized in that: The first individual benefit is expressed by the following formula: Among them, the first individual income R I1 is the driving process benefit, τ i represents the projection length of the actual driving trajectory of vehicle i on the reference path, It represents the total length of the actual driving trajectory of vehicle i from the starting point of the interaction to the target point projected on the reference path, and N represents the total number of frames simulated from the starting point of the interaction to the target point.

5. The social evaluation method in driving interaction behavior based on inverse reinforcement learning according to claim 1 is characterized in that: The second individual benefit is expressed by the following formula: Among them, the first individual income R I2 is the lane deviation gain, represents the projection point of the nth actual trajectory point of vehicle i on the reference path, It represents the nth actual trajectory point of vehicle i, and the lane offset value is the maximum distance between the actual trajectory point and the reference path.

6. The method for evaluating sociality in driving interaction behavior based on inverse reinforcement learning according to claim 1, characterized in that: The group benefit is expressed by the following formula: Among them, the group benefit R G To eliminate conflict benefits, n min represents the time frame when the interaction trajectories of vehicle i and vehicle j are closest, Indicates that vehicle i is in the nth min The position of the time frame, Indicates that vehicle j is in the nth min The position of the time frame.

7. The method for evaluating sociality in driving interaction behavior based on inverse reinforcement learning according to claim 1, characterized in that: Solving the model in step S4 includes the following steps: Step S4-1: The main vehicle plans the initial optimal trajectory based on the initial position of the interactive object with the goal of maximizing the benefit of the interactive object. Step S4-2: Based on the initial optimal trajectory, the optimal trajectory of the main vehicle is planned with the goal of maximizing the utility of the main vehicle. Step S4-3: Based on the optimal trajectory of the main vehicle, the optimal trajectory of the interactive object is updated with the goal of maximizing the utility of the interactive object. Step S4-4: The optimal trajectory of the main vehicle and the optimal trajectory of the interacting object Recorded in If the newly obtained trajectory of the main vehicle and the trajectory of the interacting object Meeting the stationing conditions: The simulation process ends and is output on p record Otherwise, execute steps S4-2 and S4-3 to obtain the new trajectory of the main vehicle. and the trajectory of the interacting object 8. The method for evaluating sociality in driving interaction behavior based on inverse reinforcement learning according to claim 1, characterized in that: The social parameter is expressed by the following formula: in, is the social parameter, ω I1 ,ω I2 ,ω G The first individual income R I1 、The second individual income R I2 and group payoff R G The weight parameter of The closer it is to π / 2, the more cooperative the driver is. The closer it is to -π / 2, the more competitive the driver is.

9. A computer system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.