Multi-target income weight dynamic updating method, device and system and storage medium
By updating the income weight based on Bayesian reasoning in intelligent vehicles, the problem that weight values in the prior art cannot accurately describe the vehicle's dynamic decision preferences is solved, and personalized updates of vehicle decision preferences in dynamic environments are realized, and decision flexibility of intelligent vehicles is improved.
Patent Information
- Application Number
- CN202510017948.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
AI Technical Summary
When building a profit function, the weight value of existing intelligent vehicles mostly depends on subjective assignment or fixed value, and cannot accurately describe the vehicle's personalized decision-making preferences during dynamic driving, and cannot express the changes in the vehicle's decision-making preferences during interaction in real time.
By judging driving tendency based on the vehicle's driving state, calculating vehicle status characterization indicators, and using natural driving data to extract the intrinsic relationship between these indicators and driving tendency, Bayesian reasoning is used to update the multi-target profit weights, and express the vehicle's dynamic decision-making preferences.
It realizes dynamic updates of vehicle decision preferences, which can more accurately reflect the personalized driving tendency of vehicles in dynamic driving environments, and improves the decision-making flexibility and anthropomorphism of intelligent vehicles in complex traffic environments.
Smart Images

Figure CN119987891A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent vehicle technology, and in particular relates to a method and device, a system, and a storage medium for dynamically updating multi-objective benefit weights. Background Art
[0002] Game decision-making is one of the effective ways to break through defensive decision-making. In the driving scenario of a lane-changing vehicle cutting in, the dynamic game between the intelligent vehicle and the cutting-in vehicle is conducive to improving the decision-making flexibility and anthropomorphism of the intelligent vehicle. In the game model, the profit function can be used to calculate the benefits and losses obtained by the participants when choosing different strategies, which is an important basis for decision-making.
[0003] At present, the definition function method is used to construct the benefit function, that is, based on the vehicle's motion state, multi-dimensional benefit indicators such as safety, efficiency, comfort and fuel economy are proposed, and finally a multi-objective benefit function is constructed by weighting the benefit indicators. The weight coefficient of each objective reflects the decision-making preference of the game participants and is related to the game result; however, in a complex and changeable driving environment, the motion state of the intelligent vehicle and the surrounding vehicles is constantly updated, and the game state of the two vehicles and the decision-making preference of the vehicle will also be affected.
[0004] The existing methods have the following shortcomings when constructing the benefit function: on the one hand, the weight values in the benefit function mostly rely on subjective assignments, or borrow weight values from existing studies, which ignores the personalized driving tendencies of vehicles and makes it difficult to accurately describe the differences in decision-making preferences of different vehicles during dynamic driving; on the other hand, the weight values in the benefit function mostly use fixed values, which obviously does not meet the dynamic and time-varying traffic environment and cannot express the changes in the decision-making preferences of vehicles during the interaction process in real time. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method and device, system and storage medium for dynamically updating multi-objective benefit weights.
[0006] To achieve the above object, the present invention adopts the following technical solution:
[0007] A multi-objective benefit weight dynamic updating method, comprising:
[0008] S1. judging the driving tendency of the vehicle based on the driving state of the vehicle;
[0009] S2. Calculating a vehicle state characterization index based on the vehicle driving state;
[0010] S3, extracting the intrinsic relationship between vehicle state representation indicators and driving tendency based on natural driving data;
[0011] S4. Update the multi-objective benefit weights based on Bayesian reasoning to express the dynamic decision-making preferences of the vehicle.
[0012] Preferably, the vehicle status characterization index includes: real-time characterization indexes of vehicle safety, efficiency and comfort;
[0013] Preferably, in S3, unsupervised clustering and frequency statistical analysis methods are used to extract the intrinsic relationship between vehicle state characterization indicators and driving tendency from natural driving data.
[0014] The present invention also provides a multi-objective benefit weight dynamic updating device, comprising:
[0015] A first processing module, used for determining a driving tendency of a vehicle based on a driving state of the vehicle;
[0016] A second processing module, used for calculating a vehicle state characterization index based on the vehicle driving state;
[0017] The third processing module is used to extract the intrinsic relationship between the vehicle state representation index and the driving tendency according to the natural driving data;
[0018] The fourth processing module is used to update the multi-objective benefit weights based on Bayesian reasoning to express the dynamic decision-making preference of the vehicle.
[0019] Preferably, the vehicle status characterization index includes: real-time characterization indexes of vehicle safety, efficiency and comfort;
[0020] Preferably, the third processing module uses unsupervised clustering and frequency statistical analysis methods to extract the intrinsic relationship between vehicle state characterization indicators and driving tendencies from natural driving data.
[0021] An embodiment of the present invention also provides a multi-objective benefit weight dynamic update system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a multi-objective benefit weight dynamic update method when executed by the processor.
[0022] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes a method for dynamically updating multi-objective benefit weights when running.
[0023] The present invention analyzes the relationship between the vehicle's decision preference and driving tendency, and realizes the dynamic update of the benefit weight through Bayesian reasoning. Based on the natural driving data set, combined with statistical analysis and clustering methods, the intrinsic relationship between driving tendency and safety, efficiency, and comfort benefit indicators is deeply explored, providing a basis for the personalized setting of benefit weights; in addition, by constructing a Bayesian reasoning model, the multi-objective benefit weights in the benefit function are dynamically updated according to the real-time driving status and driving tendency of the vehicle, solving the problem that fixed weights are difficult to adapt to dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0025] Figure 1 It is a decision-making scenario for autonomous driving vehicle behavior when a car cuts in;
[0026] Figure 2 This is a flow chart of a method for dynamically updating multi-objective benefit weights according to an embodiment of the present invention;
[0027] Figure 3 The distribution of driving tendency and each target characterization index; (a) is driving tendency, (b) is safety, (c) is efficiency, and (d) is comfort;
[0028] Figure 4 The clustering results of driving tendency and each target are shown in Figure 2. (a) is driving toughness (0.693), (b) is safety (0.724), (c) is comfort (0.803), and (d) is efficiency (0.758).
[0029] Figure 5 Update security return weights schematic for Bayesian reasoning;
[0030] Figure 6 Dynamic update results of multi-objective benefit weights for the following scenarios;
[0031] Figure 7 This is the dynamic update result of multi-objective benefit weights in scenario 2. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] Embodiment 1:
[0035] The present invention relies on common car-entry scenarios, such as Figure 1As shown in the figure, the two parties in the game are EV and LCV, where EV is the ego vehicle, which is an autonomous vehicle, and LCV is an adjacent vehicle with the intention of changing lanes, which is also the opponent vehicle of EV. EF and LF are the front vehicles of EV and LCV respectively. Considering that the ego vehicle is an autonomous vehicle, its profit weight can be set by the user and can also change according to the scene and demand. In contrast, the profit weight of the opponent vehicle is directly related to its decision preference, which in turn affects the decision of the ego vehicle. Therefore, the dynamic update method of the profit weight of the present invention is mainly aimed at the opponent vehicle (LCV).
[0036] like Figure 2 As shown, the present invention provides a method for dynamically updating multi-objective benefit weights, comprising the following steps:
[0037] S1. judging the driving tendency of the vehicle based on the driving state of the vehicle;
[0038] S2, calculate the real-time characterization index of vehicle safety, efficiency and comfort based on the vehicle driving status;
[0039] S3, using unsupervised clustering and frequency statistical analysis methods to extract the intrinsic relationship between vehicle state representation indicators and driving tendency from natural driving data;
[0040] S4. Update the multi-objective benefit weights based on Bayesian reasoning to express the dynamic decision-making preferences of the vehicle.
[0041] As an implementation method of an embodiment of the present invention, in S1, a driver's short-term driving tendency prediction model is used, and the model input is the vehicle's real-time motion state information and the driving environment information, specifically: the vehicle's yaw angular velocity, the turn signal on state and the lateral distance from the lane boundary line, the vehicle's relative longitudinal distance from its neighboring vehicles, the relative longitudinal speed and the longitudinal headway; the model output is the toughness of the vehicle driver (0-1); specifically refer to the patent document "CN117141502A, a method for quantifying and identifying instantaneous driving tendency".
[0042] As an implementation method of an embodiment of the present invention, in S2, corresponding real-time characterization indicators are proposed for the safety, efficiency and comfort goals in the benefit function, the purpose of which is to describe the safety, efficiency and comfort levels of the vehicle at different times.
[0043] 1. Real-time driving safety indicators
[0044] The driving risk model is used to calculate the real-time value of the safety target SI to express the safety loss value of the vehicle driving process, as follows:
[0045] SI(t)=R(t)
[0046] 2. Real-time characterization indicators of driving efficiency
[0047] From the aspects of speed and space, the efficiency target real-time representation index EI is proposed to describe the real-time change of efficiency during driving, as shown in the following formula:
[0048] EI(t)=(I speed (t)+I space (t)) / 2
[0049]
[0050] Among them, the subscript LF is the front vehicle in the lane of the lane-changing vehicle, and EF is the front vehicle in the target lane.
[0051] 3. Real-time characterization indicators of driving comfort
[0052] Combining the vehicle's real-time acceleration change rate jerk and the maximum acceleration change rate jerkmax specified by the ACC regulations, a real-time characterization index CI for comfort target is proposed:
[0053]
[0054] As an implementation of the embodiment of the present invention, S3 includes:
[0055] (1) Real-time characterization indicators of vehicle driving safety, efficiency, comfort and numerical distribution analysis of vehicle driving tendency.
[0056] In order to ensure the diversity of driving tendencies, we selected vehicle driving data sets under different traffic environments, NGSIM natural driving data sets, and multi-vehicle collaborative simulation data sets to enrich the types of driving tendency samples. A total of 986 vehicle lane change interaction samples were obtained from the data sets. For each sample, the driving toughness of the lane-changing vehicle and the real-time characterization index values of safety, efficiency, and comfort at the corresponding moment were calculated. The frequency distribution histogram is shown in the figure below.
[0057] (2) K-means clustering is used to divide vehicle status representation indicators and driving tendencies into intervals.
[0058] Using K-means Figure 3 The four indicators of driving tendency, safety, efficiency and comfort in (a), (b), (c) and (d) are clustered and divided into representative value ranges. The results are shown in Figure 4 The values of the silhouette coefficients of the clustering results are shown in (a), (b), (c), and (d), which prove the rationality of the clustering process. Through cluster analysis, the driving toughness and each target representation index can be divided into multiple value intervals, and several typical intervals can be used to simply represent the continuous values between 0 and 1. This step helps to more simply and quickly analyze the correlation between the vehicle driving tendency and each target.
[0059] (3) The frequency statistical analysis method is used to construct the joint probability distribution matrix of vehicle driving tendency and driving state characterization indicators.
[0060] In order to obtain the correlation between the safety, efficiency and comfort indicators and driving tendency during vehicle driving, the frequency statistical analysis method is used. Based on the clustering result of step (2), the joint probability of the vehicle driving safety real-time characterization index CI and the driving toughness H is taken as an example. The calculation method is as follows:
[0061]
[0062] Among them, h∈[h 1 ,h 2 ],s∈[s 1 ,s 2 ],N(H∈[h 1 ,h 2 ]∩SI∈[s 1 ,s 2 ]) is the driving toughness of the vehicle H in [h 1 ,h 2 ] interval and the safety index SI is in [s 1 ,s 2 ] is the frequency of , and Num is the total sample size.
[0063] The above method can be used to obtain the joint probability matrix between the safety, efficiency and comfort indicators and the vehicle driving tendency during vehicle driving, as shown in Tables 1 to 3. In the table, SI represents the vehicle driving risk and reflects the degree of safety loss. CI reflects the fluctuation level of acceleration, which directly affects the ride comfort of the vehicle. EI reflects the gains and losses of space and speed during vehicle driving. When the efficiency value EI is negative, it means that the current driving efficiency of the vehicle is much lower than the driving efficiency of the target lane, that is, the target lane has better driving space and speed conditions.
[0064] Table 1
[0065]
[0066] Table 2
[0067]
[0068] Table 3
[0069]
[0070] As an implementation method of an embodiment of the present invention, in S4, during actual driving, the vehicle's (LCV) attention to safety, efficiency and comfort is closely related to the driving attitude of the vehicle driver. Therefore, the present invention adopts Bayesian reasoning to realize the dynamic update process of safety, efficiency and comfort benefit weights along with the vehicle driving tendency.
[0071] With the safety return weight w s Take this as an example, the schematic diagram of dynamic update of income weight is as follows Figure 5 The overall idea of this method is to use the Bayesian rule, combine the motion state of the opponent vehicle (LCV) and the real-time driving tendency of the vehicle, and based on the joint probability matrix between the safety, efficiency and comfort indicators and the driving tendency obtained in step 3, use the Bayesian rule to update the prior probability of the benefit weight to obtain the posterior probability.
[0072] Bayesian reasoning is a reasoning method based on Bayes' rule. Its core idea is to modify existing beliefs (prior probabilities) based on conditional probabilities through newly observed phenomena, and then obtain updated beliefs (posterior probabilities) to continuously modify and update the probability of an event. The basic probability knowledge of Bayes' rule is briefly explained as follows:
[0073] Let Ω represent the sample space. If event A 1 , A 2 , …, A n are mutually incompatible, and A 1 UA 2 U…UA n =Ω,P(A i )>0(i=1,2,…n), then the conditional probability can be obtained:
[0074]
[0075] For P(B), there is a total probability formula:
[0076]
[0077] Substituting equation (5.37) into equation (5.36), the mathematical expression of Bayes' rule is as follows:
[0078]
[0079] Among them, P(A i ) is event A i The prior probability of i ) is the conditional probability of B after A occurs; P(A i |B) is the conditional probability of A after knowing that B has occurred, also called the posterior probability of A.
[0080] Based on the above principles, the safety benefit weight w s As an example, the weight update formula using Bayesian reasoning is as follows:
[0081]
[0082] Among them, P (focus on safety), P (focus on efficiency), and P (focus on comfort) represent the prior probability w of the weights of the multi-objective benefit function respectively. s 、w e 、w c ; P(Focus on Safety|H=h) represents the w obtained after the prior probability is corrected after adding the opponent's vehicle driving tendency information s The posterior probability w s * ; P(H=h|SI=s) is the probability of the coexistence of event A (the vehicle driving toughness value is h) and event B (the safety real-time characterization index value is s). This value can be provided by the joint probability distribution matrix of driving tendency and safety target (Table 1). Similarly, P(H=h|EI=e) and P(H=h|CI=c) can be determined in the same way.
[0083] Application examples:
[0084] Relying on the theory of dynamic game with incomplete information, the profit function and strategy space are designed, and the multi-objective profit weight dynamic update method proposed in this patent is combined to construct the game model IDGM-DD, and the game process of the model is demonstrated based on two scenarios. The basic operating state parameters of the intelligent vehicle EV and the opponent vehicle LCV at the initial moment of the game are shown in Table 4.
[0085] Table 4
[0086]
[0087]
[0088] (1) Scenario 1
[0089] In scenario 1, there is a slow vehicle LF not far in front of the LCV, causing the driving space of its lane to be smaller than that of the target lane. At the same time, the acceleration difference between the LCV and the target rear vehicle EV is not much, but its speed is greater than that of the EV, and its longitudinal position is ahead of the EV vehicle.
[0090] In this driving state, the updating process of the driving tendency and benefit weight of the opponent vehicle LCV in the game is as follows: Figure 6As shown in the figure, the lane-changing tendency of LCV continues to increase, which is consistent with its driving attitude of insisting on lane changing. At the same time, the weight values of safety and comfort decrease, while the weight value of efficiency increases, indicating that as the lane-changing attitude becomes more stringent, the vehicle pays more attention to driving efficiency, which is consistent with objective facts.
[0091] (2) Scenario 2
[0092] Compared with Scenario 1, the driving states of the two vehicles at the initial moment of the game are quite different in Scenario 2. First, the relative distance and relative speed between the LCV and the vehicle in front of it, LF, indicate that the driving condition of the LCV in the original lane is better; second, the relative motion state parameters between the LCV and the EV indicate that even though the LCV is ahead of the EV in the longitudinal position, its initial acceleration and speed are both smaller than that of the EV.
[0093] In this driving state, the updating process of the driving tendency and benefit weight of the opponent vehicle LCV in the game is as follows: Figure 7 As shown in the figure, the lane-changing tendency of LCV increases first and then decreases, which is consistent with the result of the strategy selection that the vehicle intends to change lanes but cancels the lane-changing in the fourth game stage. At the same time, the efficiency weight value shows a trend of increasing first and then decreasing, the safety weight value decreases slightly and remains unchanged in the later stage, and the comfort weight value decreases first and then increases, which is consistent with the change process of the toughness of its lane-changing attitude, indicating that the vehicle pursues efficiency more in the early stage of the game, but because the EV vehicle insists on not giving way, it changes its lane-changing tendency and gives up the intention to change lanes and drives in its own lane.
[0094] Embodiment 2:
[0095] The embodiment of the present invention also provides a multi-objective benefit weight dynamic updating device, comprising:
[0096] A first processing module, used for determining a driving tendency of a vehicle based on a driving state of the vehicle;
[0097] A second processing module, used for calculating a vehicle state characterization index based on the vehicle driving state;
[0098] The third processing module is used to extract the intrinsic relationship between the vehicle state representation index and the driving tendency from the natural driving data;
[0099] The fourth processing module is used to update the benefit weight based on Bayesian reasoning to express the dynamic decision preference of the vehicle.
[0100] As an implementation method of the embodiment of the present invention, the vehicle state characterization index includes: real-time characterization indexes of vehicle safety, efficiency and comfort;
[0101] As an implementation method of the embodiment of the present invention, the third processing module uses unsupervised clustering and frequency statistical analysis methods to extract the intrinsic relationship between the vehicle state representation index and the driving tendency from the natural driving data.
[0102] Embodiment 3:
[0103] An embodiment of the present invention also provides a multi-objective benefit weight dynamic update system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a multi-objective benefit weight dynamic update method when executed by the processor.
[0104] Embodiment 4:
[0105] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes a method for dynamically updating multi-objective benefit weights when running.
[0106] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A multi-objective benefit weight dynamic updating method, characterized in that: include: S1. judging the driving tendency of the vehicle based on the driving state of the vehicle; S2. Calculating a vehicle state characterization index based on the vehicle driving state; S3, extracting the intrinsic relationship between vehicle state representation indicators and driving tendency based on natural driving data; S4. Update the multi-objective benefit weights based on Bayesian reasoning to express the dynamic decision-making preferences of the vehicle.
2. The multi-objective benefit weight dynamic updating method according to claim 1, characterized in that: Vehicle status characterization indicators include: real-time characterization indicators of vehicle safety, efficiency and comfort.
3. The multi-objective benefit weight dynamic updating method according to claim 2, characterized in that: In S3, unsupervised clustering and frequency statistical analysis methods are used to extract the intrinsic relationship between vehicle state representation indicators and driving tendency from natural driving data.
4. A multi-objective benefit weight dynamic updating device, characterized in that: include: A first processing module, used for determining a driving tendency of a vehicle based on a driving state of the vehicle; A second processing module, used for calculating a vehicle state characterization index based on the vehicle driving state; The third processing module is used to extract the intrinsic relationship between the vehicle state representation index and the driving tendency according to the natural driving data; The fourth processing module is used to update the multi-objective benefit weights based on Bayesian reasoning to express the dynamic decision-making preference of the vehicle.
5. The multi-objective benefit weight dynamic updating device method as claimed in claim 4, characterized in that: Vehicle status characterization indicators include: real-time characterization indicators of vehicle safety, efficiency and comfort.
6. The multi-objective benefit weight dynamic updating device according to claim 5, characterized in that: The third processing module uses unsupervised clustering and frequency statistical analysis to extract the intrinsic relationship between vehicle state representation indicators and driving tendencies from natural driving data.
7. A multi-objective benefit weight dynamic update system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the method for dynamically updating the multi-objective benefit weights as described in any one of claims 1 to 3 is executed.
8. A storage medium, characterized in that: The storage medium stores a computer program, which, when running, executes the multi-objective benefit weight dynamic updating method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Instantaneous driving tendency quantification method and identification method
CN117141502A