Hybrid traffic coordination decision based on social value orientation and hybrid expert model

By combining the social intent perception module and the hybrid expert prediction module, the social behavioral preferences of human drivers are dynamically learned and predicted, enabling safe and efficient collaborative decision-making for autonomous vehicles in mixed traffic environments. This solves the problems of behavioral heterogeneity and insufficient safety mechanisms in existing technologies, and improves the collaborative smoothness and safety of the system.

CN122369290APending Publication Date: 2026-07-10HUBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610464896.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-09
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing collaborative decision-making models for autonomous vehicles lack consideration of human drivers' social behavioral preferences in mixed traffic environments, and their safety mechanisms lack forward-looking predictive capabilities, leading to a chain reaction of behavioral heterogeneity and potential conflicts, making it difficult to achieve safe and efficient collaborative decision-making.

Method used

The Social Intent Perception Module (SIPM) is used to dynamically learn the Social Value Orientation (SVO) of connected autonomous vehicles. Combined with the Hybrid Expert Prediction Module (MEPM), multi-step trajectory prediction and conflict risk assessment are performed. The decision-making module balances rewards and corrects actions, forming a decision-making closed loop that integrates social intelligence and forward-looking coordination.

Benefits of technology

It enables more human-like collaborative decision-making in complex mixed traffic scenarios, improves the system's collaborative smoothness, safety and traffic efficiency, dynamically adjusts behavior to adapt to the social value orientation of human drivers, and proactively avoids potential conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369290A_ABST
    Figure CN122369290A_ABST
Patent Text Reader

Abstract

The application relates to the field of intelligent traffic systems and automatic driving technologies, and discloses a hybrid traffic collaborative decision based on social value orientation and a hybrid expert model; through a social intention perception module, the application enables a networked automatic driving vehicle to dynamically learn and adjust the social value orientation of the vehicle, so that the vehicle adaptively balances between self-interest and altruism, and generates a more human-like and more acceptable driving strategy; meanwhile, through multi-step trajectory prediction and conflict risk assessment, a hybrid expert prediction module changes the decision mode from passive response to active coordination, and can foresee and resolve potential conflicts; the two modules work together to form a new decision closed loop with social intelligence and foresight, and finally improve the collaborative fluency, safety and traffic efficiency of the overall system in a complex hybrid traffic scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent transportation systems and autonomous driving technology, specifically to hybrid traffic collaborative decision-making based on social value orientation and hybrid expert models. Background Technology

[0002] With the development of autonomous driving technology, connected autonomous vehicles (CAVs) are expected to improve road traffic efficiency and safety through vehicle-to-vehicle (V2V) communication and collaborative decision-making. However, in mixed traffic environments where CAVs and human-driven vehicles (HDVs) coexist for a long time, especially in high-conflict scenarios such as merging zones and intersections, achieving safe, efficient, and human-understandable collaborative decision-making still faces significant challenges.

[0003] Currently, multi-agent reinforcement learning (MARL) methods have become the mainstream paradigm for solving collaborative decision-making problems in CAVs (Collaborative Action Vehicles). Among these, algorithms employing centralized training and decentralized execution (CTDE), such as Multi-Agent Proximal Policy Optimization (MAPPO), have achieved significant results in specific scenarios by introducing mechanisms such as Prior Intent Sharing and Safety Enhanced Module (SEM). However, those skilled in the art recognize that existing collaborative decision-making models still have the following limitations that urgently need to be addressed:

[0004] First, existing models generally lack consideration of human drivers' social behavioral preferences. They typically treat all traffic participants as individuals with the same cooperative inclinations, ignoring the behavioral heterogeneity in real traffic caused by drivers' different social value orientations (SVOs)—such as competitive (selfish) or cooperative (altruistic). This difference is crucial when CAVs interact with HDVs, directly affecting the effectiveness and acceptability of collaborative strategies.

[0005] Secondly, existing safety mechanisms, such as Safety Enhancement Modules (SEMs), are essentially passive response and correction mechanisms based on the current or short-term state. They lack the ability to accurately predict the future trajectory of vehicles over long periods, as well as the proactive coordination capabilities based on this prediction. This "short-sighted" behavior makes it difficult to avoid the chain reaction of potential conflicts, limiting the reliability and smoothness of the system in complex dynamic environments.

[0006] Furthermore, although existing research has attempted to introduce Social Value Orientation (SVO) into the field of autonomous driving, most studies treat it as a static, pre-set parameter, failing to simulate the ability of human drivers to dynamically adjust their social preferences based on real-time traffic conditions during interactions. Meanwhile, in terms of behavior prediction, traditional intent inference methods have limited accuracy in mixed traffic environments with high uncertainty in high-density traffic (HDV) behavior.

[0007] Therefore, there is an urgent need in this field for a new decision-making framework that can deeply integrate social behavior cognition and forward-looking prediction and coordination capabilities, so that CAVs can not only make safe and efficient decisions in mixed traffic, but also demonstrate human-like social intelligence in their behavior, thereby achieving true collaboration. To this end, a mixed traffic collaborative decision-making based on social value orientation and hybrid expert model is proposed. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a hybrid traffic collaborative decision-making system based on social value orientation and a hybrid expert model to solve the problems mentioned in the background.

[0009] To achieve the above objectives, the present invention provides the following technical solution: a hybrid traffic collaborative decision-making system based on social value orientation and a hybrid expert model, comprising:

[0010] The Social Intent Perception Module (SIPM) is used to generate and apply Social Value Orientation (SVO) values ​​for connected autonomous vehicles (CAVs) and human-driven vehicles (HDVs) in a mixed traffic environment; wherein, the social value orientation (SVO) values ​​of the connected autonomous vehicles (CAVs) are dynamically learned through a trainable network, and the social value orientation (SVO) values ​​of the HDVs follow a preset bimodal distribution static initialization.

[0011] The Hybrid Expert Prediction Module (MEPM) consists of a pre-trained hybrid expert model that receives feature vectors including vehicle state, social value orientation (SVO) value, and environmental context, and outputs predictions of future multi-step trajectories and collision risk assessments between vehicles.

[0012] The decision-making module integrates the SIPM and MEPM, and is configured to use the Social Value Orientation (SVO) value to balance rewards and adjust traffic priorities, and to perform safety supervision and action correction based on the prediction results of the MEPM, so as to output cooperative driving decisions.

[0013] Preferably, in the Social Intent Perception Module (SIPM), the Social Value Orientation (SVO) value of the connected autonomous vehicle (CAV) is dynamically generated through an independent multilayer perceptron network, and the loss function of this network is... The following weighting factors were combined:

[0014] From the advantage function Weighted Social Values ​​(SVO) ;

[0015] Team average reward Weighted Social Values ​​(SVO) ;

[0016] Traffic density factor Weighted Social Values ​​(SVO) ;

[0017] Used to encourage diversity of social values ​​(SVO) item ;

[0018] L2 regularization term used to prevent the encouragement of extreme social value orientation (SVO) values. ;

[0019] in, ; This indicates that diversity is maintained by encouraging the variance of Social Values ​​Orientation (SVO). This represents the social value orientation value of connected autonomous vehicles (i).

[0020] Preferably, in the Social Intention Perception Module (SIPM), the Social Value Orientation (SVO) value of HDV... It is statically initialized, and its value follows a preset bimodal distribution to simulate two typical human driving behavior preferences: competitive and cooperative.

[0021] Preferably, the Social Value Orientation (SVO) value is directly applied to the reward function calculation of the multi-agent reinforcement learning decision-making module to balance individual rewards. and regional rewards The specific calculation formula is as follows:

[0022] in, The Social Value Orientation (SVO) value for connected autonomous vehicles (CAVs).

[0023] Preferably, the Social Value Orientation (SVO) value is used to calculate the overall traffic priority of vehicles in conflict zones (such as ramp merging areas), and the calculation formula is as follows:

[0024] in, The location priority is calculated based on the distance from the vehicle to the conflict point. Speed ​​priority is calculated based on the ratio of the vehicle's current speed to the speed limit. , , The weighting factor is used to indicate the social value orientation (SVO) value. Vehicles with higher social value orientation (SVO) values ​​(more altruistic) will have their overall priority appropriately reduced to encourage yielding behavior.

[0025] Preferably, the hybrid expert model in the hybrid expert prediction module (MEPM) includes:

[0026] Multiple independent expert networks, each trained to focus on learning a typical driving behavior pattern, including multiple of the following: conservative, aggressive, cooperative, competitive, following, and lane-changing.

[0027] A gated network receives encoded environmental features and dynamically activates multiple expert networks with the highest weights through a Softmax function and a top-k selection strategy.

[0028] A multi-task output head includes a trajectory prediction head for predicting future multi-step trajectories and a risk assessment head for outputting the probability of collisions between vehicles.

[0029] Preferably, the Hybrid Expert Prediction Module (MEPM) is integrated into the multi-agent reinforcement learning decision-making module and serves as a safety supervisor. Its workflow includes:

[0030] At each decision step, environmental state characteristics are received in real time, and future trajectory predictions and collision risks are output.

[0031] When a potential conflict is predicted, a penalty is imposed on the agent's decision through a conflict-aware reward function, wherein the reward function is:

[0032]

[0033] in, , These represent collision penalty and conflict penalty, respectively. , These represent the penalty weights for collisions and conflicts, respectively.

[0034] At the same time, by combining the overall traffic priority of vehicles, the action masking mechanism is used to restrict low-priority vehicles from performing unsafe actions that may cause conflicts.

[0035] The action masking mechanism is configured such that when vehicle i has a conflict risk with vehicle j with higher priority and action a belongs to the set of unsafe actions, the logical value of the action is set to zero to reduce the probability of the policy network selecting the action.

[0036] Preferably, the preset bimodal distribution is a bimodal Gaussian distribution centered at -0.8 and +0.8 with a standard deviation of 0.2.

[0037] Preferably, the feature vector is 27-dimensional, consisting of 19-dimensional vehicle core state, 5-dimensional neighbor vehicle information, 1-dimensional lane information, 1-dimensional action information, and 1-dimensional social value orientation value; the future multi-step trajectory is the trajectory for the next 10 steps.

[0038] Preferably, the action masking mechanism is implemented in the following way: in the calculation of the action probability distribution of the policy network, the logical value of the insecure action is zeroed out before the calculation of the Softmax function.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] This invention enables connected autonomous vehicles to dynamically learn and adjust their social value orientation through a social intention perception module, thereby adaptively balancing self-interest and altruism to generate more human-like and more acceptable driving strategies. At the same time, a hybrid expert prediction module transforms the decision-making mode from passive response to proactive coordination through multi-step trajectory prediction and conflict risk assessment, enabling the anticipation and resolution of potential conflicts before they occur. The two work together to form a new decision-making closed loop with social intelligence and foresight, ultimately improving the overall system's collaborative smoothness, safety, and traffic efficiency in complex mixed traffic scenarios.

[0041] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the MAPPO-SOM framework of the present invention;

[0043] Figure 2 This is a framework diagram of the Social Intent Perception Module (SIPM) of this invention;

[0044] Figure 3 This is a framework diagram of the Hybrid Expert Prediction Module (MEPM) of the present invention. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] Please see Figures 1-3The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model in this invention aims to achieve safe, efficient and human-like autonomous driving decision-making by simulating human social behavior preferences and making forward-looking conflict prediction and coordination.

[0047] System initialization and environment setup:

[0048] First, a highway ramp merging scenario is constructed as the simulation environment in the Highway-Env simulator. This environment includes one main lane and one merging ramp. The total number N of interacting agents in the environment is a configurable parameter, including a certain proportion of connected autonomous vehicles (CAVs) and human-driven vehicles (HDVs). The state information of all vehicles, including position, speed, acceleration, heading angle, etc., constitutes the state space S of the environment. Each agent can only acquire its own local observations at each step. This aligns with the specification of the partially observable Markov decision process (Dec-POMDP).

[0049] The collaborative decision-making function of this invention is implemented by a decision-making module based on multi-agent reinforcement learning, which is constructed based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm framework. MAPPO adopts a centralized training distributed execution (CTDE) paradigm, and its architecture includes an actor network deployed for each CAV and a centrally trained critic network. Under this framework, the Social Intent Awareness Module (SIPM) and Hybrid Expert Prediction Module (MEPM) described below are deeply integrated: the actor network generates action policies based on local observations and its own SVO value; the critic network uses global state information (including future trajectory and risk information provided by MEPM) to more accurately evaluate state value to guide the updates of the actor network.

[0050] Module 1: Execution Process of the Social Intention Perception Module (SIPM)

[0051] This module serves as the social behavior brain of the entire system, responsible for assigning and applying Social Value Orientation (SVO) to all vehicles in the environment. Its workflow is as follows:

[0052] HDV's SVO initialization:

[0053] At the start of the simulation, each human-driven vehicle (HDV) is assigned a fixed SVO value $\varphi_j^{hdv}$. To simulate the heterogeneity of human driving behavior, this value is sampled from a pre-defined bimodal Gaussian distribution. Specifically, this distribution is a mixture of two unimodal Gaussian distributions, one centered at -0.8 (simulating competitive drivers) and the other centered at +0.8 (simulating cooperative drivers), both with a standard deviation of 0.2. This naturally results in two typical HDV behavior patterns in the environment: one inclined to cut in and the other inclined to yield.

[0054] CAV's SVO dynamic learning:

[0055] Each connected autonomous vehicle (CAV) is equipped with a separate multilayer perception machine (MLP) network as its SVO prediction network. The input to this network is the current environmental state of agent i. The output is a continuous value between (-1, 1). This represents the social preferences of the CAV at the current moment.

[0056] The training of this SVO network is performed simultaneously with the MAPPO Actor-Critic network, and its loss function is... The design guides the SVO value to evolve in a direction that benefits overall system synergy. The loss function is as follows:

[0057]

[0058] in, It is the advantage function, which guides SVO to adjust in the direction of improving the quality of individual decision-making; It is an average team reward, which encourages SVO to develop in a direction that benefits the collective. It is the normalized current traffic density that makes CAVs more inclined to cooperate during congestion; This is the mean SVO of CAVs within the batch. This term encourages SVO diversity and prevents all CAVs from converging to the same behavior pattern. The last term is L2 regularization to prevent SVO values ​​from tending to extremes, where the coefficients are 0.3, 0.2, 0.1, and... These are preset weight hyperparameters used to balance the relative importance of team rewards, traffic density, diversity incentives, and regularization penalties in the loss function. Their specific values ​​are set before training via grid search or experience; the loss function... This enables CAVs to learn to dynamically adjust their social value orientation according to environmental conditions, thereby achieving better multi-agent collaboration.

[0059] Application of SVO in decision-making:

[0060] The generated SVO value is directly used to adjust the decision logic of CAV.

[0061] Reward Balancing: At each step, the system calculates a balanced reward. The reward is given by an individual. (such as driving efficiency, comfort) and regional rewards (Such as overall traffic flow efficiency) weighted, with SVO as the weighting coefficient:

[0062]

[0063] when When approaching +1 (altruism), CAV focuses more on area rewards and tends to yield; when When the value approaches -1 (selfishness), the individual focuses more on personal rewards and behaves more aggressively.

[0064] Priority Adjustment: In the merging conflict zone, a comprehensive passage priority is calculated for each vehicle. This priority depends not only on the physical state but also incorporates social intentions:

[0065] in, It is based on the location priority according to the distance to the merging point, with higher priority for closer locations; It's based on speed priority. The key is subtracting... This means that vehicles with high SVO (cooperative) will proactively lower their priority to encourage yielding, while vehicles with low SVO (competitive) will relatively increase their priority to encourage going first. , and Preset, positive-zero normalized weighting coefficients are used to adjust the contribution ratios of position, speed, and social value orientation in the comprehensive priority calculation, respectively, to satisfy... Alternatively, it can be determined through parameter tuning. This mechanism allows CAVs to proactively adjust their behavior based on their social intentions, thereby enhancing the system's cooperation and effectively reducing potential collision risks.

[0066] Module 2: Execution Process of the Hybrid Expert Prediction Module (MEPM)

[0067] This module is the system's "forward-looking eye," responsible for predicting the future and assessing risks. Its execution is divided into two phases: pre-training and deployment.

[0068] 1. Data collection and pre-training:

[0069] Data Acquisition: Large-scale traffic flow simulations were conducted in a simulation environment to systematically collect vehicle interaction data. Each frame of data was processed into a 27-dimensional feature vector, including: 19-dimensional vehicle core state (position, speed, acceleration, etc.), 5-dimensional nearest neighbor vehicle information, 1-dimensional lane information, 1-dimensional discrete action, and 1-dimensional SVO value. To ensure data diversity and representativeness, data acquisition was conducted under different traffic densities (e.g., low, medium, and high density) and different HDV behavior ratios, collecting over 100,000 steps of interaction data. The collected raw data underwent cleaning and normalization to eliminate the influence of dimensions and improve the stability and convergence speed of model training.

[0070] Model Architecture and Training: Building a hybrid expert model, whose core components include:

[0071] Expert Networks: Six independent expert networks were designed, each focusing on learning a typical driving behavior pattern, such as conservative, aggressive, cooperative, competitive, following, and lane changing.

[0072] Gated networks: First, a feature encoder (an MLP) maps the 27-dimensional input feature x to 128-dimensional hidden features. : The gating network then... Calculate the weight of each expert : To improve efficiency, a top-2 gating strategy is adopted, which means that only the two experts with the highest weights are activated, expressed as: In the above model, , , , For the weight matrix and bias vector of the feature encoder; , The weights and biases of the gated network; This represents the processing result of the i-th expert network on input x.

[0073] Multi-task output: The model's final output is a weighted sum of two independent "heads". The trajectory prediction head outputs the mean of the trajectory over the next T=10 steps. and variance Its loss is the negative log-likelihood: The risk assessment head outputs a scalar r, predicting the probability of a collision occurring in the future, with its loss being the mean squared error. The total loss is the weighted sum of the two: ,in These are hyperparameters. After the model has been fully trained, the parameters are frozen for ensemble preparation; in the loss function, Represents the coordinates of the actual future trajectory recorded during the data collection phase, serving as a supervisory signal for trajectory prediction; Binary collision labels (1 for a collision, 0 for no collision) serve as monitoring signals for risk assessment; in the total loss The hyperparameters that balance the weights of trajectory prediction and risk assessment tasks are related to the regularization coefficient in the SVO loss function. Irrelevant.

[0074] 2. Deployment, integration, and security monitoring:

[0075] Integrate the pre-trained MOE model into the simulation loop of the MARL framework.

[0076] Real-time prediction: At each decision step, the current 27-dimensional feature vector is input into the MOE model, and the model outputs the trajectory predictions for all vehicles for the next 10 steps in real time. And the probability of collision risk r.

[0077] Multi-level safety supervision:

[0078] To combine MOE prediction capabilities with traditional security rules, a three-layer protection system is adopted.

[0079] First, the forward-looking prediction capability of the MOE model is utilized to generate multi-step trajectories after the action is executed, providing a data foundation for conflict detection. The results include geometric location information and probabilistic uncertainty, supporting multi-granularity risk assessment.

[0080] Secondly, a collision detection mechanism is implemented, based on physical geometry-based distance collision detection to ensure vehicle safety, represented as:

[0081]

[0082] in, Indicates vehicle The rotating rectangular boundary is detected by efficient calculation using the separation axis theorem, enabling adaptive safety monitoring.

[0083] Finally, the conflict detection function was set. When a potential conflict is detected, a corresponding reward or penalty will be generated. At the same time, the system will correct the action by combining the action mask with the vehicle's priority. Low-priority vehicles are restricted from changing lanes and accelerating, while high-priority vehicles can still perform all actions to ensure smooth traffic and safety.

[0084] Conflict-aware reward: Introduce a penalty term into the reward function of MARL:

[0085]

[0086] in These represent basic rewards, including standard reinforcement learning rewards such as encouraging efficient passage and maintaining comfort. Used to penalize actual collisions. Used to penalize potential conflicts predicted by the Hybrid Expert Prediction Module (MEPM). and For penalty weights.

[0087] Action masking mechanism: This is the last and most direct line of defense for security. The system defines an action masking function:

[0088]

[0089] in Represents a set of neighbors. The set of actions deemed unsafe in the current state. It is a conflict judgment function that returns true when there is a risk of intersection between the future trajectories of vehicles i and j based on the output of the Hybrid Expert Prediction Module (MEPM) (such as the minimum distance being less than the safety threshold).

[0090] When vehicle i has a prediction conflict with a higher-priority vehicle j, and the planned action a belongs to the set of unsafe actions. (e.g., when accelerating or changing lanes) This makes the logical value of the action... It is set to zero before the Softmax calculation of the policy network, as implemented below:

[0091]

[0092] in It is a state-action value function calculated by the commentator network. These are the raw motion logic values ​​output by the actor network before the Softmax function is applied;

[0093] This greatly reduces the probability of choosing this dangerous action and forces low-priority vehicles to adopt a yielding strategy.

[0094] System collaborative workflow:

[0095] At each time step of the simulation, the system initiates a collaborative decision-making loop. First, the Social Intent Perception Module (SIPM) updates the Social Value Orientation (SVO) values ​​of all vehicles, injecting social intent into the decision. Next, the policy network (Actor) in the decision module generates a preliminary action intent based on the current environmental state and the CAV's own SVO value. This preliminary intent, along with the environmental state and SVO value, is immediately encapsulated into a 27-dimensional feature vector and input to the Hybrid Expert Prediction Module (MEPM). The MEPM, acting as the system's "foresightful eye," quickly predicts the future multi-step trajectory and collision risk after executing this preliminary action. Subsequently, the safety supervisor of the decision module is activated. It integrates the MEPM's prediction results and the vehicle priorities adjusted by the SVO, performing two layers of intervention: first, reshaping the preliminary intent through a conflict-aware reward function; and second, directly rejecting high-risk actions when necessary through an action masking mechanism. Finally, the decision module outputs a collaborative driving decision that has undergone social intent trade-offs and forward-looking safety verification, thus forming a closed-loop decision-making system integrating social intelligence and foresight.

Claims

1. A hybrid traffic collaborative decision-making system based on social value orientation and hybrid expert models, characterized in that: include: The Social Intent Perception Module (SIPM) is used to generate and apply Social Value Orientation (SVO) values ​​for connected autonomous vehicles (CAVs) and human-driven vehicles (HDVs) in a mixed traffic environment; wherein, the social value orientation (SVO) values ​​of the connected autonomous vehicles (CAVs) are dynamically learned through a trainable network, and the social value orientation (SVO) values ​​of the HDVs follow a preset bimodal distribution static initialization. The Hybrid Expert Prediction Module (MEPM) consists of a pre-trained hybrid expert model that receives feature vectors including vehicle state, social value orientation (SVO) value, and environmental context, and outputs predictions of future multi-step trajectories and collision risk assessments between vehicles. The decision-making module integrates the SIPM and MEPM, and is configured to use the Social Value Orientation (SVO) value to balance rewards and adjust traffic priorities, and to perform safety supervision and action correction based on the prediction results of the MEPM, so as to output cooperative driving decisions.

2. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model as described in claim 1, characterized in that, In the Social Intent Perception Module (SIPM), the Social Value Orientation (SVO) value of the connected autonomous vehicle (CAV) is dynamically generated through an independent multilayer perceptron network, and the loss function of this network is... The following weighting factors were combined: From the advantage function Weighted Social Values ​​(SVO) ; Team average reward Weighted Social Values ​​(SVO) ; Traffic density factor Weighted Social Values ​​(SVO) ; Used to encourage diversity of social values ​​(SVO) item ; L2 regularization term used to prevent the encouragement of extreme social value orientation (SVO) values. ; in, ; This indicates that diversity is maintained by encouraging the variance of Social Values ​​Orientation (SVO). This represents the social value orientation value of connected autonomous vehicles (i).

3. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model as described in claim 1, characterized in that, In the Social Intention Perception Module (SIPM), the Social Value Orientation (SVO) value of HDV... It is statically initialized, and its value follows a preset bimodal distribution to simulate two typical human driving behavior preferences: competitive and cooperative.

4. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model as described in claim 1, characterized in that, The Social Value Orientation (SVO) value is directly applied to the reward function calculation of the multi-agent reinforcement learning decision-making module to balance individual rewards. and regional rewards The specific calculation formula is as follows: in, The Social Value Orientation (SVO) value for connected autonomous vehicles (CAVs).

5. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model according to claim 1, characterized in that, The Social Value Orientation (SVO) value is used to calculate the overall traffic priority of vehicles in conflict zones (such as ramp merging areas), and its calculation formula is as follows: in, The location priority is calculated based on the distance from the vehicle to the conflict point. Speed ​​priority is calculated based on the ratio of the vehicle's current speed to the speed limit. , , The weighting factor is used to indicate the social value orientation (SVO) value. Vehicles with higher social value orientation (SVO) values ​​(more altruistic) will have their overall priority appropriately reduced to encourage yielding behavior.

6. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model according to claim 1, characterized in that, The hybrid expert model in the Hybrid Expert Prediction Module (MEPM) includes: Multiple independent expert networks, each trained to focus on learning a typical driving behavior pattern, including multiple of the following: conservative, aggressive, cooperative, competitive, following, and lane-changing. A gated network receives encoded environmental features and dynamically activates multiple expert networks with the highest weights through a Softmax function and a top-k selection strategy. A multi-task output head includes a trajectory prediction head for predicting future multi-step trajectories and a risk assessment head for outputting the probability of collisions between vehicles.

7. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model as described in claim 6, characterized in that, The Hybrid Expert Prediction Module (MEPM) is integrated into the multi-agent reinforcement learning decision-making module and serves as a safety supervisor. Its workflow includes: At each decision step, environmental state characteristics are received in real time, and future trajectory predictions and collision risks are output. When a potential conflict is predicted, a penalty is imposed on the agent's decision through a conflict-aware reward function, wherein the reward function is: in, , These represent collision penalty and conflict penalty, respectively. , These represent the penalty weights for collisions and conflicts, respectively. At the same time, by combining the overall traffic priority of vehicles, the action masking mechanism is used to restrict low-priority vehicles from performing unsafe actions that may cause conflicts. The action masking mechanism is configured such that when vehicle i has a conflict risk with vehicle j with higher priority and action a belongs to the set of unsafe actions, the logical value of the action is set to zero to reduce the probability of the policy network selecting the action.

8. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model according to claim 3, characterized in that, The preset bimodal distribution is a bimodal Gaussian distribution centered at -0.8 and +0.8 with a standard deviation of 0.

2.

9. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model according to claim 1, characterized in that, The feature vector is 27-dimensional, consisting of 19-dimensional vehicle core state, 5-dimensional neighbor vehicle information, 1-dimensional lane information, 1-dimensional action information, and 1-dimensional social value orientation value; the future multi-step trajectory is the trajectory for the next 10 steps.

10. The hybrid traffic collaborative decision-making based on social value orientation and hybrid expert model according to claim 7, characterized in that, The action masking mechanism is implemented in the following way: in the calculation of the action probability distribution of the policy network, the logical value of the insecure action is preceded by zero in the calculation of the Softmax function.