Intelligent agent cognitive decision-making integrated method and system based on wm and mlm
Patent Information
- Application Number
- CN202610645707.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-05-12
AI Technical Summary
[0004]本申请提供了基于WM和MLM的智能体认知决策一体化方法及系统,改善了现有技术中决策模块串联执行导致误差累积、对交互场景适应性差、无法在线校准,造成自动驾驶决策可靠性不足的技术问题
本申请技术方案通过提供的基于WM和MLM的智能体认知决策一体化方法,首先,获取自车传感数据与路侧单元广播的交通参与者轨迹片段,通过交通交互态势编码网络进行掩码自注意力交叉编码,生成以自车为中心的交互对象博弈意向分布图,能够从多源异构的感知数据中提取与自车决策强相关的交通参与者及其交互意向,有效减少了无关信息对决策过程的干扰,从而提升了复杂交通场景下关键目标筛选的准确性。
Smart Images

Figure CN122166152B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to an integrated method and system for intelligent agent cognition and decision-making based on WM and MLM. Background Technology
[0002] Autonomous driving refers to the technology of vehicles perceiving their surroundings, understanding traffic scenarios, and autonomously planning paths and controlling motion through onboard sensors. Among these technologies, a World Model (WM) is a model capable of learning the internal representations of the external environment and predicting its future state; a Multimodal Large Model (MLM) typically refers to a large neural network model capable of simultaneously processing multiple modalities of information, such as text, images, audio, and video, and achieving cross-modal understanding and generation. These models can be applied to cognitive decision-making tasks for intelligent agents in complex and dynamic environments, such as autonomous driving. Traditional autonomous driving methods typically employ a modular pipeline architecture, including independent perception, prediction, planning, and control modules. These modules are executed sequentially in a fixed order, with rule-based models or simple machine learning models serving as the core driver.
[0003] However, traditional pipelined methods have significant shortcomings. On the one hand, information transmission between modules is delayed and errors are prone to accumulate, leading to a decrease in the robustness of overall decision-making. On the other hand, traditional methods have limited ability to model highly interactive dynamic traffic scenarios and struggle to effectively handle game-theoretic behavior among multiple traffic participants. Furthermore, traditional methods lack an online calibration mechanism to dynamically adjust internal model parameters based on actual execution results, failing to continuously leverage historical interaction experience to improve decision-making performance. Therefore, there is an urgent need for an agent-based cognitive decision-making method capable of co-optimizing perception, cognitive modeling, and action strategies. Summary of the Invention
[0004] This application provides an integrated method and system for intelligent agent cognitive decision-making based on WM and MLM, which improves the technical problems of insufficient reliability of autonomous driving decision-making caused by the serial execution of decision modules in the prior art, such as error accumulation, poor adaptability to interactive scenarios, and inability to be calibrated online.
[0005] This application discloses the following technical solution: In a first aspect, this application provides an integrated cognitive decision-making method for intelligent agents based on WM and MLM, the method comprising: The system acquires trajectory fragments of traffic participants from vehicle sensor data and roadside unit broadcasts, and performs masked self-attention cross-coding through a traffic interaction situation coding network to generate a distribution map of interactive object game intentions centered on the vehicle. Using the game intention distribution map of the interactive objects as the dynamic search boundary, the temporal game inference network is driven to perform multi-step inference to generate the collision time margin decay sequence and equivalent traffic deceleration corresponding to the candidate driving actions of the vehicle. Based on the collision time margin decay sequence and the equivalent passage deceleration, a conflict urgency arbitration logic is constructed. When the minimum collision time margin is lower than the linear mapping threshold of the equivalent passage deceleration, a risk avoidance instruction is output; otherwise, a compliance interaction instruction is output. Based on the risk avoidance instructions and the compliance interaction instructions, combined with the vehicle sensor data, the executable driving behavior at the current moment is generated; The actual interaction trajectory after the execution of the executable driving behavior is collected, and the cumulative deviation between the actual interaction trajectory and the inferred trajectory is used as the driving signal to reverse calibrate the temporal game inference network and the traffic interaction situation coding network.
[0006] Secondly, this application provides an integrated cognitive decision-making system for intelligent agents based on WM and MLM, the system comprising: The situational awareness module, deployed on the autonomous vehicle computing platform, is used to acquire autonomous vehicle sensor data and traffic participant trajectory segments broadcast by roadside units. It performs masked self-attention cross-coding through a traffic interaction situational coding network to generate a distribution map of interactive object game intentions centered on the autonomous vehicle. The game inference module is used to drive the time-series game inference network to perform multi-step inferences using the game intention distribution map of the interactive objects as a dynamic search boundary, and generate the collision time margin decay sequence and equivalent traffic deceleration corresponding to the candidate driving actions of the vehicle. The conflict arbitration module is used to construct conflict urgency arbitration logic based on the collision time margin decay sequence and the equivalent traffic deceleration. When the minimum collision time margin is lower than the linear mapping threshold of the equivalent traffic deceleration, a risk avoidance instruction is output; otherwise, a compliance interaction instruction is output. The behavior generation module is used to generate executable driving behaviors at the current moment based on the risk avoidance instructions and the compliance interaction instructions, combined with the vehicle sensor data. The online calibration module is used to collect the actual interaction trajectory after the execution of the executable driving behavior, and use the cumulative deviation between the actual interaction trajectory and the inferred trajectory as the driving signal to reverse calibrate the time-series game inference network and the traffic interaction situation coding network.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: The technical solution of this application provides an integrated intelligent agent cognitive decision-making method based on WM and MLM. First, it acquires the trajectory segments of traffic participants broadcast by roadside units from vehicle sensor data. Then, it performs masked self-attention cross-coding through a traffic interaction situation coding network to generate a distribution map of interactive object game intentions centered on the vehicle. This method can extract traffic participants and their interaction intentions that are strongly related to the vehicle's decision-making from multi-source heterogeneous perception data, effectively reducing the interference of irrelevant information on the decision-making process, thereby improving the accuracy of key target selection in complex traffic scenarios.
[0008] Furthermore, using the game intention distribution map of the interactive objects as a dynamic search boundary, the time-series game inference network is driven to perform multi-step inference, generating collision time margin decay sequences and equivalent traffic decelerations corresponding to candidate driving actions of the self-vehicles. The game intention information is quantified into a computable inference boundary, enabling the inference process to dynamically adapt to the behavioral tendencies of different interactive participants, thereby improving the pertinence of future state prediction and inference efficiency.
[0009] Furthermore, based on the collision time margin decay sequence and the equivalent traffic deceleration, a conflict urgency arbitration logic is constructed. When the minimum collision time margin is lower than the linear mapping threshold of the equivalent traffic deceleration, a risk avoidance instruction is output; otherwise, a compliance interaction instruction is output. This establishes an adaptive arbitration mechanism based on continuous safety margin, avoiding the bias of conservative or aggressive decision-making caused by fixed thresholds, thereby achieving a dynamic balance between safety and traffic efficiency.
[0010] Furthermore, based on the risk avoidance instructions and the compliance interaction instructions, combined with the vehicle sensor data, the executable driving behavior at the current moment is generated, and the arbitration result and the current motion state are directly mapped into specific control quantities, eliminating the semantic gap between planning and control in the traditional pipeline, thereby shortening the reaction delay from perception to execution.
[0011] Finally, the actual interaction trajectory after the execution of the executable driving behavior is collected. The cumulative deviation between the actual interaction trajectory and the inferred trajectory is used as the driving signal to reverse-calibrate the temporal game inference network and the traffic interaction situation coding network. This enables the decision-making system to make online adaptive adjustments based on the actual interaction results, continuously narrowing the gap between the prediction and the real environment, thereby improving the robustness and self-evolution capability of the system in long-term operation.
[0012] In summary, the technical solution of this application achieves end-to-end collaborative optimization and integrated decision-making of perception, cognitive modeling, and action strategies. By generating an interactive object game intention distribution map, generating collision time margin and equivalent traffic deceleration through multi-step deduction, outputting instructions through adaptive arbitration logic, generating executable driving behaviors, and performing online reverse calibration based on real trajectories, it effectively improves the technical problems in the existing modular pipeline architecture, such as information transmission delay and error accumulation, limited ability to model dynamic interactive game scenarios, and lack of online adaptive calibration mechanism, which lead to reduced decision robustness in complex traffic environments and difficulty in meeting real-time and safety requirements. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating the integrated cognitive decision-making method for intelligent agents based on WM and MLM provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of an integrated cognitive decision-making system for intelligent agents based on WM and MLM, provided in an embodiment of this application.
[0015] The components represented by each number in the attached diagram are explained as follows: Situational Awareness Module 11, Game Theory Module 12, Conflict Arbitration Module 13, Behavior Generation Module 14, and Online Calibration Module 15. Detailed Implementation
[0016] This application provides an integrated cognitive decision-making method and system for intelligent agents based on WM and MLM, which addresses the technical problems in existing technologies, such as information transmission delays and error accumulation in modular pipeline architectures, limited ability to model dynamic interactive game scenarios, and lack of online adaptive calibration mechanisms, which lead to reduced decision robustness in complex traffic environments and difficulty in meeting real-time and security requirements.
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that the numerical values in the embodiments are for illustrative purposes only and do not constitute a limitation on this application.
[0018] Example 1, as shown in the appendix Figure 1 As shown, this application provides an integrated cognitive decision-making method for intelligent agents based on WM and MLM, the method comprising the following steps: S100: Acquire the trajectory segments of traffic participants from vehicle sensor data and roadside unit broadcasts, and perform masked self-attention cross-coding through a traffic interaction situation coding network to generate a distribution map of interactive object game intentions centered on the vehicle. In this embodiment of the application, in a scenario where an autonomous vehicle needs to engage in real-time interactive game with multiple traffic participants, in order to accurately extract participants and their behavioral intentions that are strongly related to the vehicle's decision-making from heterogeneous multi-source perception data, it is necessary to simultaneously acquire the vehicle's motion state and the surrounding trajectory data broadcast by the roadside unit, and to mine the implicit interaction relationship between the vehicle and each participant through an encoding network. This solves the technical problem that traditional perception modules only output target position and speed and lack quantitative modeling of interaction intentions, making it difficult for subsequent decisions to distinguish between key game objects and non-key interference objects.
[0019] Step S100 in the method provided in this application embodiment includes: The vehicle sensing data includes at least the vehicle position sequence, the vehicle speed sequence, and the vehicle heading angle sequence, and the traffic participant trajectory segment includes at least the participant position sequence and the participant speed sequence. Using the absolute timestamp as the alignment reference, the vehicle position sequence and the participant position sequence are synchronized in time, and then a synchronized position pair sequence is generated by matching frames after synchronization. Using the vehicle position in the first synchronized frame as the origin and the vehicle heading angle in the first synchronized frame as the positive vertical axis, a vehicle coordinate system is established, and the participant position sequence is transformed frame by frame into the vehicle coordinate system. Traverse the time intervals of adjacent frames in the participant position sequence transformed to the vehicle coordinate system, and mark the intervals where the time interval exceeds the roadside unit broadcast period as communication interruption segments; Based on the participant positions and speeds of the previous frame before the communication interruption, the frames within the interruption segment are extrapolated and completed at a uniform speed, and a mask identifier is added to the completed frames. The participant position sequence after coordinate transformation and the addition of masked identifiers, the participant velocity sequence after time synchronization, and the synchronized vehicle position sequence, vehicle velocity sequence, and vehicle heading angle sequence are collectively used as the input data for the traffic interaction situation coding network. A detailed explanation follows: In this embodiment, the vehicle position sequence refers to a series of coordinate points of the autonomous vehicle in the global coordinate system recorded at a fixed sampling period, in meters (m). The vehicle speed sequence refers to a series of instantaneous speed values at the same timestamp as the vehicle position sequence, in m / s. The vehicle heading angle sequence refers to the sequence of angles between the vehicle's longitudinal axis and the north direction of the global coordinate system, in degrees. The participant position sequence refers to the sequence of coordinate points of surrounding vehicles or pedestrians broadcast by the roadside unit in the global coordinate system, in meters (m). The participant speed sequence refers to the sequence of instantaneous speeds at the same timestamp as the participant position sequence, in m / s. The absolute timestamp refers to a time identifier unified to Coordinated Universal Time (UTC), with a precision of milliseconds. Uniform extrapolation refers to an interpolation method that assumes the target maintains the same velocity vector as the previous frame during the missing time period for position estimation. The mask identifier refers to a flag bit attached to the completed data, used to indicate that the data of this frame was not actually collected but generated through completion.
[0020] In this step, firstly, to establish a basic motion state description of the vehicle and participants to support subsequent spatiotemporal alignment, the vehicle's 3D coordinate sequence, instantaneous velocity sequence, and heading angle sequence must be recorded simultaneously at each sampling moment. Simultaneously, the coordinate and velocity sequences of each participant are also recorded. This yields complete multidimensional state input for the vehicle and participants. For example, the vehicle's position sequence is [(0m,0m),(0.5m,0m)], velocity sequence [1m / s,1m / s], and heading angle sequence [0°,0°]; the participant's position sequence is [(10m,0m),(10.5m,0m)], and velocity sequence is [2m / s,2m / s]. This constraint provides the necessary and sufficient data fields for subsequent time synchronization, coordinate transformation, and interactive coding, avoiding process interruptions caused by data loss.
[0021] Furthermore, to eliminate the time misalignment between the vehicle sensors and roadside units caused by their independent clock operation, an absolute timestamp is used as a unified benchmark. Each frame in the vehicle position sequence is matched with frames in the participant position sequence that have the same absolute timestamp, generating a one-to-one corresponding sequence of synchronized position pairs. This results in strictly time-aligned bilateral trajectory point pairs, for example, pairing the vehicle position (1m, 0m) with the participant position (12m, 0m) at t=1.000s. This method solves the time deviation problem caused by different sampling start times or clock drift in multi-source heterogeneous data, ensuring that every moment in subsequent simulations corresponds to true physical synchronization.
[0022] Furthermore, to simplify interactive calculations by converting the absolute position in the global fixed coordinate system into a relative description with the vehicle itself as the origin, a moving vehicle coordinate system is established using the vehicle's position in the first frame after synchronization as the origin and the vehicle's heading angle in the first frame as the positive direction of the vertical axis. Each frame in the participant's position sequence is then transformed to this coordinate system, yielding the participant's longitudinal distance and lateral offset relative to the vehicle. For example, if the vehicle's position in the first frame is (0m, 0m) and the heading angle is 0°, the participant's global position (10m, 5m) is transformed to (10m, 5m). If the heading angle is 30°, the transformed coordinates need to be calculated using a rotation matrix. This transformation ensures that the vehicle's own motion no longer affects the relative position representation of the participants, significantly reducing the difficulty for the subsequent encoding network to extract relative relationships from absolute coordinates.
[0023] Furthermore, to identify whether packet loss or delay occurs in communication between the roadside unit and the vehicle, and to accurately determine the time period of missing trajectory data, it is necessary to traverse the participant position sequence that has been transformed to the vehicle coordinate system, calculate the time interval between each pair of adjacent sampling frames, and compare this interval with the standard broadcast period of the roadside unit. When the interval is strictly greater than the broadcast period, the interval is marked as a communication interruption segment, and the start and end frame indices of the interruption segment can be obtained. For example, if the time interval between the 3rd and 4th frames in a participant position sequence is 0.5s, while the broadcast period is 0.1s, then the interval between the 3rd and 4th frames is marked as an interruption segment. This marking provides accurate spatiotemporal positioning for subsequent data completion and confidence assessment, avoiding blindly processing all missing frames uniformly.
[0024] Furthermore, in order to maintain the continuity of participant trajectories as much as possible during communication interruptions to support stable computation of the temporal network, it is necessary to perform linear extrapolation interpolation on each frame within the interruption segment based on the participant position and velocity of the previous frame before the interruption segment. Assuming that the participants maintain uniform linear motion during the missing period, their position coordinates are completed by assuming uniform linear motion during the missing period. A special mask identifier is added to all the completed frames to obtain a complete sequence of participant positions with confidence labels. For example, if the position of the previous frame before the interruption segment is (10m, 5m) and the velocity is (2m / s, 0m / s), the broadcast period is 0.1s, the interruption length is 0.5s, and a total of 5 frames are missing, then the positions of the completed frames are (10.2m, 5m), (10.4m, 5m), (10.6m, 5m), (10.8m, 5m), and (11.0m, 5m), and a mask identifier "1" is added to each frame. This method ensures that the input sequence has the full length to meet the timing requirements of the encoding network, and also informs the network of the low confidence of the supplementary data through mask identification, thereby reducing the interference of false information on attention weight learning.
[0025] Finally, to unify all the data processed through synchronization, coordinate transformation, missing data completion, and masking into a format directly readable by the coding network, the participant position sequence (with coordinate transformation and masking), the participant velocity sequence (synchronized in time), and the synchronized vehicle position sequence, vehicle velocity sequence, and vehicle heading angle sequence are concatenated at the same time index into a multi-dimensional feature tensor. This tensor serves as the input data for the traffic interaction situation coding network, resulting in a well-structured, physically meaningful, and time-aligned multi-source input matrix. For example, the input data includes eight feature dimensions: vehicle position x, vehicle position y, vehicle velocity, vehicle heading angle, participant position x, participant position y, participant velocity, and masking. This integration step completes the entire data preparation process from raw sensing data to the network input, ensuring that the subsequent coding network can correctly extract the interaction features between the vehicle and each participant at each moment, avoiding network calculation errors or performance degradation caused by chaotic input formats.
[0026] Step S100 in the method provided in this application embodiment further includes: The synchronized vehicle position sequence, vehicle speed sequence, and vehicle heading angle sequence are extracted from the input data of the traffic interaction situation coding network and combined into a vehicle motion state segment. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle motion state segment, and use the masked vehicle motion state segment as the query sequence. Using the participant position sequence after coordinate transformation and the added mask label, and the participant speed sequence after time synchronization as key-value pairs, the traffic interaction situation coding network is input to output the attention response intensity sequence of each participant to their own vehicle. The construction steps of the traffic interaction situation coding network include: Collect historical interaction trajectory data at urban intersections, extract synchronized vehicle position sequence, vehicle speed sequence and vehicle heading angle sequence as vehicle historical motion data, and extract participant position sequence after coordinate transformation and participant speed sequence after time synchronization as participant historical interaction data. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle's historical motion data to generate masked vehicle historical motion data; Using the occluded vehicle's historical motion data as the query and the participants' historical interaction data as the key-value pairs, the inverse of the distance collision time between each participant and the vehicle is used as the attention response intensity label to construct a training sample set; A network architecture for a traffic interaction situation coding network is constructed, wherein the network architecture includes an encoder layer, a cross-attention layer, and a feedforward output layer; The mean squared error between the predicted attention response intensity and the attention response intensity label is used as the loss function. The network weights of the network architecture are updated through the backpropagation algorithm until the loss function converges. The trained network architecture is then used as the traffic interaction situation coding network. Traverse the attention response intensity sequence and extract participants whose attention response intensity is greater than an intensity threshold as game stakeholders, wherein the intensity threshold is a predetermined proportion of the maximum value in the attention response intensity sequence; Based on the attention response intensity sequence, game stakeholders are extracted, and the aggressiveness and cooperativeness coefficients of each stakeholder are calculated and combined to form the game intention distribution map of the interacting objects. A detailed explanation follows: In this embodiment, the masking code randomly selects position and velocity data from consecutive frames within a range of 15% to 25% of the input sequence, replacing their values with special masking markers. The preset proportion in this application is 20%. Query sequence and key-value pairs: The query sequence refers to the segment of the vehicle's motion state after masking, serving as the input vector in the encoding network for actively focusing on other information; the key-value pairs refer to the position and velocity data of the participants, with the key used for matching queries and the value used for outputting weighted results. Attention response intensity refers to the scalar score calculated by the output layer of the encoding network regarding the degree of influence of each participant on the vehicle's decision, ranging from 0 to 1. The reciprocal of the distance-collision time count indicates that the collision time refers to the time required for the vehicle and participants to approach zero distance at their current relative speed; its reciprocal is the reciprocal of this time, measured in seconds. The intensity threshold is set to 60% of the maximum value in the attention response intensity sequence, used to filter participants with high interaction intensity.
[0027] The network architecture of the traffic interaction situation coding network is as follows: The network adopts a transformer-based encoder structure, which includes an encoder layer, a cross-attention layer, and a feedforward output layer. The encoder layer consists of two stacked transformer encoder layers, each with four attention heads and a hidden layer dimension of 128. The cross-attention layer calculates attention weights using query sequences and key-value pairs as input. The feedforward output layer is a fully connected network with 64 neurons, outputting attention response intensity values.
[0028] The loss function uses mean squared error loss, and the formula is as follows: ,in The number of unmasked frames. For attention response intensity labels, The values are predicted; occluded frames are not included in the loss calculation.
[0029] In this step, firstly, in order to extract the vehicle motion features from the multi-source input as a query benchmark, it is necessary to separate the vehicle position, speed, and heading angle sequences from the input data of the encoding network and merge them into a vehicle motion state segment, so as to obtain a temporally continuous vehicle trajectory expression.
[0030] Furthermore, in order to simulate perceptual occlusion and train the network's completion capability, it is necessary to apply occlusion masks to the position and velocity of 20% of the consecutive frames in the vehicle's motion state segment. Using the occluded segments as query sequences, the vehicle feature vector carrying the missing label can be obtained.
[0031] Furthermore, in order to calculate the influence weight of each participant on the vehicle, the participant's position sequence and speed sequence are used as key-value pairs and input together with the query sequence into the constructed traffic interaction situation coding network. The cross-attention mechanism within the network outputs the attention response intensity corresponding to each participant, thus obtaining a quantified interaction influence score.
[0032] Furthermore, in constructing this network, historical interaction trajectories at urban intersections are first collected, and historical motion data of vehicles and historical interaction data of participants are extracted separately. The vehicle data is masked with the same proportion as the query, and the participant data is used as the key-value pair, with the reciprocal of the collision time as the supervision label, to form a training sample set. Then, the model is instantiated according to the network architecture defined above, the prediction error is calculated with the aforementioned loss function, and the network weights are iteratively updated until convergence through the backpropagation algorithm to obtain a usable encoding network.
[0033] Furthermore, after completing network inference, in order to reduce the computational complexity of subsequent game deduction, it is necessary to traverse the attention response intensity of all participants, retain only the participants whose intensity value is greater than 60% of the maximum value of the sequence, and mark them as game stakeholders, thus obtaining a subset of key interaction objects.
[0034] Finally, for these game stakeholders, based on the attention response intensity sequence and the pre-defined methods for calculating the aggressive tendency coefficient and the cooperative tendency coefficient, two types of coefficients for each game staker need to be obtained, and the two are combined into an interactive object game intention distribution map to provide input for the subsequent dynamic search boundary.
[0035] For example, the training process of the traffic interaction situation coding network is as follows: The network architecture adopts a 2-layer transformer encoder, with 4 attention heads in each layer, a hidden layer dimension of 128, and a feedforward output layer with 64 neurons. The training dataset contains 10,000 historical interaction samples from urban intersections, each sample containing 50 frames of motion data from the vehicle and the trajectory data of 5 corresponding participants. The position and velocity of 20% of consecutive frames in the historical motion data of the vehicle are randomly occluded, and the reciprocal of the distance to collision time is used as the supervision label. Training is performed for 200 epochs using the Adam optimizer, with an initial learning rate of 0.001, decaying by a factor of 0.1 every 50 epochs, and a batch size of 32. Only the mean squared error between the predicted value and the label of the unoccluded frame is calculated; occluded frames do not participate in gradient backpropagation. Training is stopped early when the validation set loss does not decrease for 10 consecutive epochs. The final trained network weights are used to output the attention response strength during the real-time inference stage.
[0036] Specifically, based on the attention response intensity sequence, game stakeholders are extracted, and the aggressiveness and cooperativeness coefficients of each stakeholder are calculated and combined to form the game intention distribution map of the interaction objects, including: From the input data of the traffic interaction situation coding network, extract the participant speed sequence and the participant heading angle sequence after time synchronization within a preset look-ahead time. Traverse the participant velocity sequence of the game stakeholders and calculate the range of velocity within the preset look-ahead time as the velocity fluctuation amplitude. Traverse the course angle sequence of the participants of the game stakeholders, calculate the ratio of the cumulative sum of the absolute values of the changes in course angle of adjacent frames within the preset look-ahead time to the cumulative number of frames, and use it as the course deviation rate; The speed fluctuation amplitude is mapped to an aggressive tendency coefficient, wherein the aggressive tendency coefficient increases positively with the increase of the speed fluctuation amplitude; The course deviation rate is mapped to a cooperative tendency coefficient, wherein the cooperative tendency coefficient decreases inversely as the course deviation rate increases; The radical tendency coefficient and the cooperative tendency coefficient are fused together to obtain the game intention representation value of the game stakeholders; Based on the correspondence between the identity identifiers of the game stakeholders and the game intention representation values, a game intention distribution map of the interacting objects is constructed. A detailed explanation follows: In this embodiment, the preset look-ahead duration refers to a fixed time length extending from the current moment into the future, used to extract the participant's motion data within this time period, measured in seconds (s). In this application, 3 seconds is used. The velocity fluctuation amplitude refers to the difference between the maximum and minimum values in the participant's velocity sequence within the preset look-ahead duration, measured in m / s, used to quantify the severity of the participant's acceleration and deceleration. The heading deviation rate refers to the cumulative sum of the absolute values of the angle changes in adjacent frames of the participant's heading angle sequence within the preset look-ahead duration, divided by the cumulative number of frames, measured in degrees per frame, used to quantify the frequency of the participant's directional oscillations.
[0037] The aggressive tendency coefficient is defined as follows: First, determine the upper limit of the velocity fluctuation amplitude saturation. This upper limit is taken as 80% of the maximum velocity sequence value of the game stakeholders within a preset look-ahead period. Aggressive tendency coefficient = Velocity fluctuation amplitude / Upper limit of velocity fluctuation amplitude saturation; when the velocity fluctuation amplitude is greater than or equal to the upper limit of velocity fluctuation amplitude saturation, the aggressive tendency coefficient is 1; when the velocity fluctuation amplitude is less than the upper limit of velocity fluctuation amplitude saturation, the aggressive tendency coefficient is the ratio of the velocity fluctuation amplitude to the upper limit of velocity fluctuation amplitude saturation, and the ratio increases linearly with the increase of the velocity fluctuation amplitude.
[0038] The cooperative tendency coefficient is defined as follows: First, the upper limit of the heading deviation rate saturation is determined. This upper limit is taken as the absolute value of the heading change rate corresponding to the maximum lateral acceleration maintained on the longitudinal axis of the vehicle coordinate system within a preset look-ahead time. In this application, it is taken as 5 degrees / frame. Cooperative tendency coefficient = 1 - (heading deviation rate / upper limit of heading deviation rate saturation); when the heading deviation rate is greater than or equal to the upper limit of heading deviation rate saturation, the cooperative tendency coefficient is 0; when the heading deviation rate is less than the upper limit of heading deviation rate saturation, the cooperative tendency coefficient is 1 minus the ratio of the heading deviation rate to the upper limit of heading deviation rate saturation. The ratio decreases linearly inversely as the heading deviation rate increases.
[0039] The game intention representation value = radical tendency coefficient × (1 - cooperative tendency coefficient), with a value range from 0 to 1. The larger the value, the stronger the player's game intention.
[0040] First, in order to obtain the movement trends of the game stakeholders in the future period and quantify their behavioral characteristics, it is necessary to extract the participant speed sequence and heading angle sequence that have been synchronized within a preset look-ahead time of 3 seconds from the input data of the traffic interaction situation coding network. This will give the speed change trajectory and direction change trajectory of each participant in the next 3 seconds.
[0041] Furthermore, in order to quantify the aggressiveness of participants' acceleration and deceleration, it is necessary to traverse the participant speed sequence of each game staker and calculate the difference between the maximum and minimum values of the sequence within 3 seconds as the speed fluctuation amplitude, which yields a scalar value. For example, a speed fluctuation amplitude of 2.5 m / s indicates that the participant's speed changes over a large range and the acceleration and deceleration actions are drastic.
[0042] Furthermore, in order to quantify the degree of directional swing of the participants, it is necessary to traverse the heading angle sequence of each game participant, calculate the cumulative sum of the absolute values of the heading angle changes in adjacent frames and divide it by the cumulative number of frames to obtain the heading deviation rate. This yields the average turning amplitude per frame. For example, a heading deviation rate of 3 degrees / frame indicates that the participant's direction changes frequently.
[0043] Furthermore, in order to map the speed fluctuation amplitude to an aggressive tendency coefficient between 0 and 1, it is necessary to first determine that the upper limit of the speed fluctuation amplitude is 80% of the participant's own maximum speed, and then calculate the ratio of the speed fluctuation amplitude to the upper limit of the saturation. When the speed fluctuation amplitude reaches or exceeds the upper limit of the saturation, the coefficient is taken as 1; otherwise, the ratio is taken. This yields an aggressive tendency coefficient that increases linearly with the speed fluctuation amplitude.
[0044] Furthermore, in order to map the heading deviation rate to a cooperative tendency coefficient between 0 and 1, it is necessary to first determine the upper limit of the heading deviation rate saturation as 5 degrees / frame, and then calculate 1 minus the ratio of the heading deviation rate to the upper limit of saturation. When the heading deviation rate reaches or exceeds the upper limit of saturation, the coefficient is taken as 0; otherwise, the difference is taken. This yields a cooperative tendency coefficient that decreases linearly in the opposite direction as the heading deviation rate increases, meaning that the more stable the direction, the higher the cooperativeness.
[0045] Finally, in order to comprehensively reflect the overall interaction intentions of the game stakeholders, the radical tendency coefficient and the cooperative tendency coefficient need to be fused according to the formula Game Intention Characteristic Value = Radical Tendency Coefficient × (1 - Cooperative Tendency Coefficient) to obtain a comprehensive scalar value between 0 and 1. Then, based on the correspondence between the identity identifier of each game staker and this characteristic value, a game intention distribution map of the interaction objects is constructed. For example, the characteristic value of vehicle A is 0.75 and the characteristic value of vehicle B is 0.32, thereby providing the intention intensity boundary of each participant for subsequent dynamic inference.
[0046] For example, in a city intersection scenario, the current time of the vehicle is t=0s, with a preset look-ahead time of 3s. After processing by the traffic interaction situation coding network, three stakeholders are identified: vehicle A, vehicle B, and pedestrian C. For vehicle A, its speed sequence is [8m / s, 8.5m / s, 9m / s, 9.2m / s, 9.5m / s], with a maximum speed of 9.5m / s, a saturation limit of 7.6m / s, a speed fluctuation amplitude of 1.5m / s, and an aggressive tendency coefficient of 1.5 / 7.6=0.197; its heading angle sequence is [0°, 2°, 5°, 9°, 14°], with adjacent changes of 2, 3, 4, and 5 degrees, a cumulative sum of 14 degrees, a cumulative frame count of 4, a heading deviation rate of 3.5 degrees / frame, a saturation limit of 5 degrees / frame, and a cooperative tendency coefficient of 1-3.5 / 5=0.3; its game intention representation value is 0.197×(1-0.3)=0.138. For vehicle B, the speed fluctuation amplitude is 4 m / s, the saturation limit is 6 m / s, and the aggressive coefficient is 0.667; the heading deviation rate is 1 degree / frame, and the coordination coefficient is 0.8; the characteristic value is 0.667 × 0.2 = 0.133. For pedestrian C, the speed fluctuation amplitude is 0.5 m / s, the saturation limit is 1 m / s, and the aggressive coefficient is 0.5; the heading deviation rate is 0 degrees / frame, and the coordination coefficient is 1; the characteristic value is 0.5 × 0 = 0. The final distribution map is: vehicle A: 0.138, vehicle B: 0.133, pedestrian C: 0.
[0047] In summary, this step extracts the velocity and heading sequence of the game stakeholders within a preset look-ahead time, calculates the velocity fluctuation amplitude and heading deviation rate respectively, and then obtains the aggressive tendency coefficient and cooperative tendency coefficient according to the linear mapping rule. Finally, it uses a fusion formula to obtain the game intention representation value, constructing a distribution map of the interactive objects' game intentions. Compared with existing technologies, this step has the following advantages: First, it transforms the participants' velocity fluctuations and heading swings into quantifiable aggressive and cooperative tendencies, achieving an objective characterization of the driver's behavioral style; second, it uses a normalized mapping based on a saturation upper limit, making the tendency coefficients between participants of different speed levels comparable; third, by integrating the two dimensions through the game intention representation value, it provides continuous and smooth decision input for the subsequent dynamic search boundary, avoiding decision jumps caused by the discretization of behavioral classification in traditional methods.
[0048] S200: Using the game intention distribution map of the interactive object as the dynamic search boundary, drive the time-series game deduction network to perform multi-step deduction and generate the collision time margin decay sequence and equivalent traffic deceleration corresponding to the candidate driving action of the vehicle. In this embodiment of the application, in a scenario where the autonomous vehicle has obtained the distribution map of the game intentions of the surrounding participants but has not yet determined the specific driving action, in order to transform the abstract intention information into a computable dynamic deduction boundary and quantify the safety and traffic efficiency of different candidate actions, it is necessary to constrain the multi-step prediction of the temporal game deduction network with the game intention distribution map, and output the collision time margin decay process and equivalent traffic deceleration corresponding to each candidate action, so as to solve the technical problems of traditional methods that only perform collision detection based on the predicted position and lack continuous safety margin assessment and traffic efficiency quantification of actions.
[0049] Step S200 in the method provided in this application embodiment includes: Using the current speed of the vehicle as the base speed, the upper limit of the vehicle's acceleration capability and the lower limit of its braking deceleration are obtained. The upper limit of the vehicle's acceleration capability is the maximum positive acceleration that the vehicle's power system can output under the current operating conditions, and the lower limit of its braking deceleration is the negative absolute value of the maximum deceleration that the vehicle's braking system can output under the current operating conditions. Within the acceleration / deceleration range defined by the lower limit of braking deceleration and the upper limit of the vehicle's acceleration capability, several candidate acceleration values are generated discretely using an equal-interval sampling method. The reference speed is paired with the plurality of candidate acceleration values one by one to generate a set of candidate driving actions for the vehicle. Each candidate driving action of the vehicle is characterized as a motion mode with the reference speed as the initial speed and the corresponding candidate acceleration value as a constant acceleration. The game intention representation values of each game-related party in the interaction object game intention distribution map are bound to the corresponding identity identifiers, which serve as the dynamic search boundary of the time-series game inference network. The game intention characterization value constrains the upper limit of speed extrapolation and the range of heading deviation for each game stakeholder. The larger the game intention characterization value, the higher the upper limit of speed extrapolation and the wider the range of heading deviation. The time-series game inference network is driven by taking the candidate driving actions of the vehicle as input and the dynamic search boundary as constraint. The future positions of the vehicle and each game stakeholders are gradually inferred along the inference step. At each step, the collision time margin is calculated based on the relative distance and relative approach speed between the vehicle and each game stakeholders, forming the collision time margin decay sequence. The reference passage time is taken from the moment when the vehicle passes through the center section of the intersection at a constant current speed. The equivalent traffic deceleration is calculated as the ratio of the difference between the estimated time when the vehicle reaches the center section of the intersection after performing the candidate driving action and the baseline travel time, to a preset look-ahead time. A detailed explanation follows: In this embodiment, candidate acceleration values refer to discrete acceleration values sampled at equal intervals within the acceleration / deceleration interval, with a sampling interval of 0.2 m / s². For example, within the interval [-3 m / s², 2 m / s²], 26 candidate values [-3, -2.8, ..., 2] can be sampled. The set of candidate driving actions for the vehicle refers to multiple motion patterns generated by pairing a base speed with each candidate acceleration value. Each pattern is considered a candidate action, and the set size is equal to the number of candidate acceleration values. The dynamic search boundary refers to the upper limit of speed and the possible deviation range of heading for each participant in the spatiotemporal deduction, based on the intention representation values of each game staker in the interactive object game intention distribution map, used to constrain the prediction output of the deduction network. The collision time margin refers to the remaining time required for the vehicle to approach a certain game staker to zero distance along the current relative speed direction during the deduction process, measured in seconds, and calculated considering both relative distance and relative approach speed. The baseline travel time refers to the time required for a vehicle to reach the center section of the intersection from its current position, assuming it travels at a constant instantaneous speed. It serves as a reference baseline for measuring traffic efficiency. The equivalent travel deceleration is the difference between the actual time it takes for the vehicle to reach the center section of the intersection after performing a candidate driving action and the baseline travel time, divided by a preset look-ahead time. It is dimensionless and used to quantify the degree of traffic delay caused by the candidate action.
[0050] The temporal game theory inference network is constructed as follows: First, historical interaction trajectory data of urban intersections is collected. The game intention representation values and identity identifiers of each game staker are extracted as historical game intention data. Candidate driving actions of the vehicle are extracted as historical candidate action data. The sequence of the vehicle's actual future positions with each game staker after executing a candidate driving action is extracted as the inference trajectory label. Then, the network architecture is constructed, consisting of a cross-attention encoding layer, a temporal recursion layer, and a dual-output layer. The input of the cross-attention encoding layer receives historical game intention data and historical candidate action data. The input of the temporal recursion layer is connected to the output of the cross-attention encoding layer, and the input of the dual-output layer is connected to the output of the temporal recursion layer. The first output of the dual-output layer outputs the predicted value of the collision time margin decay sequence, and the second output outputs the predicted value of the equivalent traffic deceleration. The loss function is composed of the sum of a first mean squared error and a second mean squared error: the first mean squared error is the mean squared error between the predicted value of the collision time margin decay sequence and the collision time margin decay sequence label calculated from the inferred trajectory label; the second mean squared error is the mean squared error between the predicted value of the equivalent traffic deceleration and the equivalent traffic deceleration label calculated from the inferred trajectory label. The network weights are updated using the backpropagation algorithm until the loss function converges, and the trained network architecture is used as the temporal game inference network.
[0051] In this step, firstly, in order to construct the acceleration space of candidate driving actions, the current speed of the vehicle is used as the reference speed. The maximum positive acceleration that the vehicle can output under the current working conditions is obtained as the upper limit of acceleration capability, and the absolute value of the maximum braking deceleration is negative as the lower limit of braking deceleration. The boundary range of the action space can be obtained, for example, the upper limit of acceleration capability is 2m / s², and the lower limit of braking deceleration is -3m / s².
[0052] Furthermore, in order to generate discrete action options for simulation and comparison, several candidate acceleration values need to be discretized within the closed interval formed by the lower limit of braking deceleration and the upper limit of acceleration capability, using an equal-interval sampling method of 0.2 m / s², thus obtaining a finite set of acceleration candidate values.
[0053] Furthermore, in order to form a complete and executable action definition, the reference speed needs to be paired with each candidate acceleration value. Each pairing represents a motion pattern with the reference speed as the initial speed and the candidate acceleration value as a constant acceleration. All pairs together constitute the set of candidate driving actions for the vehicle.
[0054] Then, in order to incorporate game intention information into the deduction process, the game intention representation value of each game stakeholder in the game intention distribution map of the interactive objects needs to be bound to its identity identifier, which serves as the dynamic search boundary of the time-series game deduction network, so that the network knows the range of allowed behavior of each participant during the deduction.
[0055] Furthermore, in order to specifically constrain the extrapolation space of each participant, the upper limit of speed extrapolation and the range of heading offset need to be adjusted according to the magnitude of the game intention representation value: the larger the representation value, the higher the upper limit of speed extrapolation and the wider the range of heading offset.
[0056] Next, the driving time-series game simulation network takes each action in the candidate driving action set of the vehicle as input in turn, uses the dynamic search boundary as constraint, and gradually recursively deduces the future position sequence of the vehicle and each game staker according to a fixed simulation step size. In each step, the collision time margin is calculated based on the relative distance and relative approach speed between the vehicle and each participant. As the simulation progresses, the collision time margin gradually changes to form a decay sequence, which is used to reflect the changing trend of safety risk over time.
[0057] Next, in order to evaluate the traffic efficiency of each candidate action, the time when the vehicle passes through the center section of the intersection at its current constant speed needs to be calculated as the baseline passage time.
[0058] Finally, the difference between the estimated time when the vehicle actually arrives at the center section of the intersection after performing the candidate action and the baseline passage time is divided by the preset look-ahead time to obtain the equivalent passage deceleration. This value quantifies the passage delay of the candidate action compared to constant speed driving.
[0059] Step S200 in the method provided in this application embodiment further includes: Extract the game intention representation values that are bound to the identity identifiers of each game participant in the dynamic search boundary; Using the last frame velocity of the participant velocity sequence after time synchronization within the preset look-ahead time as the base velocity, the product of the base velocity and the game intention representation value and the vehicle acceleration capability upper limit is added to obtain the velocity extrapolation upper limit of each game participant. Using the heading angle of the last frame in the participant heading angle sequence after time synchronization within the preset look-ahead time as the reference heading, the maximum heading deviation angle corresponding to the heading deviation rate is multiplied by the game intention representation value to obtain the heading deviation range of each game participant. The aforementioned speed extrapolation upper limit is taken as the maximum instantaneous speed that each game participant is allowed to achieve in the time-series simulation; The aforementioned heading deviation range is defined as the allowable angular range within which each game participant deviates from the baseline heading during time-series analysis. A detailed explanation follows: In this embodiment, the speed extrapolation upper limit refers to the maximum instantaneous speed that a certain game staker is allowed to reach during the time-series simulation, measured in m / s. It is obtained by adding the intended representation value to the base speed and the vehicle's acceleration capability upper limit. The heading deviation range refers to the angular range that a certain game staker is allowed to deviate from its baseline heading during the time-series simulation, measured in degrees. It is obtained by multiplying the maximum heading deviation angle by the intended representation value, with the maximum heading deviation angle set at 30 degrees.
[0060] First, in order to obtain the intention strength of each game staker to constrain the deduction boundary, it is necessary to extract the game intention representation value bound to the identity identifier of each game staker from the dynamic search boundary, so as to obtain the intention score of 0 to 1 for each participant.
[0061] Furthermore, in order to determine the maximum speed that each participant is allowed to reach in future simulations, the speed of the last frame of the speed sequence of each game stakeholder within a preset look-ahead time period is taken as the base speed. This base speed is then added to the product of the game intention representation value and the upper limit of the vehicle's acceleration capability to obtain the speed extrapolation upper limit. That is, the stronger the intention, the higher the acceleration upper limit.
[0062] Furthermore, in order to determine the allowable course swing range for each participant in future simulations, the course angle of the last frame of the course angle sequence within the preset look-ahead time for each game stakeholder is taken as the baseline course. The preset maximum course offset angle is multiplied by the game intention representation value to obtain the course offset range, that is, the stronger the intention, the greater the allowable range of direction change.
[0063] Furthermore, the calculated speed extrapolation upper limit is taken as the maximum instantaneous speed that the parties involved in the game are allowed to reach in the time series simulation. Any speed prediction value generated by the network during the simulation process that exceeds this upper limit is truncated.
[0064] Finally, the calculated heading deviation range is taken as the allowable angle range for the parties involved in the game to deviate from the baseline heading in the time series simulation. The change in heading angle during the simulation must not exceed this range, thereby ensuring that the simulation trajectory conforms to the intended boundary constraints.
[0065] For example, in a city intersection scenario, vehicle A's game intention representation value is 0.138, vehicle B's is 0.133, and pedestrian C's is 0. The vehicles' current speed is 10 m / s, their acceleration capacity upper limit is 2 m / s², and their braking deceleration lower limit is -3 m / s². 26 candidate acceleration values are obtained by sampling at equal intervals of 0.2 m / s², generating 26 candidate actions. For vehicle A, its preset look-ahead time and final frame speed are 12 m / s, with an upper limit for speed extrapolation = 12 + 0.138 × 2 = 12.276 m / s; the final frame heading angle is 5°, the maximum heading offset angle is 30°, and the heading offset range = 30 × 0.138 ≈ 4.14°. For vehicle B, the speed in the last frame is 11 m / s, and the upper limit of the speed extrapolation is 11 + 0.133 × 2 = 11.266 m / s; the heading angle in the last frame is -2°, and the heading offset range is 30 × 0.133 ≈ 4.0°. For pedestrian C, the speed in the last frame is 1.5 m / s, the intentional representation value is 0, the upper limit of the speed extrapolation is 1.5 + 0 × 2 = 1.5 m / s, and the heading offset range is 0°. The extrapolation step size is 0.1 s, and the total duration is 3 s. For the vehicle's action with a candidate acceleration value of 1 m / s², the extrapolated collision time margin sequence with vehicle A is [2.1 s, 1.8 s, 1.4 s, 0.9 s, 0.4 s], with a minimum value of 0.4 s; the baseline passage time is 2.5 s, the actual arrival time at the intersection center section is 2.7 s, and the equivalent passage deceleration is (2.7 - 2.5) / 3 ≈ 0.067.
[0066] In summary, this step uses the game intention distribution map as a dynamic search boundary, combines the vehicle's acceleration and deceleration capabilities to generate a set of candidate driving actions, and drives a time-series game inference network to perform multi-step inferences, outputting the collision time margin decay sequence and equivalent traffic deceleration for each action. Compared with existing technologies, this step has the following advantages: First, the intention representation value is directly used to constrain the inference boundary of speed and heading, making the prediction more consistent with the actual behavioral style of the participants; second, by calculating the collision time margin decay sequence, continuous monitoring of safety risks over time is achieved, rather than single-point judgment; third, the introduction of equivalent traffic deceleration quantifies the loss of traffic efficiency, providing an intuitive cost indicator for subsequent arbitration.
[0067] S300: Based on the collision time margin decay sequence and the equivalent passage deceleration, construct the conflict urgency arbitration logic. When the minimum collision time margin is lower than the linear mapping threshold of the equivalent passage deceleration, output the risk avoidance instruction; otherwise, output the compliance interaction instruction. In this embodiment of the application, in a scenario where the collision time margin decay sequence and equivalent traffic deceleration corresponding to each candidate driving action have been obtained but the action has not yet been finally selected, in order to dynamically evaluate the safety risk and traffic cost of each candidate action and make the optimal decision, it is necessary to construct an arbitration logic based on the minimum collision time margin and equivalent traffic deceleration. The logic adaptively determines whether the risk should be avoided through a linear mapping threshold, so as to solve the technical problem that the traditional method uses a fixed safety time threshold, which is too aggressive or too conservative in high-speed or emergency conditions and cannot balance safety and traffic efficiency.
[0068] Step S300 in the method provided in this application embodiment includes: From the multiple collision time margin decay sequences corresponding to the candidate driving actions of the autonomous vehicle, the minimum collision time margin of each collision time margin decay sequence is extracted to form a set of minimum collision time margins for candidate actions. Based on the vehicle's current speed and the lower limit of braking deceleration, the time required for the vehicle to decelerate from its current speed to a standstill through full braking is determined as the full braking collision avoidance time. For each candidate driving action of the autonomous vehicle, a linear mapping threshold is determined based on the equivalent traffic deceleration corresponding to the candidate action and the full braking collision avoidance time. The linear mapping threshold decreases as the equivalent traffic deceleration corresponding to the candidate action increases and increases as the full braking collision avoidance time increases. Traverse the set of minimum collision time margins for the candidate actions, and compare each minimum collision time margin with the corresponding linear mapping threshold for each candidate action one by one. When all the minimum collision time margins in the set of minimum collision time margins for candidate actions are less than their respective linear mapping thresholds, a risk avoidance instruction is output. When there is a minimum collision time margin in the set of candidate actions that is greater than or equal to its corresponding linear mapping threshold, a compliance interaction command is output, marking the candidate driving action of the vehicle corresponding to the minimum collision time margin that meets the condition as an available compliance action. A detailed explanation follows: In this embodiment, the set of minimum collision time margins for candidate actions refers to the set of minimum values extracted from the collision time margin decay sequence corresponding to each action in the candidate driving action set of the vehicle, arranged in the order of actions. Full braking collision avoidance time refers to the time required for the vehicle to decelerate fully from its current speed to a standstill at the lower limit of braking deceleration, calculated as: Full braking collision avoidance time = Current speed of the vehicle / |Lower limit of braking deceleration|, in seconds. The linear mapping threshold refers to the dynamic threshold used for comparison with the minimum collision time margin, calculated as: Linear mapping threshold = Full braking collision avoidance time × (1 - Equivalent traffic deceleration), in the same unit as time. The risk avoidance command refers to the control command output when all candidate actions fail to meet safety conditions, instructing the vehicle to perform emergency braking or other avoidance maneuvers to avoid a collision. The compliance interaction command refers to the control command output when at least one candidate action meets safety conditions, instructing the vehicle to select one action from those that meet the conditions to continue driving. Available compliance actions refer to the autonomous vehicle candidate driving actions in the candidate action set that satisfy the minimum collision time margin being greater than or equal to its corresponding linear mapping threshold, and are marked as selectable safety actions.
[0069] First, in order to assess the most dangerous moment for each candidate action, the minimum value of each sequence needs to be extracted from the multiple collision time margin decay sequences corresponding to the candidate driving action set of the vehicle. This forms a set of minimum collision time margin values for the candidate actions, resulting in a set of scalar values, where each value represents the minimum safety margin that the corresponding action may encounter during the simulation period.
[0070] Furthermore, in order to quantify the vehicle's emergency avoidance capability at the current speed, it is necessary to calculate the full braking collision avoidance time based on the vehicle's current speed and the lower limit of braking deceleration. That is, the time required for the vehicle to decelerate from the current speed to a standstill at the maximum deceleration, which can provide a benchmark time that reflects the vehicle's inherent braking performance.
[0071] Furthermore, in order to generate an adaptive safety judgment threshold for each candidate action, the threshold corresponding to each action needs to be calculated using the formula linear mapping threshold = full braking collision avoidance time × (1 - equivalent traffic deceleration) based on the equivalent traffic deceleration and full braking collision avoidance time corresponding to the candidate action. This yields a dynamic threshold that decreases as the equivalent traffic deceleration increases and increases as the full braking collision avoidance time increases.
[0072] Furthermore, in order to determine whether each candidate action is safe, it is necessary to traverse the set of minimum collision time margins of candidate actions and compare each minimum value with the linear mapping threshold of its corresponding action one by one to obtain the safety determination result of each action.
[0073] Furthermore, when the minimum collision time margin of all candidate actions is less than their respective linear mapping thresholds, it means that no action can avoid a collision under the safety margin requirement. At this time, a risk avoidance instruction is output, instructing the vehicle to abandon its intention to pass.
[0074] Finally, when there is at least one candidate action whose minimum collision time margin is greater than or equal to its corresponding linear mapping threshold, the compliant interaction instruction is output, and all candidate actions that meet this condition are marked as available compliant actions for subsequent selection.
[0075] For example, in a city intersection scenario, if the vehicle's current speed is 10 m / s and the lower limit of braking deceleration is -3 m / s², then the full braking collision avoidance time is approximately 10 / 3 s. In the preceding steps, for the action with a candidate acceleration value of 1 m / s², the equivalent traffic deceleration is approximately 0.067. Therefore, its linear mapping threshold is approximately 3.33 × (1 - 0.067) ≈ 3.11 s, while the minimum collision time margin for this action is 0.4 s. Since 0.4 < 3.11, this action does not meet the safety requirements. Similarly, assuming the equivalent traffic deceleration of the other 25 candidate actions is distributed between -0.5 and 0.5, the corresponding linear mapping threshold is between 3.33×(1+0.5)=5.0s and 3.33×(1-0.5)=1.67s, and the minimum collision time margin of all actions is distributed between 0.3s and 0.8s, all of which are less than their respective linear mapping thresholds, so a risk avoidance instruction is output.
[0076] In summary, this step extracts the minimum collision time margin of each candidate action, calculates the full braking collision avoidance time by combining the vehicle's current speed and the lower limit of braking deceleration, and then dynamically generates a linear mapping threshold for each candidate action. By comparing each candidate action, it determines whether a safe action exists, and finally outputs a risk avoidance or compliance interaction command. Compared with existing technologies, this step has the following advantages: First, it uses an adaptive threshold based on equivalent traffic deceleration, allowing the safety judgment standard to adjust with the traffic efficiency of the action, avoiding the problem of a fixed threshold being too conservative at low speeds and too aggressive at high speeds; second, it uses the full braking collision avoidance time as a benchmark, linking the threshold to the vehicle's actual braking capacity, improving the physical rationality of the arbitration logic; third, by traversing all candidate actions and marking available compliance actions, it provides multiple options rather than a single decision point for subsequent behavior generation, enhancing the system's flexibility.
[0077] S400: Based on the risk avoidance command and the compliance interaction command, and combined with the vehicle sensor data, generate the driving behavior that can be executed at the current moment.
[0078] In this embodiment of the application, in scenarios where the arbitration logic has output risk avoidance instructions or complied with interactive instructions but has not yet been converted into specific control quantities, in order to directly map the decision result and the current motion state of the vehicle into executable physical actions, it is necessary to generate the acceleration control quantity at the current moment based on the instruction type and information such as position, speed, and heading angle in the vehicle's sensor data. This is to solve the technical problem of semantic gap and large response delay caused by the separation of planning and control in traditional architecture.
[0079] Step S400 in the method provided in this application embodiment includes: When a risk avoidance command is output, the current speed of the vehicle is used as the initial speed and the lower limit of braking deceleration is used as the constant acceleration to generate an emergency braking driving behavior, which is the driving behavior that can be executed at the current moment. When the output complies with the interactive command, from the available complies, select the candidate driving action of the vehicle with the smallest equivalent deceleration and the sum of the candidate acceleration value of the corresponding candidate driving action and the current speed of the vehicle within the speed variation range formed by the upper limit of the vehicle's acceleration capability and the lower limit of the braking deceleration, and use it as the selected complies action. Using the vehicle's current speed as the initial speed, and the selected candidate acceleration value for the follow-up action as a constant acceleration, the executable driving behavior for the current moment is generated by combining the vehicle's position sequence and heading angle sequence. A detailed explanation follows: In this embodiment, emergency braking driving behavior refers to a deceleration mode with the vehicle's current speed as the initial speed and the lower limit of braking deceleration as the constant acceleration, used to minimize braking distance under risk avoidance commands. Selecting a follow-up action refers to selecting the vehicle's candidate driving action from the available follow-up action set according to the principle of minimizing the equivalent traffic deceleration and ensuring that the sum of the candidate acceleration value and the current speed is within the speed variation range. The speed variation range refers to the adjustable speed interval formed by the lower limit of braking deceleration and the upper limit of the vehicle's acceleration capability, i.e., the range between the minimum and maximum speeds allowed to be reached by the vehicle after one control cycle. The lower limit is the sum of the vehicle's current speed and the lower limit of braking deceleration, and the upper limit is the sum of the vehicle's current speed and the upper limit of its acceleration capability.
[0080] First, when the arbitration logic outputs a risk avoidance instruction, in order to decelerate the vehicle to a safe state in the shortest distance, the current speed of the vehicle must be used as the initial speed and the lower limit of braking deceleration must be used as a constant acceleration to generate an emergency braking driving behavior, which can result in a full braking acceleration instruction. For example, if the current speed is 10m / s and the lower limit of deceleration is -3m / s², the behavior description is to continuously decelerate at -3m / s².
[0081] Furthermore, when outputting compliance commands, to maximize traffic efficiency while ensuring safety, the optimal action needs to be selected from the available compliance action set. First, the candidate action with the smallest equivalent traffic deceleration is selected, as a smaller equivalent deceleration indicates less traffic delay. Simultaneously, it's necessary to verify whether the sum of the candidate acceleration value of this action and the vehicle's current speed falls within the speed variation range, specifically the interval formed by the lower limit of braking deceleration and the upper limit of the vehicle's acceleration capability. If this condition is met, the action is selected as the compliance action; otherwise, the next smallest equivalent traffic deceleration action is selected for verification until an action that meets the criteria is found.
[0082] Finally, in order to determine the specific driving behavior corresponding to the selected action, the current speed of the vehicle is used as the initial speed, the candidate acceleration value of the selected action is used as a constant acceleration, and the vehicle position sequence and the vehicle heading angle sequence are combined to obtain the executable driving behavior containing spatial direction and speed change information. For example, if the current speed is 10m / s, the selected acceleration is 1m / s², and the heading angle is 0°, then the behavior is described as accelerating at 1m / s² along the 0° direction.
[0083] For example, in a city intersection scenario, since none of the candidate actions meet the safety conditions in the aforementioned steps, a risk avoidance instruction is output. The vehicle's current speed is 10 m / s, the lower limit of braking deceleration is -3 m / s², the vehicle's position sequence current frame coordinates are (0m, 0m), and the heading angle is 0°. Based on the risk avoidance instruction, an emergency braking driving behavior is generated: decelerating along the current heading with an initial speed of 10 m / s and a constant acceleration of -3 m / s² until the speed drops to 0. This behavior is directly output to the underlying controller as the currently executable driving behavior. If it is assumed that there is an available follow-up action, for example, if the equivalent deceleration of a candidate action is 0.05, the candidate acceleration value is 0.5 m / s², the vehicle's current speed is 10 m / s, the lower limit of the speed change range is 10 - 3 = 7 m / s, the upper limit is 10 + 2 = 12 m / s, and 10 + 0.5 = 10.5 m / s is within the range, then this action is selected as the selected follow-up action. With an initial speed of 10 m / s and an acceleration of 0.5 m / s², combined with the vehicle's heading angle of 0°, a driving behavior of uniform acceleration along the 0° direction is generated.
[0084] In summary, this step generates emergency braking driving behavior directly based on risk avoidance instructions, or selects the action with the minimum equivalent deceleration and acceleration value within the speed variation range from available compliance actions under compliance interaction instructions, and generates driving behavior with constant acceleration accordingly, achieving a seamless mapping from decision-making to execution. Compared with existing technologies, this step has the following advantages: First, it uses maximum braking capacity during risk avoidance, ensuring safety in emergency situations; second, it selects the action with the minimum equivalent deceleration, prioritizing traffic efficiency; third, it adds speed variation range verification to avoid generating unexecutable actions that exceed the vehicle's dynamic limits; and fourth, it combines position and heading angle to provide a complete spatial description of the driving behavior, facilitating direct tracking by the underlying controller.
[0085] S500: Collect the actual interaction trajectory after the execution of the executable driving behavior, and use the cumulative deviation between the actual interaction trajectory and the inferred trajectory as the driving signal to reverse calibrate the temporal game inference network and the traffic interaction situation coding network.
[0086] In this embodiment of the application, in the scenario where the autonomous vehicle has executed decision-making actions and obtained feedback from the real environment, in order to continuously narrow the gap between the prediction model and the real environment and improve the decision-making accuracy in long-term operation, it is necessary to collect the real interaction trajectory and compare it with the inferred trajectory, calculate the cumulative deviation as an error signal, and use it to adjust the internal parameters of the inferred network and the coding network in reverse, so as to solve the technical problems of traditional pipeline methods lacking online adaptive calibration mechanism, being unable to learn and improve from actual interactions, and causing the prediction deviation to accumulate and expand after long-term operation.
[0087] Step S500 in the method provided in this application embodiment includes: After executing the currently executable driving behavior, the real position sequence of the vehicle within a preset look-ahead time period is collected, as well as the real position sequence of the game stakeholders within the preset look-ahead time period; Before executing the currently executable driving action, the temporal game inference network generates the vehicle's inferred position sequence and the inferred position sequence of the game stakeholders for the selected driving action; The actual vehicle position sequence is compared frame by frame with the estimated vehicle position sequence to calculate the cumulative vehicle position deviation. The actual position sequence of each game staker is compared frame by frame with the inferred position sequence of the corresponding game staker, and the cumulative position deviation of the game staker is calculated. With the goal of reducing the cumulative deviation of the vehicle's position, the mapping parameters used to generate the upper limit of speed extrapolation and the mapping parameters used to generate the range of heading offset in the time-series game inference network are adjusted. The weights of the encoder layer and the cross-attention layer in the traffic interaction situation coding network are adjusted with the aim of reducing the cumulative positional bias of the game stakeholders. A detailed explanation follows: In this embodiment, the real interaction trajectory refers to the sequence of actual positions of the vehicle and other game stakeholders within a preset look-ahead time, obtained by sensors in a real environment after the vehicle performs an executable driving action. The projected trajectory refers to the sequence of positions of the vehicle and other game stakeholders predicted by the temporal game projection network for the same time period before the action is executed. Mapping parameters refer to the linear scaling coefficients used in the temporal game projection network to convert game intention representation values into speed extrapolation upper limits, and the gain coefficients used to convert them into heading offset ranges. Network weights refer to the learnable parameter matrices in the encoder layer and cross-attention layer of the traffic interaction situation coding network.
[0088] Accumulated deviation of vehicle position Where: i is the frame index within the preset look-ahead time, taking a positive integer sequence i=1,2,...,n, and n is the total number of frames within the preset look-ahead time; To collect the actual position coordinates of the vehicle in the i-th frame of the vehicle coordinate system after executing the currently executable driving behavior; The position coordinates of the vehicle in the vehicle coordinate system are predicted by the temporal game inference network at the i-th frame of the inference step; Cumulative positional bias of game stakeholders Where: j is the index of the game stakeholders, which is traversed one by one; i is the frame index within the preset look-ahead time, which is a positive integer sequence i=1,2,...,n, where n is the total number of frames within the preset look-ahead time; To collect the actual position coordinates of the j-th game player in the i-th frame of the vehicle coordinate system after executing the currently executable driving behavior; Let be the position coordinates of the j-th player in the vehicle coordinate system as predicted by the temporal game simulation network in the i-th frame of the simulation step; First, in order to obtain real-world feedback to assess prediction errors, after executing the currently executable driving behavior, the vehicle's real position sequence within a preset look-ahead time period needs to be collected using onboard sensors, while the real position sequences of each game participant within the same time period also need to be collected.
[0089] Furthermore, in order to obtain the prediction results to be calibrated, it is necessary to extract the self-vehicle position sequence generated by the network simulation before executing the driving behavior and the position sequences of each game staker from the memory of the time-series game simulation network.
[0090] Furthermore, according to the definition formula of the cumulative position deviation of the vehicle, the actual position sequence of the vehicle and the estimated position sequence of the vehicle are compared frame by frame and the Euclidean distance is accumulated to obtain the cumulative position deviation of the vehicle.
[0091] Furthermore, according to the definition formula of the cumulative position deviation of the game stakeholders, each game staker is traversed, and its actual position sequence is compared with the inferred position sequence frame by frame and the Euclidean distance is accumulated. Then, the independent cumulative deviations of all game stakeholders are summed to obtain a unified cumulative position deviation of the game stakeholders.
[0092] Furthermore, with the goal of reducing the cumulative deviation of the vehicle's position, the mapping parameters used to generate the upper limit of speed extrapolation and the mapping parameters used to generate the heading offset range in the time-series game simulation network are adjusted by gradient descent, so that the predicted trajectory of the vehicle in the next simulation is closer to the actual trajectory.
[0093] Finally, with the optimization goal of reducing the cumulative bias in the positions of the game stakeholders, the network weights of the encoder layer and the cross-attention layer in the traffic interaction situation coding network are adjusted through the backpropagation algorithm. This makes the network more accurate in extracting the interaction features of the participants, thereby improving the prediction accuracy of subsequent inferences.
[0094] For example, in a city intersection scenario, after an emergency braking driving action, real data is collected within a preset look-ahead time of 3 seconds. The actual vehicle position sequence (frame interval 0.1s) is: Frame 1 (0,0), Frame 2 (0.45,0), Frame 3 (0.85,0), ..., Frame 30 (8.5,0); the projected vehicle position sequence is: Frame 1 (0,0), Frame 2 (0.5,0), Frame 3 (0.95,0), ..., Frame 30 (9.0,0). The Euclidean distance is calculated frame by frame and accumulated, resulting in a cumulative deviation of 3.2m for the vehicle position. For vehicle A, a stakeholder in the game, the cumulative deviation of the Euclidean distance between the actual and projected position sequences is 1.5m; for vehicle B, the cumulative deviation is 0.8m; for pedestrian C, the cumulative deviation is 0.2m; the cumulative deviation of the stakeholder's position is 1.5 + 0.8 + 0.2 = 2.5m. To reduce the cumulative bias of the vehicle, the mapping parameter for the speed extrapolation upper limit in the time-series game inference network was adjusted from 1.0 to 0.95, and the gain coefficient for the heading offset range was adjusted from 1.0 to 0.98. To reduce the cumulative bias of the game stakeholders, the weights of the encoder layer and cross-attention layer in the traffic interaction situation coding network were adjusted, and the learning rate was set to 0.001.
[0095] In summary, this step collects real interaction trajectories and inferred trajectories, calculates the cumulative position deviation of the vehicle and the cumulative position deviation of the game stakeholders according to defined formulas, and then calibrates the mapping parameters of the time-series game inference network and the network weights of the traffic interaction situation coding network with the goal of reducing the deviation. Compared with existing technologies, this step has the following advantages: First, it achieves online adaptive calibration, enabling the decision-making system to continuously learn from actual interactions and improve its prediction capabilities, effectively suppressing error accumulation in long-term operation; second, it designs two calibration mechanisms for vehicle motion prediction and participant motion prediction respectively. The former adjusts the output boundary parameters of the inference network, while the latter updates the feature extraction weights of the coding network. The two mechanisms have clear division of labor and do not interfere with each other; third, the cumulative deviation is accumulated frame by frame using Euclidean distance accumulation, which comprehensively reflects the overall error over the entire prediction period and avoids the randomness of single-point judgments.
[0096] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects: This application proposes an integrated cognitive decision-making method for intelligent agents based on WM and MLM. First, it acquires trajectory fragments of traffic participants from vehicle sensor data and roadside unit broadcasts. A traffic interaction situation coding network generates a distribution map of interactive object game intentions centered on the vehicle. Then, using this distribution map as a dynamic search boundary, it drives a temporal game inference network to generate collision time margin decay sequences and equivalent traffic decelerations corresponding to candidate driving actions of the vehicle. It also constructs a conflict urgency arbitration logic to output risk avoidance instructions or compliance instructions, combining this with vehicle sensor data to generate executable driving behaviors at the current moment. Finally, it collects the actual interaction trajectories after the executable driving behaviors are executed. The cumulative deviation between the actual interaction trajectory and the inferred trajectory is used as a driving signal to reverse-calibrate the temporal game inference network and the traffic interaction situation coding network. This method solves the technical problems of information transmission delays and error accumulation in modular pipeline architectures, limited modeling capabilities for dynamic interactive game scenarios, and the lack of online adaptive calibration mechanisms, which lead to reduced decision robustness in complex traffic environments.
[0097] Example 2, as shown in the appendix Figure 2 As shown, based on the inventive concept of the integrated cognitive decision-making method for intelligent agents based on WM and MLM provided in Embodiment 1, this application also provides an integrated cognitive decision-making system for intelligent agents based on WM and MLM, specifically including: Situational awareness module 11, deployed on the autonomous vehicle computing platform, is used to acquire autonomous vehicle sensor data and traffic participant trajectory segments broadcast by roadside units. It performs masked self-attention cross-coding through a traffic interaction situation coding network to generate a distribution map of interactive object game intentions centered on the autonomous vehicle. The game inference module 12 is used to drive the time-series game inference network to perform multi-step inference with the game intention distribution map of the interactive object as the dynamic search boundary, and generate the collision time margin decay sequence and equivalent traffic deceleration corresponding to the candidate driving action of the vehicle. The conflict arbitration module 13 is used to construct conflict urgency arbitration logic based on the collision time margin decay sequence and the equivalent traffic deceleration. When the minimum collision time margin is lower than the linear mapping threshold of the equivalent traffic deceleration, a risk avoidance instruction is output; otherwise, a compliance interaction instruction is output. The behavior generation module 14 is used to generate the driving behavior that can be executed at the current moment based on the risk avoidance instruction and the compliance interaction instruction, combined with the vehicle sensor data; The online calibration module 15 is used to collect the actual interaction trajectory after the execution of the executable driving behavior, and use the cumulative deviation between the actual interaction trajectory and the inferred trajectory as the driving signal to reverse calibrate the time-series game inference network and the traffic interaction situation coding network.
[0098] In one embodiment, the situation awareness module 11 is further configured to: The vehicle sensing data includes at least the vehicle position sequence, the vehicle speed sequence, and the vehicle heading angle sequence, and the traffic participant trajectory segment includes at least the participant position sequence and the participant speed sequence. Using the absolute timestamp as the alignment reference, the vehicle position sequence and the participant position sequence are synchronized in time, and then a synchronized position pair sequence is generated by matching frames after synchronization. Using the vehicle position in the first synchronized frame as the origin and the vehicle heading angle in the first synchronized frame as the positive vertical axis, a vehicle coordinate system is established, and the participant position sequence is transformed frame by frame into the vehicle coordinate system. Traverse the time intervals of adjacent frames in the participant position sequence transformed to the vehicle coordinate system, and mark the intervals where the time interval exceeds the roadside unit broadcast period as communication interruption segments; Based on the participant positions and speeds of the previous frame before the communication interruption, the frames within the interruption segment are extrapolated and completed at a uniform speed, and a mask identifier is added to the completed frames. The participant position sequence after coordinate transformation and the addition of masking labels, the participant speed sequence after time synchronization, and the synchronized vehicle position sequence, vehicle speed sequence, and vehicle heading angle sequence are used together as the input data of the traffic interaction situation coding network.
[0099] Furthermore, the situational awareness module 11 is also used for: The synchronized vehicle position sequence, vehicle speed sequence, and vehicle heading angle sequence are extracted from the input data of the traffic interaction situation coding network and combined into a vehicle motion state segment. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle motion state segment, and use the masked vehicle motion state segment as the query sequence. Using the participant position sequence after coordinate transformation and the added mask label, and the participant speed sequence after time synchronization as key-value pairs, the traffic interaction situation coding network is input to output the attention response intensity sequence of each participant to their own vehicle. The construction steps of the traffic interaction situation coding network include: Collect historical interaction trajectory data at urban intersections, extract synchronized vehicle position sequence, vehicle speed sequence and vehicle heading angle sequence as vehicle historical motion data, and extract participant position sequence after coordinate transformation and participant speed sequence after time synchronization as participant historical interaction data. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle's historical motion data to generate masked vehicle historical motion data; Using the occluded vehicle's historical motion data as the query and the participants' historical interaction data as the key-value pairs, the inverse of the distance collision time between each participant and the vehicle is used as the attention response intensity label to construct a training sample set; A network architecture for a traffic interaction situation coding network is constructed, wherein the network architecture includes an encoder layer, a cross-attention layer, and a feedforward output layer; The mean squared error between the predicted value of the attention response intensity and the label of the attention response intensity is used as the loss function. The network weights of the network architecture are updated through the backpropagation algorithm until the loss function converges. The trained network architecture is then used as the traffic interaction situation coding network. Traverse the attention response intensity sequence and extract participants whose attention response intensity is greater than an intensity threshold as game stakeholders, wherein the intensity threshold is a predetermined proportion of the maximum value in the attention response intensity sequence; Based on the attention response intensity sequence, game stakeholders are extracted, and the aggressive tendency coefficient and cooperative tendency coefficient of each game stakeholder are calculated and combined to form the game intention distribution map of the interaction objects.
[0100] Furthermore, the situational awareness module 11 is also used for: From the input data of the traffic interaction situation coding network, extract the participant speed sequence and the participant heading angle sequence after time synchronization within a preset look-ahead time. Traverse the participant velocity sequence of the game stakeholders and calculate the range of velocity within the preset look-ahead time as the velocity fluctuation amplitude. Traverse the course angle sequence of the participants of the game stakeholders, calculate the ratio of the cumulative sum of the absolute values of the changes in course angle of adjacent frames within the preset look-ahead time to the cumulative number of frames, and use it as the course deviation rate; The speed fluctuation amplitude is mapped to an aggressive tendency coefficient, wherein the aggressive tendency coefficient increases positively with the increase of the speed fluctuation amplitude; The course deviation rate is mapped to a cooperative tendency coefficient, wherein the cooperative tendency coefficient decreases inversely as the course deviation rate increases; The radical tendency coefficient and the cooperative tendency coefficient are fused together to obtain the game intention representation value of the game stakeholders; Based on the correspondence between the identity identifiers of the game stakeholders and the game intention representation values, a game intention distribution map of the interactive objects is constructed.
[0101] In one embodiment, the game theory deduction module 12 is further configured to: Using the current speed of the vehicle as the base speed, the upper limit of the vehicle's acceleration capability and the lower limit of its braking deceleration are obtained. The upper limit of the vehicle's acceleration capability is the maximum positive acceleration that the vehicle's power system can output under the current operating conditions, and the lower limit of its braking deceleration is the negative absolute value of the maximum deceleration that the vehicle's braking system can output under the current operating conditions. Within the acceleration / deceleration range defined by the lower limit of braking deceleration and the upper limit of the vehicle's acceleration capability, several candidate acceleration values are generated discretely using an equal-interval sampling method. The reference speed is paired with the plurality of candidate acceleration values one by one to generate a set of candidate driving actions for the vehicle. Each candidate driving action of the vehicle is characterized as a motion mode with the reference speed as the initial speed and the corresponding candidate acceleration value as a constant acceleration. The game intention representation values of each game-related party in the interaction object game intention distribution map are bound to the corresponding identity identifiers, which serve as the dynamic search boundary of the time-series game inference network. The game intention characterization value constrains the upper limit of speed extrapolation and the range of heading deviation for each game stakeholder. The larger the game intention characterization value, the higher the upper limit of speed extrapolation and the wider the range of heading deviation. The time-series game inference network is driven by taking the candidate driving actions of the vehicle as input and the dynamic search boundary as constraint. The future positions of the vehicle and each game stakeholders are gradually inferred along the inference step. At each step, the collision time margin is calculated based on the relative distance and relative approach speed between the vehicle and each game stakeholders, forming the collision time margin decay sequence. The reference passage time is taken from the moment when the vehicle passes through the center section of the intersection at a constant current speed. The ratio of the difference between the estimated time when the vehicle arrives at the center section of the intersection after performing the candidate driving action and the reference passage time to the preset look-ahead time is used as the equivalent passage deceleration.
[0102] Furthermore, the game theory simulation module 12 is also used for: Extract the game intention representation values that are bound to the identity identifiers of each game participant in the dynamic search boundary; Using the last frame velocity of the participant velocity sequence after time synchronization within the preset look-ahead time as the base velocity, the product of the base velocity and the game intention representation value and the vehicle acceleration capability upper limit is added to obtain the velocity extrapolation upper limit of each game participant. Using the heading angle of the last frame in the participant heading angle sequence after time synchronization within the preset look-ahead time as the reference heading, the maximum heading deviation angle corresponding to the heading deviation rate is multiplied by the game intention representation value to obtain the heading deviation range of each game participant. The aforementioned speed extrapolation upper limit is taken as the maximum instantaneous speed that each game participant is allowed to achieve in the time-series simulation; The heading offset range is used as the allowable angular range for each game participant to deviate from the baseline heading during the time-series simulation.
[0103] In one embodiment, the conflict arbitration module 13 is further configured to: From the multiple collision time margin decay sequences corresponding to the candidate driving actions of the autonomous vehicle, the minimum collision time margin of each collision time margin decay sequence is extracted to form a set of minimum collision time margins for candidate actions. Based on the vehicle's current speed and the lower limit of braking deceleration, the time required for the vehicle to decelerate from its current speed to a standstill through full braking is determined as the full braking collision avoidance time. For each candidate driving action of the autonomous vehicle, a linear mapping threshold is determined based on the equivalent traffic deceleration corresponding to the candidate action and the full braking collision avoidance time. The linear mapping threshold decreases as the equivalent traffic deceleration corresponding to the candidate action increases and increases as the full braking collision avoidance time increases. Traverse the set of minimum collision time margins for the candidate actions, and compare each minimum collision time margin with the corresponding linear mapping threshold for each candidate action one by one. When all the minimum collision time margins in the set of minimum collision time margins for candidate actions are less than their respective linear mapping thresholds, a risk avoidance instruction is output. When there is a minimum collision time margin in the set of candidate actions that is greater than or equal to its corresponding linear mapping threshold, the follow interaction instruction is output, and the candidate driving action of the vehicle corresponding to the minimum collision time margin that meets the condition is marked as an available follow action.
[0104] In one embodiment, the behavior generation module 14 is further configured to: When a risk avoidance command is output, the current speed of the vehicle is used as the initial speed and the lower limit of braking deceleration is used as the constant acceleration to generate an emergency braking driving behavior, which is the driving behavior that can be executed at the current moment. When the output complies with the interactive command, from the available complies, select the candidate driving action of the vehicle with the smallest equivalent deceleration and the sum of the candidate acceleration value of the corresponding candidate driving action and the current speed of the vehicle within the speed variation range formed by the upper limit of the vehicle's acceleration capability and the lower limit of the braking deceleration, and use it as the selected complies action. Using the vehicle's current speed as the initial speed, and the candidate acceleration value of the selected follow-up action as a constant acceleration, the executable driving behavior at the current moment is generated by combining the vehicle's position sequence and heading angle sequence.
[0105] In one embodiment, the online calibration module 15 is further configured to: After executing the currently executable driving behavior, the real position sequence of the vehicle within a preset look-ahead time period is collected, as well as the real position sequence of the game stakeholders within the preset look-ahead time period; Before executing the currently executable driving action, the temporal game inference network generates the vehicle's inferred position sequence and the inferred position sequence of the game stakeholders for the selected driving action; The actual vehicle position sequence is compared frame by frame with the estimated vehicle position sequence to calculate the cumulative vehicle position deviation. The actual position sequence of each game staker is compared frame by frame with the inferred position sequence of the corresponding game staker, and the cumulative position deviation of the game staker is calculated. With the goal of reducing the cumulative deviation of the vehicle's position, the mapping parameters used to generate the upper limit of speed extrapolation and the mapping parameters used to generate the range of heading offset in the time-series game inference network are adjusted. The weights of the encoder layer and the cross-attention layer in the traffic interaction situation coding network are adjusted with the aim of reducing the cumulative positional bias of the game stakeholders.
[0106] The intelligent agent cognitive decision-making integrated system based on WM and MLM provided in this application can achieve intelligent cognitive decision management in scenarios such as daily driving of autonomous vehicles, complex intersection interaction and game, roadside communication interruption, and mixed traffic of multiple traffic participants. This management process involves fusion of multi-source heterogeneous perception data, modeling of interactive object game intentions, and dynamic decision-making with online adaptive calibration. It can be integrated into onboard computing platforms or roadside edge computing units, effectively improving decision robustness and safety in complex traffic environments, reducing collision risks and the probability of unreasonable driving behaviors, while providing reliable data support for continuous optimization of driving behavior and online updates of predictive models. For the specific workflow and optimization details of this system, please refer to Example 1.
[0107] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
Claims
1. An integrated cognitive decision-making method for intelligent agents based on WM and MLM, characterized in that, The method includes: The system acquires trajectory fragments of traffic participants from vehicle sensor data and roadside unit broadcasts, and performs masked self-attention cross-coding through a traffic interaction situation coding network to generate a distribution map of interactive object game intentions centered on the vehicle, including: The synchronized vehicle position sequence, vehicle speed sequence, and vehicle heading angle sequence are extracted from the input data of the traffic interaction situation coding network and combined into a vehicle motion state segment. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle motion state segment, and use the masked vehicle motion state segment as the query sequence. Using the participant position sequence after coordinate transformation and the added mask label, and the participant speed sequence after time synchronization as key-value pairs, the traffic interaction situation coding network is input to output the attention response intensity sequence of each participant to their own vehicle. The construction steps of the traffic interaction situation coding network include: Collect historical interaction trajectory data at urban intersections, extract synchronized vehicle position sequence, vehicle speed sequence and vehicle heading angle sequence as vehicle historical motion data, and extract participant position sequence after coordinate transformation and participant speed sequence after time synchronization as participant historical interaction data. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle's historical motion data to generate masked vehicle historical motion data; Using the occluded vehicle's historical motion data as the query and the participants' historical interaction data as the key-value pairs, the inverse of the distance collision time between each participant and the vehicle is used as the attention response intensity label to construct a training sample set; A network architecture for a traffic interaction situation coding network is constructed, wherein the network architecture includes an encoder layer, a cross-attention layer, and a feedforward output layer; The mean squared error between the predicted value of the attention response intensity and the label of the attention response intensity is used as the loss function. The network weights of the network architecture are updated through the backpropagation algorithm until the loss function converges. The trained network architecture is then used as the traffic interaction situation coding network. Traverse the attention response intensity sequence and extract participants whose attention response intensity is greater than an intensity threshold as game stakeholders, wherein the intensity threshold is a predetermined proportion of the maximum value in the attention response intensity sequence; Based on the attention response intensity sequence, game stakeholders are extracted, and the aggressive tendency coefficient and cooperative tendency coefficient of each game stakeholder are calculated and combined to form the game intention distribution map of the interaction object. Using the game intention distribution map of the interactive objects as a dynamic search boundary, the temporal game inference network is driven to perform multi-step inference to generate collision time margin decay sequences and equivalent traffic decelerations corresponding to candidate driving actions of the self-vehicle, including: Using the current speed of the vehicle as the base speed, the upper limit of the vehicle's acceleration capability and the lower limit of its braking deceleration are obtained. The upper limit of the vehicle's acceleration capability is the maximum positive acceleration that the vehicle's power system can output under the current operating conditions, and the lower limit of its braking deceleration is the negative absolute value of the maximum deceleration that the vehicle's braking system can output under the current operating conditions. Within the acceleration / deceleration range defined by the lower limit of braking deceleration and the upper limit of the vehicle's acceleration capability, several candidate acceleration values are generated discretely using an equal-interval sampling method. The reference speed is paired with the plurality of candidate acceleration values one by one to generate a set of candidate driving actions for the vehicle. Each candidate driving action of the vehicle is characterized as a motion mode with the reference speed as the initial speed and the corresponding candidate acceleration value as a constant acceleration. The game intention representation values of each game-related party in the interaction object game intention distribution map are bound to the corresponding identity identifiers, which serve as the dynamic search boundary of the time-series game inference network. The game intention representation value constrains the speed extrapolation upper limit and heading deviation range of each game staker. The speed extrapolation upper limit is equal to the speed of the last frame of the speed sequence of the game staker within a preset look-ahead time plus the product of the game intention representation value and the vehicle acceleration capability upper limit. The larger the game intention representation value, the higher the speed extrapolation upper limit and the wider the heading deviation range. The time-series game inference network is driven by taking the candidate driving actions of the vehicle as input and the dynamic search boundary as constraint. The future positions of the vehicle and each game stakeholders are gradually inferred along the inference step. At each step, the collision time margin is calculated based on the relative distance and relative approach speed between the vehicle and each game stakeholders, forming the collision time margin decay sequence. The reference passage time is taken from the moment when the vehicle passes through the center section of the intersection at a constant current speed. The ratio of the difference between the estimated time when the vehicle arrives at the center section of the intersection after performing the candidate driving action and the reference passage time to the preset look-ahead time is used as the equivalent passage deceleration. Based on the collision time margin decay sequence and the equivalent traffic deceleration, a conflict urgency arbitration logic is constructed. When the minimum collision time margin is lower than the linear mapping threshold of the equivalent traffic deceleration, a risk avoidance instruction is output; otherwise, a compliance interaction instruction is output, including: From the multiple collision time margin decay sequences corresponding to the candidate driving actions of the autonomous vehicle, the minimum collision time margin of each collision time margin decay sequence is extracted to form a set of minimum collision time margins for candidate actions. Based on the vehicle's current speed and the lower limit of braking deceleration, the time required for the vehicle to decelerate from its current speed to a standstill through full braking is determined as the full braking collision avoidance time. For each candidate driving action of the autonomous vehicle, a linear mapping threshold is determined based on the equivalent traffic deceleration corresponding to the candidate action and the full braking collision avoidance time. The linear mapping threshold decreases as the equivalent traffic deceleration corresponding to the candidate action increases and increases as the full braking collision avoidance time increases. Traverse the set of minimum collision time margins for the candidate actions, and compare each minimum collision time margin with the corresponding linear mapping threshold for each candidate action one by one. When all the minimum collision time margins in the set of minimum collision time margins for candidate actions are less than their respective linear mapping thresholds, a risk avoidance instruction is output. When there is a minimum collision time margin in the set of candidate action collision time margins that is greater than or equal to its corresponding linear mapping threshold, the follow interaction instruction is output, and the candidate driving action of the vehicle corresponding to the minimum collision time margin that meets the condition is marked as an available follow action. Based on the risk avoidance instructions and the compliance interaction instructions, combined with the vehicle sensor data, the executable driving behavior at the current moment is generated; The actual interaction trajectory after the execution of the executable driving behavior is collected, and the cumulative deviation between the actual interaction trajectory and the inferred trajectory is used as the driving signal to reverse calibrate the temporal game inference network and the traffic interaction situation coding network.
2. The integrated cognitive decision-making method for intelligent agents based on WM and MLM according to claim 1, characterized in that, Acquire vehicle sensor data and roadside unit broadcast traffic participant trajectory segments, including: The vehicle sensing data includes at least the vehicle position sequence, the vehicle speed sequence, and the vehicle heading angle sequence, and the traffic participant trajectory segment includes at least the participant position sequence and the participant speed sequence. Using the absolute timestamp as the alignment reference, the vehicle position sequence and the participant position sequence are synchronized in time, and then a synchronized position pair sequence is generated by matching frames after synchronization. Using the vehicle position in the first synchronized frame as the origin and the vehicle heading angle in the first synchronized frame as the positive vertical axis, a vehicle coordinate system is established, and the participant position sequence is transformed frame by frame into the vehicle coordinate system. Traverse the time intervals of adjacent frames in the participant position sequence transformed to the vehicle coordinate system, and mark the intervals where the time interval exceeds the roadside unit broadcast period as communication interruption segments; Based on the participant positions and speeds of the previous frame before the communication interruption, the frames within the interruption segment are extrapolated and completed at a uniform speed, and a mask identifier is added to the completed frames. The participant position sequence after coordinate transformation and the addition of masking labels, the participant speed sequence after time synchronization, and the synchronized vehicle position sequence, vehicle speed sequence, and vehicle heading angle sequence are used together as the input data of the traffic interaction situation coding network.
3. The integrated cognitive decision-making method for intelligent agents based on WM and MLM according to claim 1, characterized in that, Based on the attention response intensity sequence, game stakeholders are extracted, and the aggressiveness and cooperativeness coefficients of each stakeholder are calculated and combined to form the game intention distribution map of the interaction objects, including: From the input data of the traffic interaction situation coding network, extract the participant speed sequence and the participant heading angle sequence after time synchronization within a preset look-ahead time. Traverse the participant velocity sequence of the game stakeholders and calculate the range of velocity within the preset look-ahead time as the velocity fluctuation amplitude. Traverse the course angle sequence of the participants of the game stakeholders, calculate the ratio of the cumulative sum of the absolute values of the changes in course angle of adjacent frames within the preset look-ahead time to the cumulative number of frames, and use it as the course deviation rate; The speed fluctuation amplitude is mapped to an aggressive tendency coefficient, wherein the aggressive tendency coefficient increases positively with the increase of the speed fluctuation amplitude; The course deviation rate is mapped to a cooperative tendency coefficient, wherein the cooperative tendency coefficient decreases inversely as the course deviation rate increases; The radical tendency coefficient and the cooperative tendency coefficient are fused together to obtain the game intention representation value of the game stakeholders; Based on the correspondence between the identity identifiers of the game stakeholders and the game intention representation values, a game intention distribution map of the interactive objects is constructed.
4. The integrated cognitive decision-making method for intelligent agents based on WM and MLM according to claim 1, characterized in that, The upper limit of velocity extrapolation and the range of course deviation for each game participant are constrained by the game intention representation value, including: Extract the game intention representation values that are bound to the identity identifiers of each game participant in the dynamic search boundary; Using the last frame velocity of the participant velocity sequence after time synchronization within the preset look-ahead time as the base velocity, the product of the base velocity and the game intention representation value and the vehicle acceleration capability upper limit is added to obtain the velocity extrapolation upper limit of each game participant. Using the heading angle of the last frame in the participant heading angle sequence after time synchronization within the preset look-ahead time as the reference heading, the maximum heading deviation angle corresponding to the heading deviation rate is multiplied by the game intention representation value to obtain the heading deviation range of each game participant. The aforementioned extrapolated speed upper limit is taken as the maximum instantaneous speed that each game participant is allowed to achieve in the time-series simulation; The heading offset range is used as the allowable angular range for each game participant to deviate from the baseline heading during the time-series simulation.
5. The integrated cognitive decision-making method for intelligent agents based on WM and MLM according to claim 1, characterized in that, Based on the risk avoidance command and the compliance interaction command, and combined with the vehicle sensor data, the currently executable driving behavior is generated, including: When a risk avoidance command is output, the current speed of the vehicle is used as the initial speed and the lower limit of braking deceleration is used as the constant acceleration to generate an emergency braking driving behavior, which is the driving behavior that can be executed at the current moment. When the output complies with the interactive command, from the available complies, select the candidate driving action of the vehicle with the smallest equivalent deceleration and the sum of the candidate acceleration value of the corresponding candidate driving action and the current speed of the vehicle within the speed variation range formed by the upper limit of the vehicle's acceleration capability and the lower limit of the braking deceleration, and use it as the selected complies action. Using the vehicle's current speed as the initial speed, and the candidate acceleration value of the selected follow-up action as a constant acceleration, the executable driving behavior at the current moment is generated by combining the vehicle's position sequence and heading angle sequence.
6. The integrated cognitive decision-making method for intelligent agents based on WM and MLM according to claim 1, characterized in that, Collect the actual interaction trajectory after the execution of the executable driving behavior, and use the cumulative deviation between the actual interaction trajectory and the inferred trajectory as the driving signal to reverse-calibrate the temporal game inference network and the traffic interaction situation coding network, including: After executing the currently executable driving behavior, the real position sequence of the vehicle within a preset look-ahead time period is collected, as well as the real position sequence of the game stakeholders within the preset look-ahead time period; Before executing the currently executable driving action, the temporal game inference network generates the vehicle's inferred position sequence and the inferred position sequence of the game stakeholders for the selected driving action; The actual position sequence of the vehicle is compared frame by frame with the estimated position sequence of the vehicle to calculate the cumulative position deviation of the vehicle. The actual position sequence of each game staker is compared frame by frame with the inferred position sequence of the corresponding game staker, and the cumulative position deviation of the game staker is calculated. With the goal of reducing the cumulative deviation of the vehicle's position, the mapping parameters used to generate the upper limit of speed extrapolation and the mapping parameters used to generate the range of heading offset in the time-series game inference network are adjusted. The weights of the encoder layer and the cross-attention layer in the traffic interaction situation coding network are adjusted with the aim of reducing the cumulative positional bias of the game stakeholders.
7. An integrated cognitive decision-making system for intelligent agents based on WM and MLM, characterized in that, The system is used to execute the integrated agent cognitive decision-making method based on WM and MLM as described in any one of claims 1-6, and the system comprises: The situational awareness module, deployed on the autonomous vehicle computing platform, acquires trajectory fragments of traffic participants from autonomous vehicle sensor data and roadside unit broadcasts. It then uses a traffic interaction situational coding network to perform masked self-attention cross-coding to generate a distribution map of the interactive object's game intentions centered on the autonomous vehicle. This includes: The synchronized vehicle position sequence, vehicle speed sequence, and vehicle heading angle sequence are extracted from the input data of the traffic interaction situation coding network and combined into a vehicle motion state segment. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle motion state segment, and use the masked vehicle motion state segment as the query sequence. Using the participant position sequence after coordinate transformation and the added mask label, and the participant speed sequence after time synchronization as key-value pairs, the traffic interaction situation coding network is input to output the attention response intensity sequence of each participant to their own vehicle. The construction steps of the traffic interaction situation coding network include: Collect historical interaction trajectory data at urban intersections, extract synchronized vehicle position sequence, vehicle speed sequence and vehicle heading angle sequence as vehicle historical motion data, and extract participant position sequence after coordinate transformation and participant speed sequence after time synchronization as participant historical interaction data. Apply a masking code to the position and speed of the preset proportion of continuous frames in the vehicle's historical motion data to generate masked vehicle historical motion data; Using the occluded vehicle's historical motion data as the query and the participants' historical interaction data as the key-value pairs, the inverse of the distance collision time between each participant and the vehicle is used as the attention response intensity label to construct a training sample set; A network architecture for a traffic interaction situation coding network is constructed, wherein the network architecture includes an encoder layer, a cross-attention layer, and a feedforward output layer; The mean squared error between the predicted value of the attention response intensity and the label of the attention response intensity is used as the loss function. The network weights of the network architecture are updated through the backpropagation algorithm until the loss function converges. The trained network architecture is then used as the traffic interaction situation coding network. Traverse the attention response intensity sequence and extract participants whose attention response intensity is greater than an intensity threshold as game stakeholders, wherein the intensity threshold is a predetermined proportion of the maximum value in the attention response intensity sequence; Based on the attention response intensity sequence, game stakeholders are extracted, and the aggressive tendency coefficient and cooperative tendency coefficient of each game stakeholder are calculated and combined to form the game intention distribution map of the interaction object. The game theory deduction module is used to drive a time-series game theory deduction network to perform multi-step deductions using the game intention distribution map of the interactive objects as a dynamic search boundary, generating collision time margin decay sequences and equivalent traffic decelerations corresponding to candidate driving actions of the self-vehicle, including: Using the current speed of the vehicle as the base speed, the upper limit of the vehicle's acceleration capability and the lower limit of its braking deceleration are obtained. The upper limit of the vehicle's acceleration capability is the maximum positive acceleration that the vehicle's power system can output under the current operating conditions, and the lower limit of its braking deceleration is the negative absolute value of the maximum deceleration that the vehicle's braking system can output under the current operating conditions. Within the acceleration / deceleration range defined by the lower limit of braking deceleration and the upper limit of the vehicle's acceleration capability, several candidate acceleration values are generated discretely using an equal-interval sampling method. The reference speed is paired with the plurality of candidate acceleration values one by one to generate a set of candidate driving actions for the vehicle. Each candidate driving action of the vehicle is characterized as a motion mode with the reference speed as the initial speed and the corresponding candidate acceleration value as a constant acceleration. The game intention representation values of each game-related party in the interaction object game intention distribution map are bound to the corresponding identity identifiers, which serve as the dynamic search boundary of the time-series game inference network. The game intention representation value constrains the speed extrapolation upper limit and heading deviation range of each game staker. The speed extrapolation upper limit is equal to the speed of the last frame of the speed sequence of the game staker within a preset look-ahead time plus the product of the game intention representation value and the vehicle acceleration capability upper limit. The larger the game intention representation value, the higher the speed extrapolation upper limit and the wider the heading deviation range. The time-series game inference network is driven by taking the candidate driving actions of the vehicle as input and the dynamic search boundary as constraint. The future positions of the vehicle and each game stakeholders are gradually inferred along the inference step. At each step, the collision time margin is calculated based on the relative distance and relative approach speed between the vehicle and each game stakeholders, forming the collision time margin decay sequence. The reference passage time is taken from the moment when the vehicle passes through the center section of the intersection at a constant current speed. The ratio of the difference between the estimated time when the vehicle arrives at the center section of the intersection after performing the candidate driving action and the reference passage time to the preset look-ahead time is used as the equivalent passage deceleration. The conflict arbitration module is used to construct conflict urgency arbitration logic based on the collision time margin decay sequence and the equivalent traffic deceleration. When the minimum collision time margin is lower than the linear mapping threshold of the equivalent traffic deceleration, a risk avoidance instruction is output; otherwise, a compliance interaction instruction is output, including: From the multiple collision time margin decay sequences corresponding to the candidate driving actions of the autonomous vehicle, the minimum collision time margin of each collision time margin decay sequence is extracted to form a set of minimum collision time margins for candidate actions. Based on the vehicle's current speed and the lower limit of braking deceleration, the time required for the vehicle to decelerate from its current speed to a standstill through full braking is determined as the full braking collision avoidance time. For each candidate driving action of the autonomous vehicle, a linear mapping threshold is determined based on the equivalent traffic deceleration corresponding to the candidate action and the full braking collision avoidance time. The linear mapping threshold decreases as the equivalent traffic deceleration corresponding to the candidate action increases and increases as the full braking collision avoidance time increases. Traverse the set of minimum collision time margins for the candidate actions, and compare each minimum collision time margin with the corresponding linear mapping threshold for each candidate action one by one. When all the minimum collision time margins in the set of minimum collision time margins for candidate actions are less than their respective linear mapping thresholds, a risk avoidance instruction is output. When there is a minimum collision time margin in the set of candidate action collision time margins that is greater than or equal to its corresponding linear mapping threshold, the follow interaction instruction is output, and the candidate driving action of the vehicle corresponding to the minimum collision time margin that meets the condition is marked as an available follow action. The behavior generation module is used to generate executable driving behaviors at the current moment based on the risk avoidance instructions and the compliance interaction instructions, combined with the vehicle sensor data. The online calibration module is used to collect the actual interaction trajectory after the execution of the executable driving behavior, and use the cumulative deviation between the actual interaction trajectory and the inferred trajectory as the driving signal to reverse calibrate the time-series game inference network and the traffic interaction situation coding network.
Citation Information
Patent Citations
Game theory-based surrounding vehicle interaction behavior prediction method
CN111267846A
Local obstacle avoidance decision-making method and system based on dynamic risk map and real-time game
CN121626108A