Vehicle control method and device for intersection without traffic control and storage medium
Patent Information
- Application Number
- CN202610937198.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]现有决策方法多采用基于安全距离的硬性规则,当多车同时到达路口时,若各方系统都判定“有碰撞风险”而采取等待策略,就会形成“机器人死锁”,存在车辆在路口长期期停滞的风险
[0032]本发明通过从环境信息中提取与本车存在博弈的目标对象及目标对象的意图动作,实现了 “物理感知”到“语义意图理解”的跨越,使车辆具备了对复杂路口交互对象的社会属性预判能力,此外,引入“拟人化蠕行试探”机制,使得自动驾驶车辆能够像老司机一样,通过微小位移与对方进行博弈,从而打破路口停车等待的僵局,显著提升了通行效率。
Smart Images

Figure CN122821787A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of autonomous driving technology, and more specifically, this invention relates to a vehicle control method, device and storage medium for intersections without signal control. Background Technology
[0002] As autonomous driving technology evolves towards urban navigation-assisted driving (urban NOA), vehicles will inevitably need to handle a large number of unsignalized (no traffic lights) intersections, where right-of-way is often unclear.
[0003] Existing decision-making methods mostly adopt hard rules based on safe distance. When multiple vehicles arrive at the intersection at the same time, if all systems determine that there is a "collision risk" and adopt a waiting strategy, a "robot deadlock" will be formed, which poses a risk that vehicles will be stuck at the intersection for a long time. Summary of the Invention
[0004] In view of this, this application provides a vehicle control method for intersections without signal control, which aims to improve at least one of the above-mentioned problems.
[0005] Specifically, the following technical solutions are included:
[0006] On the one hand, embodiments of this application provide a vehicle control method for intersections without signal control, the method comprising:
[0007] (1) When the vehicle enters the interaction area without signal control, read the surrounding environment information perceived by the current vehicle, and extract the target object and the intention action of the target object that is playing a game with the vehicle from the environmental information.
[0008] (2) Match the vehicle’s optimal decision action in the current cycle based on the target object’s intention action. The decision action is to accelerate, crawl at low speed or brake.
[0009] In some embodiments of the present invention, the process of extracting the target object is as follows:
[0010] The interaction area without signal control is taken as the target area. Based on the past t consecutive environmental video frames, the driving trajectory of each object in the current environmental video frame in the target area is predicted. Based on radar data, the collision time or relative distance between the current vehicle and each object is predicted. Objects whose driving trajectory overlaps with the vehicle's driving trajectory in the target area and whose collision time or relative distance is less than the preset safe distance threshold are taken as target objects.
[0011] In some embodiments of the present invention, environmental video including the target object and the current kinematic state of the vehicle are input into the VLM model, and the VLM model identifies the intended action of the target object, wherein the intended action includes the intention to cut in, the intention to give way, and the undefined intention.
[0012] In some embodiments of the present invention, the VLM model outputs the intent probability distribution vector of each target object, wherein the intent probability distribution vector of the k-th target object is... It is expressed as follows:
[0013] ;
[0014] in, This represents the probability that the k-th target object intends to steal. This represents the probability that the k-th target object is hesitant or has unclear intentions; This represents the probability that the k-th target object has a yielding intention.
[0015] In some embodiments of the present invention, the attention map output by the VLM model is used for filtering target objects, retaining the target objects corresponding to the high-scoring regions in the attention map.
[0016] In some embodiments of the present invention, an expected benefit function is constructed based on the safety, traffic efficiency and interaction matching degree of the decision action, and the expected benefit score of the current vehicle under each decision action is calculated. The decision action with the highest expected benefit score is taken as the optimal decision action of the vehicle in the current period.
[0017] In some embodiments of the present invention, the expected return score The calculation formula is as follows:
[0018] ;
[0019] in, , , Indicates the weighting coefficient. Indicates the decision action taken by the current vehicle. Safety factors at that time Indicates the decision action taken by the current vehicle. The efficiency factor at that time This represents the probability distribution vector of the intent of the k-th target object. The j-th intentional action, Indicates the decision action currently taken by the vehicle. With intentional action The degree of interaction and matching between them.
[0020] In some embodiments of the present invention, the decision action taken by the current vehicle is calculated. The collision time with each target object is used to calculate the minimum collision time, which is then used to generate the safety factor.
[0021] On the other hand, embodiments of this application provide a vehicle control device for intersections without signal control, the device comprising:
[0022] The sensing unit is used to sense environmental information around the vehicle, including environmental images and radar data;
[0023] The target object extraction unit is used to extract target objects that are in a game with the vehicle from the environmental information around the vehicle when the vehicle is in an interaction area without signal control.
[0024] The intent prediction unit is used to predict the intended actions of the target object.
[0025] The decision unit is used to match the vehicle's optimal decision action in the current cycle based on the intentional actions of each target object. The decision action is to accelerate, crawl at low speed, or brake.
[0026] In some embodiments of the present invention, the intent prediction unit inputs the environmental video of the target object and the kinematic state of the current vehicle into the VLM model. The VLM model identifies the graph probability distribution vector of the target object, which includes the probability of the target object being in various intent actions, including: intention to cut in, intention to yield, and unclear intention.
[0027] In some embodiments of the present invention, the decision-making unit constructs an expected benefit function based on the safety, traffic efficiency and interaction matching degree of the decision action, calculates the expected benefit score of the current vehicle under each decision action, and takes the decision action with the highest expected benefit score as the optimal decision action of the vehicle in the current period.
[0028] In some embodiments of the present invention, the expected return score The calculation formula is as follows:
[0029] ;
[0030] in, , , Indicates the weighting coefficient. Indicates the decision action taken by the current vehicle. Safety factors at that time Indicates the decision action taken by the current vehicle. The efficiency factor at that time This represents the probability distribution vector of the intent of the k-th target object. The j-th intentional action, Indicates the decision action currently taken by the vehicle. With intentional action The degree of interaction and matching between them.
[0031] On the other hand, embodiments of this application provide a storage medium storing a computer program, which is executed by a processor to implement the vehicle control method for signalless intersections as described above.
[0032] This invention achieves a leap from "physical perception" to "semantic intent understanding" by extracting the target object and its intended action from environmental information that is interacting with the vehicle. This enables the vehicle to predict the social attributes of the interactive objects at complex intersections. In addition, the introduction of a "human-like crawling probing" mechanism allows autonomous vehicles to interact with other vehicles through small displacements, just like experienced drivers, thereby breaking the deadlock of waiting at intersections and significantly improving traffic efficiency. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 A flowchart of a vehicle control method for intersections without signal control provided in an embodiment of the present invention;
[0035] Figure 2 A schematic diagram of the structure of a vehicle control device for intersections without signal control provided in an embodiment of the present invention;
[0036] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise defined, all technical terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the art.
[0038] This invention extracts target objects that pose a game risk to the vehicle from environmental information, predicts the intentions of the target objects, matches the intentions of each target object in the interaction area without control to determine the optimal decision action of the vehicle in the current period, and controls the vehicle in the interaction area without control based on the current optimal decision action. Figure 1 A flowchart of a vehicle control method for intersections without signal control provided in an embodiment of the present invention is shown below:
[0039] (1) When the vehicle enters the interaction area without signal control, read the surrounding environment information perceived by the current vehicle, and extract the target object and the intention action of the target object that is playing a game with the vehicle from the environmental information.
[0040] When a vehicle enters an intersection without traffic lights, i.e., when the vehicle enters an interaction area without traffic control, the vehicle's surrounding environment video and radar data are read from the onboard vision sensors. Based on the surrounding environment video and radar data, the target object in the environment video that is interacting with the current vehicle is identified. The environment video including the target object and the kinematic state of the current vehicle are input into the VLM model. The VLM model recognizes the target object's intentional action. The current vehicle's kinematic state includes: current vehicle speed, longitudinal acceleration, lateral acceleration, steering wheel angle, and yaw rate.
[0041] In this embodiment of the invention, the method for determining the target object is as follows:
[0042] The uncontrolled interactive area is taken as the target area, and objects in the environmental video that pose a collision risk with the current vehicle within the target area are taken as the target objects. The objects are traffic participants, including motor vehicles, non-motor vehicles and pedestrians.
[0043] In this embodiment of the invention, the process of determining the target object in the environmental video is as follows:
[0044] Based on the past t consecutive environmental video frames, predict the driving trajectory of each object in the current environmental video frame within the target area. Based on radar data, predict the collision time or relative distance between the current vehicle and each object. Define the object whose driving trajectory overlaps with the vehicle's driving trajectory within the target area and whose collision time (TTC) or relative distance is less than a preset safe distance threshold (e.g., 50 meters) as the target object that is in a game with the current vehicle.
[0045] The VLM model extracts micro-semantic features of target objects through an attention mechanism, including vehicle wheel yaw angle and vehicle pitch attitude; pedestrian head orientation and whether they are looking at a mobile phone, etc. Instead of outputting a single physical trajectory, the VLM model outputs the intent probability distribution vector for each target object. The intent actions of the target objects include: intention to cut in, intention to yield, and unclear intent. The probability distribution vector of the intent of the k-th target object is shown below. The specific details are as follows:
[0046] ;
[0047] in, This represents the probability that the k-th target object intends to steal. This represents the probability that the k-th target object is hesitant or has unclear intentions; Let represent the probability that the k-th target object has a yielding intention, and .
[0048] In this embodiment of the invention, in the interaction area without signal control, it is necessary to identify target objects in the environment based on visual images captured by cameras and radar data captured by radar. However, the data acquisition frequencies of cameras and radar are different, requiring temporal and spatial alignment of the acquired visual images and radar data. Spatial alignment involves transforming the visual images and radar data to the same global coordinate system. When processing parallel inputs from multiple cameras, the spatial attention mechanism in the feature encoding and fusion network layer of the VLM model fuses environmental images from different perspectives to form a compact semantic descriptor with multi-view features. This descriptor includes camera image sequences from the forward-facing main view, left-front-side view, and right-front-side view. This semantic descriptor not only contains the coordinate vectors of obstacles but also includes a Transformer-based attention map, which is extracted from the self-attention layer within the VLM model. The intermediate feature output of the VLM model (Layer) represents the probability matrix of attention to each pixel region of the image during inference. During real-time inference, the VLM model dynamically calculates the feature attention score generated by the spatial attention mechanism. This is a dynamic matrix that changes with the real-time image. The feature matrix values of regions with high interactive value in the video are significantly amplified, while the values of irrelevant regions approach 0. Based on this, after the target object is detected based on video images and radar data, the target object is further filtered based on the attention map output by the VLM model. Only the target objects corresponding to the high-score regions in the attention map are retained. This ensures that the VLM model can accurately focus on "target objects" with high interactive value when facing complex intersections with mixed pedestrians and vehicles and filtering out the background corresponding to non-interactive regions. This effectively reduces the computational cost and improves the accuracy of intent decoding.
[0049] (2) Match the vehicle’s optimal decision action in the current cycle based on the target object’s intention action, and control the current vehicle based on the optimal decision action.
[0050] In this embodiment of the invention, decision-making actions within the action space include: acceleration, slow creeping, and braking, wherein acceleration corresponds to cutting in, slow creeping corresponds to probing / unclear intent, and braking corresponds to yielding. , Indicates an acceleration action. This indicates slow, creeping motion. This indicates a braking action.
[0051] To break rule deadlock, this invention constructs an expected reward function based on the safety of decision actions, traffic efficiency, and interaction matching degree with the target object. This function is applied to any decision action of the current vehicle. Give an expected return score. Calculate the expected reward score for each decision action of the vehicle in the action space, and take the decision action with the highest expected reward score as the optimal decision action for the vehicle in the current cycle. The expected reward score is... The specific calculation formula is as follows:
[0052] ;
[0053] in, , , Indicates the weighting coefficient. Indicates the decision action taken by the current vehicle. The higher the security level, the larger the corresponding security factor. Indicates the decision action taken by the current vehicle. The efficiency factor is determined by the traffic flow efficiency; the higher the traffic flow efficiency, the larger the corresponding efficiency factor. This represents the probability distribution vector of the intent of the k-th target object. The j-th intention action has an intention probability distribution vector output by the VLM model. Indicates the decision action currently taken by the vehicle. With intentional action The degree of interaction and matching between them.
[0054] In this embodiment of the invention, the decision action taken by the current vehicle is calculated. The collision time with each target object is used to calculate the minimum collision time. This minimum collision time is used to generate the safety factor. The higher the minimum collision time, the higher the corresponding safety factor. The larger. The current vehicle's decision-making action. The faster the corresponding action speed, the higher the efficiency factor. The larger the value, the more likely the vehicle is to accelerate. If the current traffic flow is efficient, then the efficiency factor of the decision-making action is high. Interaction matching degree The determination is made by looking up a table, which records the various decision-making actions of the current vehicle. Various intentions of the target object In terms of matching degree during interaction, the higher the security of the interaction between the two parties, the higher the matching degree. For example, if the target object's intent... It's "rushing," and the current vehicle's decision-making action... When set to "accelerate", the interaction matching degree is... It is a huge negative number; if the target object's intention It is "give way", and the current vehicle's decision action When set to "accelerate", the interaction matching degree is... It is positive.
[0055] When the VLM model determines that the target object's intention is "hesitant" (i.e., the probability in the intention probability distribution vector is...), At its maximum, after calculation using the above formula, the current vehicle's decision-making action... Benefits of (slow crawling) At its highest level, a small speed limit command of 2 km / h to 5 km / h is issued to the chassis, causing the current vehicle to adopt a slow, probing forward posture. During this creeping probing period, the VLM model is invoked at a higher frequency (e.g., 20Hz) to monitor the microscopic reactions of the target object. If the target object exhibits significant braking deceleration and accompanying changes in the vehicle's longitudinal pitch angle (i.e., a braking "nodding" posture caused by suspension compression), the yield probability in the intent probability distribution vector output by the VLM model will be calculated. This will increase rapidly, along with the yield probability in the intention probability distribution vector. Increased probability, decision-making action The expected return score of (acceleration) surpasses that of the previous state, and the current vehicle automatically and smoothly transitions its decision state from "creeping" to "acceleration" to complete efficient interaction within the target area. If, during the current vehicle's exploration, a target object suddenly accelerates and approaches, and the physical safety distance is detected to be insufficient, the VLM model's output decision action is directly blocked, and automatic emergency braking (AEB) is directly triggered to ensure the physical safety baseline.
[0056] Intent Probability Calibration Mechanism: To reduce the output digital display jumps of the VLM model under adverse weather or lighting conditions, this invention introduces a spatiotemporal consistency calibration strategy, specifically: using Long Short-Term Memory (LSTM) to adjust the intent probability distribution vector. Temporal smoothing is performed to suppress output intent jitter caused by perceptual noise, maximizing the smoothness of vehicle acceleration and deceleration and reducing the possibility of "twitchy" driving. Confidence dynamic weighting: A confidence assessment model is constructed to monitor the confidence level of the intent probability distribution vector output by the VLM model in real time. Information entropy is used to calculate the uncertainty of the VLM model output. If the confidence level is below a threshold, the weight of physical rule decisions is automatically increased, i.e., the weight of VLM model output decisions is decreased. This is achieved by slowing down the vehicle speed (from crawling to a standstill) to "gain more perception time," ensuring that driving safety is prioritized when perception is ambiguous.
[0057] To verify the anthropomorphic interactive decision-making logic of the present invention, the following typical interactive scenarios were selected for specific implementation:
[0058] Scenario A (Oncoming Traffic Response in a Narrow Single-Lane Alley): When the vehicle detects an oncoming vehicle (target) and perceives that it is moving at a very low speed and is pausing, the VLM model recognizes the "hesitant intention" of the oncoming vehicle. At this time, the vehicle's decision action is "creeping probing," moving forward at a speed of 0.5 m / s. During the movement, the brake light status and suspension drop of the oncoming vehicle (vehicle attitude characteristics) are continuously monitored. If the oncoming vehicle's suspension is detected to be tilting forward, it is determined that the oncoming vehicle has braked successfully, and the vehicle's decision action is automatically switched to accelerate through. If the oncoming vehicle continues to approach slowly, and the physical safety distance between the two is insufficient, the VLM model's output decision action is directly blocked, and automatic emergency braking (AEB) is directly triggered. Scenario B (Blind Spot Detection in Mixed Pedestrian and Vehicle Traffic): At intersections where pedestrians (targets) are obscured by stationary vehicles, pedestrians cannot be directly observed. At this point, leveraging the inference capabilities of the VLM model, the "partial human afterimage" is used as input. Based on the VLM model, the pedestrian's intention probability distribution vector is obtained, triggering an "active detection action." This involves changing the vehicle's line-of-sight angle through low-speed creeping to fill in blind spots, and updating the information in the game matrix in real time. Scenario C (Robust switching in extreme weather scenarios): When the perception system determines that it is in a low-visibility environment such as rain or snow, the sensor confidence decreases. At this time, the safety weight W1 in the comprehensive benefit function is automatically increased. In the game decision-making, the original "creeping probing" strategy is directly downgraded to a "conservative waiting" strategy, prioritizing the vehicle's static safety within the physical safety boundary and avoiding decision-making errors caused by blurred visual semantic information.
[0059] To ensure real-time interactive decision-making, this invention employs an asynchronous pipelined architecture, distributing perceptual computation and game optimization computation across different hardware processing units, communicating at high speed via a shared memory mechanism. More importantly, this system integrates a functional safety monitoring module (Safety Arbiter) compliant with the ISO 26262 standard. This module operates independently of the perceptual game decision-making logic, monitoring the dynamic feasibility of game actions in real time. If the action command obtained from solving the game payoff function violates the underlying physical safety barrier (e.g., causing the TTC to exceed a preset limit), the functional safety monitoring module will have the highest priority to interrupt, forcibly blocking the game decision command and triggering the minimum risk maneuver (MRM) action for the current scenario, achieving a zero-latency switch from "intelligent AI interaction" to "safe physical braking."
[0060] Figure 2 This is a schematic diagram of a vehicle control device for intersections without signal control provided in an embodiment of the present invention. For ease of explanation, only the parts related to the embodiment of the present invention are shown. The device includes:
[0061] The system comprises a perception unit, a target object extraction unit, an intent prediction unit, and a decision-making unit. The perception unit is used to perceive environmental information around the vehicle, including environmental images and radar data. When the vehicle is in an interaction area without signal control, the target object extraction unit is used to extract target objects that are interacting with the vehicle from the environmental information around the vehicle. The intent prediction unit is used to predict the intent actions of the target objects. The decision-making unit is used to match the optimal decision action of the vehicle in the current period based on the intent actions of each target object.
[0062] When a vehicle enters an intersection without traffic lights, i.e., when the vehicle enters an uncontrolled interaction area, the sensing unit reads environmental video and radar data from the surrounding environment. The target object extraction unit extracts target objects that are at odds with the vehicle from the surrounding video and radar data. The uncontrolled interaction area is designated as the target area, and objects in the environmental video that pose a collision risk to the vehicle within the target area are identified as target objects. These objects are traffic participants, including motor vehicles, non-motor vehicles, and pedestrians. The target object extraction unit predicts the trajectory of each object in the current environmental video frame within the target area based on t consecutive environmental video frames from the past. It also predicts collision events or relative distances between the vehicle and each object based on radar data. Objects whose trajectories overlap with the vehicle's trajectory within the target area and whose time to collision (TTC) or relative distance is less than a preset safe distance threshold (e.g., 50 meters) are defined as target objects at odds with the vehicle.
[0063] The intent prediction unit inputs environmental video of the target object and the current kinematic state of the vehicle into the VLM model. The VLM model identifies the target object's action intent. The current vehicle kinematic state includes: current speed, longitudinal acceleration, lateral acceleration, steering wheel angle, and yaw rate. The VLM model extracts micro-semantic features of the target object through an attention mechanism, including the vehicle's wheel yaw angle and vehicle pitch attitude; the pedestrian's head orientation, whether they are looking at a mobile phone, etc. The VLM model no longer outputs a single physical trajectory, but instead outputs the intent probability distribution vector of each target object. The target object's intent actions include: cutting in, yielding, and unclear intent. Among them, the intent probability distribution vector of the k-th target object is... The specific details are as follows:
[0064] ;
[0065] in, This represents the probability that the k-th target object intends to steal. This represents the probability that the k-th target object is hesitant or has unclear intentions; Let represent the probability that the k-th target object has a yielding intention, and .
[0066] In the uncontrolled interaction area, the target object extraction unit needs to identify target objects in the environment based on visual images captured by cameras and radar data collected by radar. However, the data acquisition frequencies of cameras and radar are different, requiring temporal and spatial alignment of the acquired visual images and radar data. Spatial alignment involves transforming the visual images and radar data to the same global coordinate system. When processing parallel inputs from multiple cameras, the spatial attention mechanism in the feature encoding and fusion network layer of the VLM model fuses environmental images from different perspectives to form a compact semantic descriptor with multi-view features. This includes camera image sequences from the front-facing main view, left front-side view, and right front-side view. This semantic descriptor not only contains the coordinate vectors of obstacles but also includes a Transformer-based attention map, which is extracted from the self-attention layer within the VLM model. The intermediate feature output of the Layer represents the attention probability matrix of each pixel region of the image during inference by the VLM model. During real-time inference, the spatial attention mechanism of the VLM model dynamically calculates the feature attention score, which is a dynamic matrix that changes with the real-time image. The feature matrix values of regions with high interactive value in the video are significantly amplified, while the values of irrelevant regions approach 0. Based on this, after the target object extraction unit completes the detection of target objects based on video images and radar data, it further filters the target objects based on the attention map output by the VLM model, retaining only target objects with high feature attention scores. This ensures that the VLM model can accurately focus on "target objects" with high interactive value when facing complex intersections with mixed pedestrians and vehicles, filtering out the background corresponding to non-interactive regions, effectively reducing the computational cost and improving the accuracy of intent decoding.
[0067] In this embodiment of the invention, decision-making actions within the action space include: acceleration, slow creeping, and braking, wherein acceleration corresponds to cutting in, slow creeping corresponds to probing / unclear intent, and braking corresponds to yielding. , Indicates an acceleration action. This indicates slow, creeping motion. This indicates a braking action.
[0068] To break rule deadlock, the decision-making unit constructs an expected reward function based on the safety of the decision action, traffic efficiency, and the degree of interaction matching with the target object. This function is applied to any decision action taken by the current vehicle. Give an expected return score. Calculate the expected reward score for each decision action of the vehicle in the action space, and take the decision action with the highest expected reward score as the optimal decision action for the vehicle in the current period. The expected reward score is... The specific calculation formula is as follows:
[0069] ;
[0070] in, , , Indicates the weighting coefficient. Indicates the decision action taken by the current vehicle. The higher the security level, the larger the corresponding security factor. Indicates the decision action taken by the current vehicle. The efficiency factor is determined by the traffic flow efficiency; the higher the traffic flow efficiency, the larger the corresponding efficiency factor. This represents the probability distribution vector of the intent of the k-th target object. The j-th intent is represented by the VLM model output, and the intent probability distribution vector is the value of the j-th intent. Indicates the decision action currently taken by the vehicle. With intention Interaction matching degree.
[0071] In this embodiment of the invention, the decision action taken by the current vehicle is calculated. The collision time with each target object is used to calculate the minimum collision time. This minimum collision time is used to generate the safety factor. The higher the minimum collision time, the higher the corresponding safety factor. The larger the value, the better. The vehicle is currently making a decision. The faster the corresponding action speed, the higher the efficiency factor. The larger the value, the more likely the vehicle is to accelerate. If the current traffic flow is efficient, then the efficiency factor of the decision-making action is high. Interaction matching degree The determination is made by looking up a table, which records the various decision-making actions of the current vehicle. Various intentions of the target object In terms of matching degree during interaction, the higher the security of the interaction between the two parties, the higher the matching degree. For example, if the target object's intent... It's "rushing," and the current vehicle's decision-making action... When set to "accelerate", the interaction matching degree is... It is a huge negative number; if the target object's intention It is "yielding," and the current vehicle's decision-making action. When set to "accelerate", the interaction matching degree is... It is positive.
[0072] When the VLM model determines that the target object's intention is "hesitant" (i.e., the probability in the intention probability distribution vector is...), At its maximum, after calculation using the above formula, the current vehicle's decision-making action... Benefits of (slow crawling) At its highest level, a small speed limit command of 2 km / h to 5 km / h is issued to the chassis, causing the current vehicle to adopt a slow, probing forward posture. During this creeping probing period, the VLM model is invoked at a higher frequency (e.g., 20Hz) to monitor the microscopic reactions of the target object. If the target object exhibits significant braking deceleration and accompanying changes in the vehicle's longitudinal pitch angle (i.e., a braking "nodding" posture caused by suspension compression), the yield probability in the intent probability distribution vector output by the VLM model will be calculated. This will increase rapidly, along with the yield probability in the intention probability distribution vector. Increased probability, decision-making action The expected benefit score of (acceleration) surpasses that of the previous one, and the current vehicle automatically and smoothly transitions the decision state from "crawl" to "acceleration" to complete the efficient interaction within the target area.
[0073] In this embodiment of the invention, the device further includes: a kinematic safety verification unit, used to directly forcibly block the VLM model's output decision action and directly trigger automatic emergency braking (AEB) when the physical safety distance between the target object and the current vehicle is insufficient, ensuring the physical safety baseline. If, during the current vehicle's probing period, a target object suddenly accelerates and approaches, and insufficient physical safety distance is detected, the VLM model's output decision action is directly blocked, and automatic emergency braking (AEB) is directly triggered, ensuring the physical safety baseline.
[0074] In this embodiment of the invention, the device further includes: an intent probability calibration unit, used to calibrate the intent probability distribution vector output by the intent prediction unit. To smooth the time dimension and reduce the output digital display jumps of the VLM model under adverse weather or lighting conditions, this invention introduces a spatiotemporal consistency calibration strategy. Specifically, it uses Long Short-Term Memory (LSTM) pairs to suppress output intention jitter caused by perceptual noise, thereby improving the smoothness of vehicle acceleration and deceleration as much as possible and reducing the possibility of "twitchy" driving.
[0075] In this embodiment of the invention, the device further includes:
[0076] The centroid assessment unit is used to evaluate the confidence level of the intent probability distribution vector output by the VLM model. It uses information entropy to calculate the uncertainty of the VLM model output. If the confidence level is lower than the threshold, it automatically increases the weight of the physical rule decision, that is, reduces the weight of the VLM model output decision. By slowing down the vehicle speed (from crawling to standing still), it "gains more perception time" to ensure that driving safety is prioritized when perception is ambiguous.
[0077] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0078] This device achieves a leap from "physical perception" to "semantic intent understanding" by extracting the target objects and their intentional actions from environmental information that are interacting with the vehicle. This enables the vehicle to predict the social attributes of interactive objects at complex intersections and introduces a "human-like crawling probing" mechanism, allowing autonomous vehicles to interact with other vehicles through minute movements, just like experienced human drivers, thereby breaking the deadlock of waiting at intersections and significantly improving traffic efficiency.
[0079] One embodiment of this application provides a terminal device including a processor and a memory. The processor may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor may also include a main processor and a coprocessor. The main processor is used to process data in the wake-up state, also known as a central processing unit (CPU); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, the processor may also include an AI processor, which is used to handle computational operations related to machine learning. The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, a non-transitory computer-readable storage medium in the memory is used to store a computer program configured to be executed by one or more processors to implement the above-described vehicle control method for intersections without signal control.
[0080] In some embodiments, the terminal device may also optionally include: a peripheral device interface and at least one peripheral device. The processor, memory, and peripheral device interface can be connected via a bus or signal lines. Each peripheral device can be connected to the peripheral device interface via a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: radio frequency circuitry, a display screen, audio circuitry, and a power supply. Those skilled in the art will understand that the above structure does not constitute a limitation on the terminal device, and may include more or fewer components than illustrated, or combine certain components, or employ different component arrangements.
[0081] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, and the computer program, when executed by a processor, implements the aforementioned vehicle control method for intersections without signal control. Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drive (SSD), or optical disk, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).
[0082] In an exemplary embodiment, a computer program product is also provided, comprising a computer program stored in a computer-readable storage medium. A processor of a terminal device reads the computer program from the computer-readable storage medium and executes the computer program, causing the terminal device to perform the aforementioned vehicle control method for intersections without signal control.
[0083] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only.
[0084] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A vehicle control method for signal-free intersections, characterized in that, The method includes: (1) When the vehicle enters the interaction area without signal control, read the surrounding environment information perceived by the current vehicle, and extract the target object and the intention action of the target object that is playing a game with the vehicle from the environmental information. (2) Match the vehicle’s optimal decision action in the current cycle based on the target object’s intention action. The decision action is to accelerate, crawl at low speed or brake.
2. The vehicle control method for intersections without signal control as described in claim 1, characterized in that, The process of extracting the target object is as follows: The interaction area without signal control is taken as the target area. Based on the past t consecutive environmental video frames, the driving trajectory of each object in the current environmental video frame in the target area is predicted. Based on radar data, the collision time or relative distance between the current vehicle and each object is predicted. Objects whose driving trajectory overlaps with the vehicle's driving trajectory in the target area and whose collision time or relative distance is less than the preset safe distance threshold are taken as target objects.
3. The vehicle control method for intersections without signal control as described in claim 1, characterized in that, The VLM model inputs environmental video of the target object and the current kinematic state of the vehicle. The VLM model identifies the target object's intentional actions, which include the intention to cut in, the intention to yield, and unclear intentions.
4. The vehicle control method for intersections without signal control as described in claim 3, characterized in that, The VLM model outputs the intent probability distribution vector for each target object, where the intent probability distribution vector for the k-th target object is... It is expressed as follows: ; in, This represents the probability that the k-th target object intends to steal. This represents the probability that the k-th target object is hesitant or has unclear intentions; This represents the probability that the k-th target object has a yielding intention.
5. The vehicle control method for intersections without signal control as described in claim 3, characterized in that, The attention map output by the VLM model is used to filter target objects, retaining the target objects corresponding to the high-scoring regions in the attention map.
6. The vehicle control method for intersections without signal control as described in claim 1, characterized in that, Based on the safety, traffic efficiency, and interaction matching degree of the decision action, an expected benefit function is constructed. The expected benefit score of the current vehicle under each decision action is calculated, and the decision action with the highest expected benefit score is taken as the optimal decision action of the vehicle in the current period.
7. The vehicle control method for intersections without signal control as described in claim 6, characterized in that, Expected return score The calculation formula is as follows: ; in, , , Indicates the weighting coefficient. Indicates the decision action taken by the current vehicle. Safety factors at that time Indicates the decision action taken by the current vehicle. The efficiency factor at that time This represents the probability distribution vector of the intent of the k-th target object. The j-th intentional action, Indicates the decision action currently taken by the vehicle. With intentional action The degree of interaction and matching between them.
8. The vehicle control method for intersections without signal control as described in claim 7, characterized in that, Calculate the decision action to be taken by the current vehicle The collision time with each target object is used to calculate the minimum collision time, which is then used to generate the safety factor.
9. A vehicle control device for intersections without signal control, characterized in that, The device includes: The sensing unit is used to sense environmental information around the vehicle, including environmental images and radar data; The target object extraction unit is used to extract target objects that are in a game with the vehicle from the environmental information around the vehicle when the vehicle is in an interaction area without signal control. The intent prediction unit is used to predict the intended actions of the target object. The decision unit is used to match the vehicle's optimal decision action in the current cycle based on the intentional actions of each target object. The decision action is to accelerate, crawl at low speed, or brake.
10. The vehicle control device for intersections without signal control as described in claim 9, characterized in that, The intent prediction unit inputs the environmental video of the target object and the current kinematic state of the vehicle into the VLM model. The VLM model identifies the graph probability distribution vector of the target object, which includes the probability of the target object being in various intent actions, including: cutting in, yielding, and unclear intent.
11. The vehicle control device for intersections without signal control as described in claim 9, characterized in that, The decision-making unit constructs an expected benefit function based on the safety, traffic efficiency, and interaction matching degree of the decision action with the target object, calculates the expected benefit score of the current vehicle under each decision action, and takes the decision action with the highest expected benefit score as the optimal decision action of the vehicle in the current period.
12. The vehicle control device for intersections without signal control as described in claim 11, characterized in that, Expected return score The calculation formula is as follows: ; in, , , Indicates the weighting coefficient. Indicates the decision action taken by the current vehicle. Safety factors at that time Indicates the decision action taken by the current vehicle. The efficiency factor at that time This represents the probability distribution vector of the intent of the k-th target object. The j-th intentional action, Indicates the decision action currently taken by the vehicle. With intentional action The degree of interaction and matching between them.
13. A storage medium, characterized in that, The storage medium stores a computer program, which is executed by a processor to implement the vehicle control method for unsignalized intersections as described in any one of claims 1 to 8.