A hierarchical retrieval and guidance visual world model planning method and system for automatic parking
Patent Information
- Application Number
- CN202611033359.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-07-13
AI Technical Summary
[0008]本申请实施例提供一种面向自动泊车的分层检索引导视觉世界模型规划方法及系统,解决现有视觉世界模型直接用于自动泊车时存在长程预测误差大、随机采样规划效率低以及偏离专家轨迹后恢复能力不足的问题
本申请实施例提供一种面向自动泊车的分层检索引导视觉世界模型规划方法及系统,包括三个阶段,分别为:阶段一、获取自动泊车离线轨迹数据,构建自动泊车动作条件视觉世界模型,并根据泊车任务进程生成多个阶段子目标,得到分层泊车子目标序列;阶段二、在车辆在线泊车过程中,根据当前车辆状态和阶段一生成的与当前车辆状态对应的阶段子目标,从离线泊车轨迹检索库中检索与当前车辆状态和当前阶段子目标相匹配的动作片段,将检索得到的动作序列作为模型预测路径积分控制 MPPI 的采样先验,并利用阶段一构建的动作条件视觉世界模型预测候选动作序列的未来状态,输出最优动作序列并执行,得到闭环执行结果;阶段三、根据阶段二的闭环执行结果检测车辆是否处于偏离状态,若处于偏离状态,则生成恢复轨迹、终端姿态对齐轨迹和带噪声纠偏轨迹,并基于生成的轨迹对所述动作条件视觉世界模型和/或所述离线泊车轨迹检索库进行增强。
Smart Images

Figure CN122519256B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to a hierarchical retrieval-guided visual world model planning method and system for automatic parking. Background Technology
[0002] Automated parking is a crucial function for intelligent vehicles in low-speed, structured scenarios. Unlike conventional road driving, automated parking typically requires the vehicle to approach a parking space, reverse into the space, straighten the steering wheel, align the vehicle's final position, and maintain a stop within a confined space. This process is characterized by low speed, large turning angles, frequent switching between forward and reverse maneuvers, and high requirements for final position and heading accuracy.
[0003] Existing automated parking methods mainly include geometric rule-based parking path planning methods, model predictive control methods based on vehicle dynamics models, and data-driven learning control methods. Geometric rule-based methods typically rely on predefined curves and strong scenario assumptions. When there are significant changes in initial pose, complex parking space structures, or deviations in vehicle execution, problems such as discontinuous paths, unreasonable reversing timing, or large terminal posture errors can easily occur. While explicit vehicle model-based control methods offer some interpretability, they are still susceptible to the effects of low-speed tire characteristics, control delays, scenario constraints, and the accumulation of local deviations during real-world vehicle execution.
[0004] In recent years, visual world models have been able to learn the dynamic relationship between visual states and actions using offline trajectory data, and optimize actions through model predictive control during testing. This type of method provides a new technical path for automated parking, where vehicles can predict future parking states in the visual feature space and select control actions based on the prediction results. However, directly applying existing visual world models to automated parking still has the following shortcomings: First, automated parking tasks have a clear long-term phased nature. If the planner plans directly from the initial state to the final parking space, the world model needs to maintain accurate predictions over a long time scale, which can easily lead to the accumulation of prediction errors, thus affecting the reversing into the parking space and the terminal alignment effect.
[0005] Second, automatic parking actions are highly structured. Effective actions such as reversing, large turning angles, low-speed fine adjustments, and steering wheel straightening account for a relatively small proportion of the action space. If only random sampling model predictive control is used, it is difficult to efficiently search for action sequences that meet parking requirements.
[0006] Third, offline expert trajectories typically only cover the ideal parking process. When the vehicle deviates from the expert trajectory during closed-loop execution due to model errors, sampling errors, or control errors, the visual world model lacks the ability to predict and recover from the deviation, which can easily lead to further failures in subsequent planning.
[0007] Therefore, there is an urgent need for a visual world model planning method and system for automatic parking scenarios, which can reduce the difficulty of long-distance parking planning, improve the search efficiency of key parking actions, and enhance the vehicle's recovery control capability in closed-loop deviation states. Summary of the Invention
[0008] This application provides a hierarchical retrieval-guided visual world model planning method and system for automatic parking, which solves the problems of large long-range prediction errors, low efficiency of random sampling planning, and insufficient recovery ability after deviating from expert trajectories when existing visual world models are directly used for automatic parking.
[0009] To address the aforementioned technical problems, in a first aspect, embodiments of this application provide a hierarchical retrieval-guided visual world model planning method for automatic parking, comprising: first, acquiring offline trajectory data for automatic parking, constructing an automatic parking action condition visual world model, and generating multiple stage sub-objectives according to the parking task progress to obtain a hierarchical parking sub-objective sequence; second, during the online parking process, based on the current vehicle state and the stage sub-objectives generated in stage one corresponding to the current vehicle state, retrieving action segments matching the current vehicle state and the current stage sub-objectives from the offline parking trajectory retrieval library, using the retrieved action sequences as sampling priors for the Model Predicted Path Integral Control (MPPI), and using the action condition visual world model constructed in stage one to predict the future state of candidate action sequences, outputting the optimal action sequence and executing it to obtain a closed-loop execution result; finally, detecting whether the vehicle is in a deviated state based on the closed-loop execution result of stage two, if it is in a deviated state, generating a recovery trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory, and enhancing the action condition visual world model and / or the offline parking trajectory retrieval library based on the generated trajectories.
[0010] In some exemplary embodiments, constructing an automatic parking action-conditional visual world model includes: acquiring automatic parking offline trajectory data; the offline trajectory data includes vehicle visual observations, vehicle status, vehicle control actions, and target parking space information; based on the offline trajectory data, converting the parking scene into a unified visual observation representation, and extracting latent visual features using a pre-trained visual encoder; based on the latent visual features, constructing an action-conditional visual world model, enabling it to predict future latent visual features based on historical latent visual features, vehicle status, and control actions; and dividing the complete parking task into multiple stage sub-objectives based on the parking task progress, expert trajectory keyframes, the positional difference between the vehicle's current position and the target parking space position, and the heading difference between the vehicle's current heading and the target parking space's orientation, generating a hierarchical parking sub-objective sequence for stage two online parking planning.
[0011] In some exemplary embodiments, the stage sub-targets include one or more of the following: parking space entrance arrival sub-target, reversing preparation sub-target, reversing into the parking space sub-target, vehicle body straightening sub-target, and terminal posture alignment sub-target; the stage sub-targets are used to represent the visual image of the sub-target, the vehicle pose of the sub-target, or the visual latent features of the sub-target.
[0012] In some exemplary embodiments, an action-conditional visual world model is constructed. Based on historical visual latent features, historical vehicle states, and historical actions, predict the visual latent features for the next moment:
[0013] in, Indicates the current moment. Indicates the length of the history window. Represents a visual world model with action conditions. Represents the parameters of the visual world model. Indicates from time At that time Historical visual latent feature sequences Indicates from time At that time The historical sequence of vehicle control actions, Indicates from time At that time Historical vehicle state sequence This represents the predicted visual latent features for the next moment.
[0014] In some exemplary embodiments, during training, latent features are used to predict the loss:
[0015] in, This represents the latent feature prediction loss. This represents the time step in the trajectory data. This represents the summation of prediction errors at each time step in the trajectory data. This represents the visual latent features predicted by the action-conditional visual world model for the next moment. This represents the real latent visual features extracted by the pre-trained visual encoder from real visual observations at the next time step. Let L2 be the squared norm; by minimizing the above loss, the visual world model learns the dynamic relationship between visual state and vehicle movement during automatic parking.
[0016] In some exemplary embodiments, MPPI parking planning based on offline trajectory retrieval priors includes: segmenting multiple action segments from offline parking trajectories to establish a parking trajectory retrieval library; during online parking, retrieving action segments matching the current vehicle state and current stage sub-objective from the parking trajectory retrieval library based on the current vehicle state and current stage sub-objective, and using the retrieved action sequences as sampling priors for model prediction path integral control (MPPI); adding disturbances near the priors to generate multiple sets of candidate action sequences; using a visual world model to predict the future visual potential state corresponding to each candidate action sequence, and calculating MPPI weights based on the cost between the predicted state and the current stage sub-objective to obtain the optimal action sequence; the vehicle executing the first action or the first macro action in the optimal action sequence using a rolling time-domain control method, and replanning in the next control cycle.
[0017] In some exemplary embodiments, each action segment includes a segment start state, a segment end state, the parking stage to which it belongs, an action sequence, and a success level label.
[0018] In some exemplary embodiments, data augmentation based on closed-loop deviation states includes: during the vehicle's closed-loop execution, if the vehicle deviates from the expert trajectory, fails to complete a stage sub-objective, overshoots the target, fails to make a large low-speed turn, or has excessive terminal posture error, the corresponding state is recorded as the starting point for data recovery; a recovery trajectory, a terminal posture alignment trajectory, and a noisy correction trajectory are generated based on the recovery data starting point, and these are added to the data recovery augmentation dataset to enhance the visual world model.
[0019] In some exemplary embodiments, when enhancing the training of the visual world model, recovered trajectory segments that meet the success or near-success conditions are added to the parking trajectory retrieval library to update the trajectory retrieval library and provide more reliable action priors for subsequent MPPI planning.
[0020] Secondly, embodiments of this application also provide a hierarchical retrieval-guided visual world model planning system for automatic parking. This system is used to implement the hierarchical retrieval-guided visual world model planning method for automatic parking described in the above embodiments. The system includes: an automatic parking offline modeling module, an MPPI online planning module, and a closed-loop deviation state recovery enhancement module connected in sequence. The automatic parking offline modeling module is used to construct a visual world model of automatic parking action conditions and generate multiple stage sub-objectives according to the parking task progress to obtain a hierarchical parking sub-objective sequence. The MPPI online planning module is used to execute MPPI parking planning based on the prior information retrieved from the offline trajectory. The closed-loop deviation state recovery enhancement module is used to perform recovery data enhancement based on the closed-loop deviation state.
[0021] The technical solution provided in this application has at least the following advantages: This application provides a hierarchical retrieval-guided visual world model planning method and system for automatic parking, comprising three stages: Stage 1: Acquiring offline trajectory data for automatic parking, constructing an automatic parking action condition visual world model, and generating multiple stage sub-objectives according to the parking task progress to obtain a hierarchical parking sub-objective sequence; Stage 2: During the online parking process, based on the current vehicle state and the stage sub-objectives generated in Stage 1 corresponding to the current vehicle state, retrieving action segments matching the current vehicle state and the current stage sub-objectives from the offline parking trajectory retrieval library, using the retrieved action sequences as sampling priors for the Model Predicted Path Integral Control (MPPI), and using the action condition visual world model constructed in Stage 1 to predict the future state of candidate action sequences, outputting the optimal action sequence and executing it to obtain a closed-loop execution result; Stage 3: Detecting whether the vehicle is in a deviated state based on the closed-loop execution result of Stage 2. If it is in a deviated state, generating a recovery trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory, and enhancing the action condition visual world model and / or the offline parking trajectory retrieval library based on the generated trajectories.
[0022] This application utilizes a motion-conditional visual world model to predict the future parking state of a vehicle in a latent visual feature space; it divides the complete parking process into multiple stage sub-objectives to reduce the difficulty of long-term planning; during online planning, it retrieves motion segments from an offline parking trajectory library that match the current vehicle state and stage sub-objectives, and uses these segments as priors for MPPI sampling optimization; simultaneously, it generates recovery data based on deviation states during closed-loop execution to enhance the visual world model and trajectory retrieval library, thereby improving the stability of automatic parking closed-loop control.
[0023] The hierarchical retrieval-guided visual world model planning method and system for automatic parking provided in this application have the following advantages compared with the prior art.
[0024] (1) Reduce the difficulty of long-range prediction for automatic parking.
[0025] This application divides the complete parking task into multiple stage sub-objectives, so that the visual world model only needs to plan to the current stage objective in each control cycle, reducing the long-range error accumulation caused by directly predicting from the initial state to the final parking space, and improving the planning stability of the reversing into the parking space and the terminal alignment stage.
[0026] (2) Improve the sampling efficiency of key parking actions.
[0027] This application introduces offline parking trajectory retrieval priors into MPPI planning. Based on the current vehicle state and stage sub-objectives, it retrieves action segments that match the current vehicle state and stage sub-objectives, using the retrieved action sequences as the sampling distribution center. Compared to direct sampling from a random distribution, this method increases the probability of sampling key actions such as reversing, large turns, low-speed fine-tuning, and steering correction, thereby improving parking action search efficiency.
[0028] (3) Enhance the recovery capability under closed-loop deviation state.
[0029] This application addresses potential issues during closed-loop vehicle execution, such as deviation from the expert trajectory, over-reversing, failure to complete stage objectives, and excessive terminal posture errors. It generates recovered trajectories, terminal posture aligned trajectories, and noisy correction trajectories, which are then used for visual world model enhancement training and retrieval library updates. As a result, the model can cover deviation state distributions beyond the expert trajectory, improving the vehicle's ability to recover from non-ideal states to the correct parking stage.
[0030] (4) Improve the success rate of automatic parking closed loop and the accuracy of terminal alignment.
[0031] This application improves the planning stability of key stages in automatic parking tasks, such as low speed, large turning angle, reversing into a parking space, and terminal attitude alignment, by combining "visual world model prediction - hierarchical sub-objective constraint - retrieval of prior MPPI planning - closed-loop deviation state recovery enhancement". It can reduce terminal position error and heading error and improve the success rate of automatic parking closed-loop execution. Attached Figure Description
[0032] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments, and unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0033] Figure 1 This application provides a flowchart illustrating the specific architecture of a hierarchical retrieval-guided visual world model planning method and system for automatic parking, as part of an embodiment of this application. Detailed Implementation
[0034] As can be seen from the background technology, when existing visual world models are directly used for automatic parking, they suffer from problems such as large long-range prediction errors, low efficiency of random sampling planning, and insufficient recovery ability after deviating from the expert trajectory.
[0035] To address the aforementioned technical problems, this application provides a hierarchical retrieval-guided visual world model planning method for automatic parking. This method comprises three stages: Stage 1: Acquiring offline trajectory data for automatic parking, constructing an automatic parking action condition visual world model, and generating multiple stage sub-objectives based on the parking task progress to obtain a hierarchical parking sub-objective sequence; Stage 2: During online parking, based on the current vehicle state and the stage sub-objectives generated in Stage 1 corresponding to the current vehicle state, retrieving action segments matching the current vehicle state and the current stage sub-objectives from the offline parking trajectory retrieval library, using the retrieved action sequences as sampling priors for the Model Predicted Path Integral Control (MPPI), and using the action condition visual world model constructed in Stage 1 to predict the future state of candidate action sequences, outputting the optimal action sequence and executing it to obtain a closed-loop execution result; Stage 3: Detecting whether the vehicle is in a deviated state based on the closed-loop execution result of Stage 2. If it is in a deviated state, generating a recovery trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory, and enhancing the action condition visual world model and / or the offline parking trajectory retrieval library based on the generated trajectories. The visual world model of automatic parking action conditions constructed in this application not only learns the ideal expert parking process, but also learns the dynamic law of recovering from the deviation state to the stage sub-objective or the final parking space; at the same time, MPPI can retrieve more restorative action priors in subsequent planning, thereby improving the stability of automatic parking closed-loop control.
[0036] The embodiments of this application will now be described in detail with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0037] like Figure 1As shown in the embodiments of this application, a hierarchical retrieval-guided visual world model planning method and system for automatic parking is provided. The overall process includes three parts: offline modeling for automatic parking, online planning of MPPI (Multi-Performance Indicator) guided by retrieval, and closed-loop deviation state recovery enhancement. Specifically, the system first constructs an action-conditional visual world model based on offline trajectory data for automatic parking and generates a hierarchical parking sub-target sequence. Subsequently, during the online parking process, based on the current vehicle state and the current stage sub-target, it retrieves action segments matching the current vehicle state and the current stage sub-target from the offline parking trajectory retrieval library. The retrieved action sequences are used as MPPI sampling priors, and the action-conditional visual world model is used to predict the future state of candidate action sequences to output the optimal action sequence and drive the vehicle to execute. Finally, based on the closed-loop execution result, it detects whether the vehicle has deviated from its target state. If a deviation occurs, a recovery trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory are generated to form a recovery enhancement dataset, which is used to enhance the action-conditional visual world model. Recovery trajectory segments that meet the success or near-success conditions are added to the offline parking trajectory retrieval library to provide more reliable action priors in subsequent online planning processes.
[0038] The above three parts will be further explained below.
[0039] Phase 1: Acquire offline trajectory data for automatic parking, construct a visual world model of automatic parking action conditions, and generate multiple stage sub-objectives corresponding to the vehicle state according to the progress of the parking task, thus obtaining a hierarchical parking sub-objective sequence.
[0040] Acquire offline trajectory data for automatic parking, including vehicle visual observations, vehicle status, vehicle control actions, and target parking space information. Convert the parking scene into a unified visual observation representation and extract latent visual features using a pre-trained visual encoder.
[0041] A motion-conditional visual world model is constructed, which predicts future visual potential features based on historical visual potential features, vehicle state, and control actions, thereby establishing the state transition prediction capability in automatic parking scenarios.
[0042] Based on the parking task progress, expert trajectory keyframes, the position difference between the vehicle's current position and the target parking space position, and the heading difference between the vehicle's current heading and the target parking space's orientation, the complete parking task is divided into multiple stage sub-objectives.
[0043] The stage sub-objectives include one or more of the following: parking space entrance arrival sub-objective, reversing preparation sub-objective, reversing into the parking space sub-objective, vehicle straightening sub-objective, and terminal posture alignment sub-objective. Each stage sub-objective can be represented as a sub-objective visual image, a sub-objective vehicle pose, or a sub-objective visual latent feature.
[0044] Through the hierarchical sub-goal mechanism, vehicles only need to be planned to the current stage sub-goal in each planning cycle, rather than being directly planned to the final parking state.
[0045] Phase Two: During the online parking process, based on the current vehicle state and the phase sub-objectives generated in Phase One corresponding to the current vehicle state, action segments matching the current vehicle state and the current phase sub-objectives are retrieved from the offline parking trajectory retrieval library. The retrieved action sequences are used as the sampling priors for the model predicts path integral control (MPPI). The future states of candidate action sequences are predicted using the action-conditional visual world model constructed in Phase One. The optimal action sequence is output and executed to obtain the closed-loop execution result.
[0046] Specifically, Phase 2 mainly involves performing MPPI parking planning based on prior offline trajectory retrieval.
[0047] First, multiple action segments are segmented from the offline parking trajectory to establish a parking trajectory retrieval library. Each action segment includes the segment's starting state, ending state, parking stage, action sequence, and success level label. In Phase Two of this application, action segments matching the current vehicle state and the current stage sub-target are retrieved from the offline parking trajectory retrieval library. Specifically, a matching action segment refers to an action segment whose starting state is close to the current vehicle state, and whose corresponding parking stage is consistent with or connected to the current stage sub-target. Specifically, action segments in the parking trajectory retrieval library can be filtered based on the vehicle's current position, current heading, current parking stage, and stage sub-target position. A starting position close to the current vehicle position can mean the difference between their positions is less than a preset position threshold, and a starting heading close to the current vehicle heading can mean the difference between their headings is less than a preset heading threshold. If the starting position of an action segment is close to the current vehicle position, the starting heading is close to the current vehicle heading, and the ending state of the action segment can guide the vehicle towards the current stage sub-target, then the action segment is considered to match the current vehicle state and the current stage sub-target. For example, when the vehicle is in the reversing preparation stage, priority can be given to retrieving action segments whose starting state is close to the current vehicle state and whose action sequence can guide the vehicle into the reversing parking stage; when the vehicle is in the terminal attitude alignment stage, priority can be given to retrieving action segments whose endpoint position and endpoint heading are close to the terminal attitude alignment sub-target.
[0048] During online parking, based on the current vehicle state and the current stage sub-objective, action segments matching the current vehicle state and the current stage sub-objective are retrieved from the parking trajectory retrieval database. These retrieved action sequences are then used as sampling priors for the model's predicted path integral control (MPPI). Perturbations are introduced near this prior to generate multiple sets of candidate action sequences.
[0049] The system uses a visual world model to predict the future visual potential state corresponding to each candidate action sequence, and calculates the MPPI weight based on the cost between the predicted state and the current stage sub-objective to obtain the optimal action sequence. The vehicle uses a rolling time-domain control method to execute the first action or the first macro action in the optimal action sequence, and replans in the next control cycle.
[0050] Phase 3: Based on the closed-loop execution results of Phase 2, detect whether the vehicle is in a deviated state. If it is in a deviated state, generate a recovery trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory, and enhance the motion condition visual world model and / or the offline parking trajectory retrieval library based on the generated trajectory.
[0051] Phase 3 primarily focuses on data augmentation to recover from closed-loop deviations.
[0052] During the closed-loop execution of the vehicle, if the vehicle deviates from the expert trajectory, fails to complete the stage sub-objective, over-reverses, fails to make a large turn at low speed, or has an excessive terminal attitude error, the corresponding state is recorded as the starting point for data recovery.
[0053] Based on the starting point of the restored data, a restored trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory are generated and added to the restored augmentation data to enhance the visual world model. Simultaneously, restored trajectory segments that meet the success or near-success criteria are added to the parking trajectory retrieval library to provide more reliable action priors for subsequent MPPI planning.
[0054] This application utilizes a motion-conditional visual world model to predict the future parking state of a vehicle in a latent visual feature space; it divides the complete parking process into multiple stage sub-objectives to reduce the difficulty of long-term planning; during online planning, it retrieves motion segments from an offline parking trajectory library that match the current vehicle state and stage sub-objectives, and uses these segments as priors for MPPI sampling optimization; simultaneously, it generates recovery data based on deviation states during closed-loop execution to enhance the visual world model and trajectory retrieval library, thereby improving the stability of automatic parking closed-loop control.
[0055] The hierarchical retrieval-guided visual world model planning method and system for automatic parking provided in this application will be described in detail below through specific embodiments. The specific implementation of the invention will be illustrated using the example of a vehicle reversing into a parking lot.
[0056] In this embodiment, the vehicle starts from an initial position in the parking lane, with the goal of entering the target parking space and ultimately satisfying constraints on position error, heading error, vehicle speed, and parking hold. The vehicle can acquire current visual observations, vehicle pose, speed, steering, throttle, braking, and reversing status.
[0057] First, a visual world model of the automatic parking action conditions is constructed, and a hierarchical parking sub-target sequence is generated.
[0058] Set time The vehicle status is:
[0059] in, Indicates the vehicle's location. Indicates the vehicle's heading angle. Indicates vehicle speed. Indicates the vehicle's steering angle.
[0060] Assume the vehicle control actions are as follows:
[0061] in, Indicates steering control quantity. This indicates the longitudinal control quantity.
[0062] Convert the parking scene into a visual observation image:
[0063] in, This represents the visual representation function for the parking scene. Indicates the target parking space position. Indicates time Parking visual observation.
[0064] Using pre-trained visual encoders Extracting latent visual features:
[0065] in, Indicates time The visual latent features.
[0066] Constructing a visual world model with action conditions Based on historical visual latent features, historical vehicle states, and historical actions, predict the visual latent features for the next moment:
[0067] in, Indicates the current moment. Indicates the length of the history window. Represents a visual world model with action conditions. Represents the parameters of the visual world model. Indicates from time At that time Historical visual latent feature sequences Indicates from time At that time The historical sequence of vehicle control actions, Indicates from time At that time Historical vehicle state sequence This represents the predicted visual latent features for the next time step. During training, the latent features are used to predict the loss:
[0068] in, This represents the latent feature prediction loss. This represents the time step in the trajectory data. This represents the summation of prediction errors at each time step in the trajectory data. This represents the visual latent features predicted by the action-conditional visual world model for the next moment. This represents the real latent visual features extracted by the pre-trained visual encoder from real visual observations at the next time step. Let L2 be the squared norm. By minimizing the above loss, the visual world model learns the dynamic relationship between visual state and vehicle movement during automatic parking.
[0069] To address the long-term and phased nature of automated parking tasks, the complete parking task is divided into multiple phased sub-objectives. Let the sequence of complete parking sub-objectives be:
[0070] in, Indicates the first Each stage's sub-goals Indicates the number of sub-targets.
[0071] In this embodiment, the sub-target sequence includes:
[0072] in, This indicates that the parking space entrance leads to the sub-target. This indicates the sub-objective of the reversing preparation phase. This indicates the sub-objective of the reversing into a parking space phase. This indicates the sub-target of the vehicle body returning to upright position. This represents a sub-target in the terminal attitude alignment phase.
[0073] For the For each stage of sub-targets, generate corresponding visual observations of the sub-targets:
[0074] And extract the visual latent features of the sub-targets:
[0075] During online execution, the current stage index is determined based on the positional error, heading error, and stage completion status between the current vehicle status and the sub-objectives of each stage.
[0076] in, This represents a stage-determining function. Indicates time The corresponding current parking phase.
[0077] Within the current control cycle, the planner uses the current stage sub-objectives. Local planning is carried out as a target.
[0078] Next, MPPI parking planning is performed based on the prior knowledge of offline trajectory retrieval.
[0079] Segmenting motion segments from offline parking trajectories to build a parking trajectory retrieval library:
[0080] in, This indicates the number of trajectory segments. Each trajectory segment... Represented as:
[0081] in, Indicates the starting state of the segment. Indicates the end state of the segment. This indicates that the segment belongs to the parking phase. This indicates the action sequence corresponding to the segment. This indicates the success level label for the segment.
[0082] Action sequence Represented as:
[0083] in, Indicates the length of the planned horizon.
[0084] During online planning, based on the current vehicle status and current stage sub-goals Calculate the matching cost of the candidate segments:
[0085] in, Indicates the current position of the sub-target. Indicates the end position of the candidate segment. Indicates the heading angle of the sub-target at the current stage. Indicates the heading angle at the end of the candidate segment. These are the weighting coefficients. This is an indicator function.
[0086] Select the trajectory segment with the minimum matching cost:
[0087] And the corresponding action sequence is used as the action prior for MPPI:
[0088] During the MPPI sampling process, the first step is generated centered on the prior of the retrieval action. Group candidate action sequence:
[0089] in:
[0090] Indicates the first Group of candidate action sequences, This indicates a random perturbation. This represents the sampling covariance matrix.
[0091] For each candidate action sequence, a latent space prediction is performed using a visual world model:
[0092] in, , Indicates the first The candidate action sequence in the group is in the first The potential state of each prediction step Indicates the corresponding action. This indicates a predicted vehicle status.
[0093] Construct the planning cost function:
[0094] The first term is the cost of visual latent feature targets, the second term is the cost of position error, the third term is the cost of heading error, and the fourth term is the cost of motion amplitude. These are the corresponding weighting coefficients.
[0095] According to the MPPI weight update rule, calculate the first... Weights of candidate action sequences:
[0096] in, Indicates the number of candidate action sequences. Indicates the temperature coefficient. This represents the minimum cost among all candidate action sequences:
[0097] Based on the candidate action sequences and their weights, the optimal action sequence for the current time step is obtained:
[0098] Using a rolling time-domain control method, only the first action or the first macro action in the optimal action sequence is executed:
[0099] Vehicle execution action Afterwards, the system reacquires the vehicle status and visual observations, and enters the next round of closed-loop planning.
[0100] Finally, data augmentation is performed based on the closed-loop deviation state.
[0101] During closed-loop execution, if the vehicle deviates from the expert trajectory, fails to complete the stage sub-objective, overshoots the target, fails to make a large turn at low speed, or has an excessive terminal attitude error, the current state is recorded as the starting point for data recovery.
[0102] The recovery trajectory generated from the starting point of the recovered data is denoted as:
[0103] For situations where the terminal is close to the target parking space but its position or heading angle does not meet the requirements, a terminal attitude alignment trajectory dataset is generated. .
[0104] To address minor disturbances or motion noise that may occur during execution, a noisy trajectory correction dataset is generated. .
[0105] The restored augmented dataset is represented as:
[0106] in, This represents a dataset of expert parking trajectories.
[0107] Continue to optimize the visual world model using restored augmented data:
[0108] For recovered trajectory segments that meet the conditions for success or near success, add them to the offline parking trajectory retrieval library:
[0109] in, This represents the action segments obtained from the recovered trajectory. This indicates the success level of the recovered fragment.
[0110] In this way, the visual world model not only learns the ideal expert parking process, but also learns the dynamic laws of recovering from the deviation state to the stage sub-objective or the final parking space; at the same time, MPPI can retrieve more recovery action priors in subsequent planning, thereby improving the stability of automatic parking closed-loop control.
[0111] In addition, this application embodiment also provides a hierarchical retrieval-guided visual world model planning system for automatic parking. This system is used to implement the hierarchical retrieval-guided visual world model planning method for automatic parking described in the above embodiments. The system includes: an automatic parking offline modeling module, an MPPI online planning module, and a closed-loop deviation state recovery enhancement module connected in sequence. The automatic parking offline modeling module is used to construct a visual world model of automatic parking action conditions and generate a hierarchical parking sub-target sequence; the MPPI online planning module is used to execute MPPI parking planning based on prior information retrieved from the offline trajectory; and the closed-loop deviation state recovery enhancement module is used to perform recovery data enhancement based on the closed-loop deviation state.
[0112] Based on the above technical solutions, this application provides a hierarchical retrieval-guided visual world model planning method and system for automatic parking, comprising three stages: Stage 1: Acquiring offline trajectory data for automatic parking, constructing an automatic parking action condition visual world model, and generating multiple stage sub-objectives according to the parking task progress to obtain a hierarchical parking sub-objective sequence; Stage 2: During the online parking process, based on the current vehicle state and the stage sub-objectives generated in Stage 1 corresponding to the current vehicle state, retrieving action segments matching the current vehicle state and the current stage sub-objectives from the offline parking trajectory retrieval library, using the retrieved action sequences as sampling priors for the Model Predicted Path Integral Control (MPPI), and using the action condition visual world model constructed in Stage 1 to predict the future state of candidate action sequences, outputting the optimal action sequence and executing it to obtain a closed-loop execution result; Stage 3: Detecting whether the vehicle is in a deviated state based on the closed-loop execution result of Stage 2. If it is in a deviated state, generating a recovery trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory, and enhancing the action condition visual world model and / or the offline parking trajectory retrieval library based on the generated trajectories.
[0113] This application utilizes a motion-conditional visual world model to predict the future parking state of a vehicle in a latent visual feature space; it divides the complete parking process into multiple stage sub-objectives to reduce the difficulty of long-term planning; during online planning, it retrieves motion segments from an offline parking trajectory library that match the current vehicle state and stage sub-objectives, and uses these segments as priors for MPPI sampling optimization; simultaneously, it generates recovery data based on deviation states during closed-loop execution to enhance the visual world model and trajectory retrieval library, thereby improving the stability of automatic parking closed-loop control.
[0114] The hierarchical retrieval-guided visual world model planning method and system for automatic parking provided in this application have the following advantages compared with the prior art.
[0115] (1) Reduce the difficulty of long-range prediction for automatic parking.
[0116] This application divides the complete parking task into multiple stage sub-objectives, so that the visual world model only needs to plan to the current stage objective in each control cycle, reducing the long-range error accumulation caused by directly predicting from the initial state to the final parking space, and improving the planning stability of the reversing into the parking space and the terminal alignment stage.
[0117] (2) Improve the sampling efficiency of key parking actions.
[0118] This application introduces an offline parking trajectory retrieval prior into MPPI planning. It retrieves matching action segments based on the current vehicle state and stage sub-objectives, using the retrieved action sequence as the sampling distribution center. Compared to direct sampling from a random distribution, this method increases the probability of sampling key actions such as reversing, large turns, low-speed fine-tuning, and steering wheel straightening, thereby improving parking action search efficiency.
[0119] (3) Enhance the recovery capability under closed-loop deviation state.
[0120] This application addresses potential issues during vehicle closed-loop execution, such as deviation from the expert trajectory, over-reversing, failure to complete stage objectives, and excessive terminal posture errors. It generates recovered trajectories, terminal posture aligned trajectories, and noisy correction trajectories, which are then used for two-stage training of the visual world model and updating the retrieval database. As a result, the model can cover deviation state distributions beyond the expert trajectory, improving the vehicle's ability to recover from non-ideal states to the correct parking stage.
[0121] (4) Improve the success rate of automatic parking closed loop and the accuracy of terminal alignment.
[0122] This application improves the planning stability of key stages in automatic parking tasks, such as low speed, large turning angle, reversing into a parking space, and terminal attitude alignment, by combining "visual world model prediction - hierarchical sub-objective constraint - retrieval of prior MPPI planning - closed-loop deviation state recovery enhancement". It can reduce terminal position error and heading error and improve the success rate of automatic parking closed-loop execution.
[0123] Those skilled in the art will understand that the above-described embodiments are specific examples of implementing this application, and in practical applications, various changes in form and detail may be made without departing from the spirit and scope of this application. Any person skilled in the art can make their own modifications and alterations without departing from the spirit and scope of this application; therefore, the scope of protection of this application should be determined by the scope defined in the claims.
Claims
1. A hierarchical retrieval-guided visual world model planning method for automated parking, characterized in that, include: Phase 1: Acquire offline trajectory data for automatic parking, construct a visual world model of automatic parking action conditions, and generate multiple stage sub-objectives according to the progress of the parking task to obtain a hierarchical parking sub-objective sequence; Phase 2: During the online parking process, based on the current vehicle state and the phase sub-objectives generated in Phase 1 corresponding to the current vehicle state, action segments matching the current vehicle state and the current phase sub-objectives are retrieved from the offline parking trajectory retrieval library. The retrieved action sequences are used as the sampling priors for the model predicts path integral control MPPI. The action condition visual world model constructed in Phase 1 is used to predict the future state of the candidate action sequences. The optimal action sequence is output and executed to obtain the closed-loop execution result. Phase 3: Based on the closed-loop execution results of Phase 2, detect whether the vehicle is in a deviated state. If it is in a deviated state, generate a recovery trajectory, a terminal attitude alignment trajectory, and a noisy correction trajectory, and enhance the action condition visual world model and / or the offline parking trajectory retrieval library based on the generated trajectory.
2. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 1, characterized in that, Constructing a visual world model of the conditions for automatic parking actions, including: Acquire offline trajectory data for automatic parking; the offline trajectory data includes vehicle visual observation, vehicle status, vehicle control actions, and target parking space information; Based on the offline trajectory data, the parking scene is converted into a unified visual observation representation, and a pre-trained visual encoder is used to extract latent visual features. Based on the aforementioned visual latent features, a motion-conditional visual world model is constructed, which can predict future visual latent features based on historical visual latent features, vehicle status, and control actions. Based on the parking task progress, expert trajectory keyframes, the position difference between the vehicle's current position and the target parking space position, and the heading difference between the vehicle's current heading and the target parking space's orientation, the complete parking task is divided into multiple stage sub-objectives, generating a hierarchical parking sub-objective sequence for stage two online parking planning.
3. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 2, characterized in that, The stage sub-objectives include one or more of the following: parking space entrance arrival sub-objective, reversing preparation sub-objective, reversing into the parking space sub-objective, vehicle straightening sub-objective, and terminal attitude alignment sub-objective; The stage sub-target is used to represent the visual image of the sub-target, the vehicle pose of the sub-target, or the visual latent features of the sub-target.
4. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 2, characterized in that, Constructing a visual world model with action conditions Based on historical visual latent features, historical vehicle states, and historical actions, predict the visual latent features for the next moment: in, Indicates the current moment. Indicates the length of the history window. Represents a visual world model with action conditions. Represents the parameters of the visual world model. Indicates from time At that time Historical visual latent feature sequences Indicates from time At that time The historical sequence of vehicle control actions, Indicates from time At that time Historical vehicle state sequence This represents the predicted visual latent features for the next moment.
5. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 2, characterized in that, During training, latent features are used to predict the loss: in, This represents the latent feature prediction loss. This represents the time step in the trajectory data. This represents the summation of prediction errors at each time step in the trajectory data. This represents the visual latent features predicted by the action-conditional visual world model for the next moment. This represents the real latent visual features extracted by the pre-trained visual encoder from real visual observations at the next time step. Let L2 be the squared norm; by minimizing the above loss, the visual world model learns the dynamic relationship between visual state and vehicle movement during automatic parking.
6. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 1, characterized in that, MPPI parking planning is performed based on offline trajectory retrieval priors, including: Multiple action segments are obtained from offline parking trajectories, and a parking trajectory retrieval library is established; During online parking, based on the current vehicle status and the current stage sub-objective, action segments matching the current vehicle status and the current stage sub-objective are retrieved from the parking trajectory retrieval library, and the retrieved action sequences are used as the sampling prior for the model's predicted path integral control (MPPI). Perturbations are added near this prior to generate multiple sets of candidate action sequences. The motion-conditional visual world model constructed in Phase 1 is used to predict the future visual potential state corresponding to each candidate motion sequence. The MPPI weight is calculated based on the cost between the predicted state and the current stage sub-objective generated in Phase 1 to obtain the optimal motion sequence. The vehicle uses a rolling time-domain control method to execute the first action or the first macro action in the optimal motion sequence and replans in the next control cycle.
7. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 6, characterized in that, Each action segment includes the segment's starting state, the segment's ending state, the parking phase it belongs to, the action sequence, and a success level label.
8. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 1, characterized in that, Data augmentation based on closed-loop deviation states includes: During the closed-loop execution of the vehicle, if the vehicle deviates from the expert trajectory, fails to complete the stage sub-objective, over-reverses, fails to make a large low-speed turn, or has an excessive terminal attitude error, the corresponding state is recorded as the starting point for data recovery. Based on the starting point of the recovery data, a recovery trajectory, an end pose alignment trajectory, and a noisy correction trajectory are generated and added to the recovery enhancement dataset to enhance the visual world model.
9. The hierarchical retrieval-guided visual world model planning method for automatic parking according to claim 8, characterized in that, When enhancing the visual world model, recovered trajectory segments that meet the success or near-success conditions are added to the parking trajectory retrieval library to update the trajectory retrieval library and provide more reliable action priors for subsequent MPPI planning.
10. A hierarchical retrieval-guided visual world model planning system for automated parking, the system being used to implement the hierarchical retrieval-guided visual world model planning method for automated parking as described in any one of claims 1 to 9, characterized in that, The system includes: an offline modeling module for automatic parking, an online planning module for MPPI (Multi-Level Processing), and a closed-loop deviation recovery enhancement module, connected in sequence; among them, The automatic parking offline modeling module is used to construct a visual world model of automatic parking action conditions and generate multiple stage sub-objectives according to the parking task process to obtain a hierarchical parking sub-objective sequence. The MPPI online planning module is used to perform MPPI parking planning based on prior information retrieved from the offline trajectory. The closed-loop deviation state recovery enhancement module is used to enhance the recovery data based on the closed-loop deviation state.
Citation Information
Patent Citations
Method and device for helping parking
CN101830226A
Automatic parking planning method based on iterative sampling optimization
CN118618349A