VLA-based low-speed park automatic driving method
By collecting various data to generate a panoramic park environment perception tensor, semantic element recognition and dynamic trend prediction are performed to optimize path planning. This solves the problems of single perception dimension and lag in decision-making in low-speed park autonomous driving, achieving higher safety and comfort, and the ability to adapt to complex scene changes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DECK SMART TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-05
AI Technical Summary
Existing low-speed autonomous driving solutions in parks lack an end-to-end fusion architecture of vision, language, and action. Visual perception is easily affected by lighting and occlusion, and voice interaction is limited to the command reception level, unable to be deeply linked with environmental perception and action status. It lacks the ability to predict the movement trend of dynamic subjects, and decision-making relies heavily on real-time perception results, resulting in a delayed response. This leads to weak adaptability and robustness of the system in complex scenarios, making it difficult to meet safety and intelligence requirements.
The system collects road image data, obstacle laser point cloud data, environmental audio data, park voice interaction data, and vehicle movement status data of the park environment to generate a panoramic park environment perception tensor. Through semantic element recognition and dynamic subject movement trend prediction, it performs early obstacle avoidance path planning and smooth path optimization, generates a multimodal-path planning collaborative optimization parameter tensor, and continuously adjusts the autonomous driving execution logic.
It achieves comprehensive integration and correlation of multimodal data, enriches the environmental perception dimension, improves the response capability to dynamic scenarios, ensures the safety and comfort of vehicles in low-speed parks, adapts to changes in complex scenarios, and improves the stability and robustness of the system.
Smart Images

Figure CN121979044A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and specifically to a low-speed campus autonomous driving method based on VLA. Background Technology
[0002] Ensuring driving safety while balancing traffic efficiency and passenger comfort places stringent demands on the system's environmental understanding, decision-making response, and adaptive adjustment capabilities. This is especially true in densely populated areas, where the unpredictability of pedestrian behavior and the sudden nature of voice interaction needs necessitate multi-dimensional perception and proactive decision-making abilities from the system.
[0003] Existing low-speed autonomous driving solutions in industrial parks primarily rely on visual sensors to collect road image information and semantic segmentation algorithms to identify roads and obstacles. They are supplemented by simple voice command receiving modules that can only respond to control commands with fixed phrases. Vehicle movement status data is only used for real-time control feedback and does not participate in the optimization of perception and decision-making logic. This solution first acquires environmental images and identifies key elements using cameras, then generates a driving path according to preset rules, and finally adjusts control commands based on the vehicle's current speed, attitude, and other status data.
[0004] However, this solution has obvious technical shortcomings: it does not build an end-to-end fusion architecture of vision, language, and action; visual perception is easily affected by lighting and occlusion; voice interaction is only at the command reception level and cannot be deeply linked with environmental perception and action status; it lacks the ability to predict the movement trend of dynamic subjects, and decision-making relies heavily on real-time perception results, resulting in a delayed response; and it lacks a closed-loop optimization mechanism, so vehicle action status data does not have a reverse effect on the perception and decision-making logic adjustment, resulting in weak adaptability and robustness of the system in complex scenarios, making it difficult to meet the safety and intelligence requirements of low-speed park autonomous driving. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a low-speed campus autonomous driving method based on VLA, which can at least alleviate the aforementioned technical problems.
[0006] The technical solutions provided in this application are as follows: A low-speed autonomous driving method for a park based on VLA includes: Step 1, collecting road image data, obstacle laser point cloud data, environmental audio data, park voice interaction data, and vehicle movement status data of the park environment to generate a panoramic park environment perception tensor; Step 2, performing semantic element recognition and dynamic subject motion trend prediction on the panoramic park environment perception tensor to generate a park environment semantic parsing topology with dynamic motion trends; Step 3, performing early obstacle avoidance path planning and smooth path optimization on the park environment semantic parsing topology with dynamic motion trends and the preset driving path to generate a park autonomous driving control instruction set with dynamic safety redundancy; Step 4, executing the park autonomous driving control instruction set with dynamic safety redundancy and collecting vehicle driving status feedback data to generate a multimodal-path planning co-optimization parameter tensor, and continuously adjusting the autonomous driving execution logic based on the multimodal-path planning co-optimization parameter tensor.
[0007] This application has the following technical advantages: The technical solution of this application achieves comprehensive integration and correlation of multimodal data by acquiring road images, obstacle laser point clouds, environmental audio, park voice interaction, and vehicle movement status data, and generating a panoramic park environmental perception tensor. Traditional solutions are mostly visual or visual + laser-based, neglecting the intent information in voice interaction, environmental cues in audio, and the linkage value between vehicle movement status and environmental perception, resulting in many blind spots in complex scenarios. This solution, however, utilizes the complementarity of different modalities through multi-data acquisition and tensor fusion. For example, voice interaction data can capture personnel collaboration intentions, environmental audio data can assist in locating abnormal risk sources, and vehicle movement status data can calibrate the relative relationships of environmental perception. This makes the environmental perception dimensions richer and more correlated, better meeting the specific needs of low-speed parks with mixed pedestrian and vehicle traffic and frequent interactions, effectively solving the problem of insufficient scenario adaptability caused by the single perception dimension of traditional solutions. Furthermore, this application's solution performs semantic element recognition and dynamic subject movement trend prediction on the panoramic park environmental perception tensor to generate a park environmental semantic parsing topology with dynamic movement trends, and then performs advance obstacle avoidance path planning and stability optimization based on this topology. Traditional solutions rely solely on real-time obstacle avoidance decisions based on perceived obstacle locations. When faced with dynamic scenarios such as pedestrians suddenly crossing or vehicles making unexpected turns, the response time is insufficient, leading to sudden braking or abrupt avoidance maneuvers. This solution, however, uses dynamic trend prediction to identify the potential trajectories of dynamic subjects in advance. Combined with pre-planned obstacle avoidance, this allows the vehicle more time to adjust its driving state, resulting in smoother avoidance maneuvers. Furthermore, the optimized stability further adapts to the comfort requirements of low-speed park driving, precisely addressing the safety and comfort imbalance caused by the decision lag in traditional solutions. Moreover, by collecting vehicle driving status feedback data, a multimodal-path planning collaborative optimization parameter tensor is generated, continuously adjusting the autonomous driving execution logic. Traditional solutions have fixed parameters for perception, decision-making, and execution modules, making each module relatively independent and unable to adapt to dynamic changes in the park environment (such as differences in pedestrian density at different times, the appearance of temporary construction areas, and changes in lighting conditions) based on actual driving performance. The closed-loop adjustment mechanism of this solution can optimize the perception fusion logic and decision planning parameters in reverse based on feedback data such as trajectory deviation, obstacle avoidance effect and prediction accuracy during actual driving. This enables the system to continuously adapt to scene changes during long-term operation and improve operational stability and robustness. Attached Figure Description
[0008] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0009] Figure 1 This is a flowchart of a low-speed campus autonomous driving method based on VLA according to an embodiment of the present invention. Detailed Implementation
[0010] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. Various aspects are provided by way of explanation and not limitation of the invention. Indeed, those skilled in the art will recognize that modifications and variations can be made to the invention without departing from its scope or spirit. For example, a feature represented or described as part of one embodiment may be used in another embodiment to produce yet another embodiment. Therefore, it is desirable that the invention encompass such modifications and variations falling within the scope of the appended claims and their equivalents.
[0011] In the description of this invention, the terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," and "bottom," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and do not require the invention to be constructed and operated in a specific orientation; therefore, they should not be construed as limitations on the invention. The terms "connected," "linked," and "set up" used in this invention should be interpreted broadly. For example, they can refer to a fixed connection or a detachable connection; they can refer to a direct connection or an indirect connection through intermediate components. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0012] like Figure 1As shown, this application provides a low-speed autonomous driving method for a park based on VLA, including: Step 1, collecting road image data, obstacle laser point cloud data, environmental audio data, park voice interaction data, and vehicle movement status data of the park environment to generate a panoramic park environment perception tensor; Step 2, performing semantic element recognition and dynamic subject motion trend prediction on the panoramic park environment perception tensor to generate a park environment semantic parsing topology with dynamic motion trends; Step 3, performing early obstacle avoidance path planning and smoothness path optimization on the park environment semantic parsing topology with dynamic motion trends and the preset driving path to generate a park autonomous driving control instruction set with dynamic safety redundancy; Step 4, executing the park autonomous driving control instruction set with dynamic safety redundancy and collecting vehicle driving status feedback data to generate a multimodal-path planning co-optimization parameter tensor, and continuously adjusting the autonomous driving execution logic based on the multimodal-path planning co-optimization parameter tensor. Optionally, step 1 includes: Step 11, aligning the road image data, obstacle laser point cloud data, environmental audio data, park voice interaction data, and vehicle action status data collected by the VLA multimodal heterogeneous sensing fusion unit with spatiotemporal dimension benchmarks to generate a multimodal initial calibration data tensor; Step 12, performing multi-sensor accuracy weight assignment and scene adaptability calibration on the multimodal initial calibration data tensor to generate multimodal weight adaptation data; Step 13, performing multi-source conflict detection and weighted resolution on the multimodal weight adaptation data to fuse and generate a panoramic park environment perception tensor. The panoramic park environment perception tensor links and binds the road marking features, obstacle 3D features, abnormal sound spatial features, road marking visual features, voice interaction semantic features, and vehicle action status features in tensor dimensions, and the feature dimensions correspond one-to-one with the park's autonomous driving perception requirements.
[0013] Optionally, step 11 includes: Step 111, extracting frame time stamps from road image data, acquisition time stamps from laser point cloud data, sampling time stamps from environmental audio data, recording time stamps from park voice interaction data, and acquisition time stamps from vehicle movement status data, constructing a unified spatiotemporal reference axis and simultaneously generating a unified spatiotemporal reference system for the park; Step 112, based on the unified spatiotemporal reference system for the park, mapping the spatial coordinates of each modality data to the park's global coordinate system to complete spatial dimension alignment, obtaining spatiotemporally aligned multimodal data; Step 113, based on the spatiotemporally aligned multimodal data, generating a multimodal initial calibration data tensor, wherein the temporal and spatial dimensions of the multimodal initial calibration data tensor are consistent with the park's global reference, and provide a unified data basis for weight assignment.
[0014] Preferably, the specific implementation process of step 111 is as follows: The original temporal information of the five types of modal data is extracted and standardized to generate a unified temporal label set. Specifically, for road image data, the hardware trigger timestamp of the camera output frame is extracted. This timestamp is generated by the image sensor and the synchronous trigger module, with an accuracy of microseconds (e.g., ±1μs). Simultaneously, the frame number and exposure time parameters are recorded to form an image temporal record. For obstacle laser point cloud data, the starting acquisition timestamp of each rotation of the lidar is extracted. Combined with the angular resolution of the point cloud (e.g., 0.1° / point), a corresponding temporal sub-label is assigned to each set of point cloud data to ensure a precise correlation between the point cloud and the time dimension. For environmental audio data, due to its sampling rate being much higher than other modalities (e.g., 48kHz), a sliding window-based temporal aggregation processing method is adopted. Continuous audio sampling points are divided into audio frames according to a fixed duration (e.g., 20ms), and the start and end timestamps of each frame are extracted. Simultaneously, the average sound pressure level within the frame is calculated as an auxiliary feature for temporal correlation. For park voice interaction data, the activation and termination timestamps of the voice signal are extracted. These timestamps are generated by the voice wake-up module and are synchronously correlated with the signal-to-noise ratio (SNR) parameter of the voice signal. For vehicle movement status data, the acquisition timestamps of parameters such as vehicle speed and steering angle are extracted. These timestamps are synchronized with the communication cycle of the vehicle's CAN bus (e.g., 10ms / time). The temporal records of the above five modalities are standardized according to the timestamp format (e.g., YYYY-MM-DD-HH-MM-SS-μs) to generate a unified temporal tag set. Each tag contains a modal type identifier, core timestamp information, and an auxiliary parameter index. Preferably, in a further specific implementation of step 111, a global time series reference axis is constructed based on a unified time series label set, thereby generating a unified spatiotemporal reference system for the park; specifically, using the standard time provided by a preset GNSS (Global Navigation Satellite System) reference station within the park as a reference, all timestamps in the unified time series label set are calibrated for deviation, and the deviation of the calibrated timestamps is controlled within a preset range (e.g., ±5μs); based on the calibrated timestamps, a global time series reference axis is constructed, which divides time slices at fixed time intervals (e.g., 10ms), and each time slice corresponds to a unique time series index value; Simultaneously, a global coordinate system for the park is established, with a fixed landmark within the park (such as the center of the park entrance) as the origin. The x-axis extends along the main road of the park, the y-axis is perpendicular to the x-axis and points to the right side of the road, and the z-axis is perpendicular to the ground and points upward. The coordinate unit is meters. This coordinate system is calibrated by fusing GNSS positioning data and IMU (Inertial Measurement Unit) data to ensure the stability of the spatial coordinates. The global temporal reference axis is associated with the global coordinate system of the park to form a unified spatiotemporal reference system for the park. This system includes slicing index rules for the temporal dimension and coordinate transformation standards for the spatial dimension, providing a unified basis for the spatiotemporal alignment of subsequent multimodal data.
[0015] Preferably, in the specific technical implementation of step 112, based on the spatial coordinate transformation standard of the unified spatiotemporal reference system of the park, coordinate mapping processing is performed on the original spatial data of each modality to generate single-modality spatial mapping data; specifically, for road image data, the designed image pixel-spatial coordinate mapping model is used for transformation. This model includes camera intrinsic and extrinsic parameter matrices. The intrinsic parameter matrix is obtained through checkerboard calibration and includes parameters such as focal length (e.g., 1200 pixels) and principal point coordinates (e.g., (640, 360) pixels). The extrinsic parameter matrix is determined through the position calibration of the camera and the park's global coordinate system and includes rotation and translation vectors. During the transformation process, the image pixel coordinates are mapped to the three-dimensional coordinates under the park's global coordinate system through a perspective projection algorithm, while correcting the errors caused by lens distortion; for obstacle laser point cloud data, its original data is the three-dimensional coordinates under the local coordinate system of the lidar, and is processed through the calibration parameters (rotation matrix and translation vector) of the lidar and the park's global coordinate system. Linear coordinate transformation preserves the reflection intensity information and spatial distribution characteristics of the point cloud, ensuring the consistency of point cloud coordinates for the same obstacle in the global coordinate system. For environmental audio data, based on the spatial layout parameters of the microphone array (e.g., a circular array of 6 microphones with an array radius of 5cm), combined with the Time Difference of Arrival (TDOA) algorithm and the phase difference algorithm, the spatial azimuth and distance of the audio signal are calculated and mapped to the spatial region in the global coordinate system of the park, forming audio spatial positioning data. For park voice interaction data, based on the fixed installation coordinates of the voice acquisition device in the global coordinate system of the park (e.g., (x0, y0, z0)), the voice signal is associated with the coordinate point, and the propagation direction vector of the voice signal is recorded. For vehicle movement status data, the vehicle's own positioning coordinates (obtained by fusion of GNSS and IMU), attitude angles (roll angle, pitch angle, yaw angle), and other parameters are directly mapped to the global coordinate system of the park, forming vehicle state space data.Preferably, in a further specific implementation of step 112, based on the temporal slice indexing rules of the unified spatiotemporal reference system of the park, the single-modal spatial mapping data is subjected to temporal correlation processing to obtain spatiotemporally aligned multimodal data; specifically, the single-modal spatial mapping data of each modality is associated with the corresponding time slice of the global temporal reference axis according to its corresponding calibrated timestamp, and the spatial mapping data of all modalities at that moment is integrated within each time slice; for modal data with a collection frequency lower than the time slice interval (e.g., lidar collection frequency of 10Hz, time slice interval of 10ms), linear interpolation is used to supplement adjacent times. Missing data between slices ensures data continuity; for modal data with a sampling frequency higher than the time slice interval (e.g., audio sampling frequency of 48kHz, time slice interval of 10ms), multiple sets of data within the same time slice are aggregated for features, such as audio data being aggregated into features like maximum intra-frame sound pressure level and frequency peak, ensuring that the data volume is compatible with other modalities; after temporal correlation and data integration, spatiotemporally aligned multimodal data is obtained, in which the spatial feature information of all modalities is contained within the same time slice, and the spatial coordinates uniformly follow the global coordinate system of the park, and the temporal information uniformly follows the global temporal reference axis.
[0016] Preferably, in a scenario, step 113 is specifically implemented by constructing tensor dimensions and performing data filling processing based on the spatiotemporally aligned multimodal data to generate a multimodal initial calibration data tensor. Specifically, the dimensional structure of the multimodal initial calibration data tensor is determined, which includes four core dimensions: temporal dimension, modal dimension, feature dimension, and spatial dimension. The temporal dimension corresponds to the time slice index of the global temporal reference axis. The modal dimension includes five sub-dimensions: road image, laser point cloud, environmental audio, park voice interaction, and vehicle movement status. The feature dimension is set separately for different modalities. For example, the feature dimension of the road image includes pixel grayscale value, edge features, texture features, etc., and the feature dimension of the laser point cloud includes three-dimensional coordinates, reflection intensity, point cloud density, etc. The spatial dimension corresponds to the x, y, and z coordinate axes of the park's global coordinate system. The aforementioned dimensional structure fills the modal feature values in the spatiotemporally aligned multimodal data according to their corresponding dimensional positions. During the filling process, a confidence marker is added to each feature value. This marker is determined based on the sensor's acquisition accuracy and environmental adaptability (for example, the confidence marker value of a visual sensor is higher in sunny conditions than in rainy conditions). For the few missing values that remain after filling, a completion algorithm based on modal correlation is used to supplement them. For example, the spatial correlation between laser point cloud data and road image data is used to complete the feature values of occluded areas in the image. The temporal and spatial dimensions of the generated multimodal initial calibration data tensor strictly follow the unified spatiotemporal benchmark system of the park. The modal data are distributed in an orderly manner according to the dimensions in the tensor, providing a unified data basis for weight assignment in subsequent steps, ensuring that feature data of different modalities can be weighted and fused within the same tensor framework.
[0017] Optionally, step 12 includes: Step 121, based on the characteristics of the park scene, assigning high weights to laser sensors and vision sensors for near-range obstacle recognition and road environment semantic recognition; assigning high weights to vision sensors and voice sensors for long-range semantic recognition and interactive intent parsing; assigning auxiliary weights to audio sensors and motion state sensors for abnormal sound localization and auxiliary vehicle posture perception, thus forming a park scene sensor weight configuration rule; Step 122, based on the park scene sensor weight configuration rule, assigning weights to each modal feature in the multimodal initial calibration data tensor, thus generating modal weight annotation data; Step 123, based on the modal weight annotation data, generating multimodal weight adaptation data, wherein the weight allocation of the multimodal weight adaptation data is bound to the sensor adaptation characteristics of different areas of the park, and provides a weight basis for conflict resolution.
[0018] Optionally, step 13 includes: step 131, performing feature conflict detection on the multimodal weighted adaptation data, identifying conflict items such as visual-laser semantic contradictions, audio-visual spatial misalignment, visual-speech semantic contradictions, and action-visual state misalignment, and forming a list of multimodal feature conflict items; step 132, performing weighted resolution on the list of multimodal feature conflict items based on preset weights, retaining the feature results of high-weight modalities, and obtaining a single-modal feature array after weight resolution; step 133, fusing the single-modal feature array after weight resolution to generate a panoramic park environment perception tensor, the feature credibility of which is obtained by weighted calculation of each modality weight, and the three types of environmental features are based on a unified dimension.
[0019] Preferably, the specific implementation process of step 121 is as follows: The park scene is decomposed into multi-dimensional characteristics and processed by perception requirement mapping to generate a scene-perception-sensor association matrix; specifically, the park scene is divided into sub-scenes such as teaching area, office area, living area, road intersection, and green belt perimeter according to functional areas. Key characteristic parameters are extracted from each sub-scene, including personnel density (e.g., high, medium, and low levels, corresponding to numerical ranges of >3 people / square meter, 1-3 people / square meter, and <1 person / square meter, respectively), obstacle type (static obstacles such as streetlights, dynamic obstacles such as pedestrians / vehicles), lighting conditions (strong light, weak light, normal light), and noise level (high, medium, and low); for each... Sub-scenes are defined to clarify core perception requirements. For example, the teaching area needs to focus on pedestrian recognition and voice interaction intent parsing, while road intersections need to focus on long-distance obstacle recognition and turning intent judgment. Based on sensor performance characteristics (laser sensors have high accuracy in short-range ranging, visual sensors have strong long-range semantic recognition capabilities, audio sensors are sensitive to abnormal sound location, voice sensors accurately parsing interaction intent, and motion status sensors directly perceive vehicle posture), a scene-perception-sensor association matrix is constructed. The rows of this matrix represent sub-scenes in the park, the columns represent sensor types, and the matrix elements represent the degree of adaptation of a certain sensor to the corresponding perception requirements in a certain sub-scene (the value ranges from 0 to 1, with higher values indicating stronger adaptation).Preferably, in a further specific implementation of step 121, based on the scene-perception-sensor association matrix, a multi-dimensional weight allocation rule is formulated to form a park scene sensor weight configuration rule. Specifically, for the near-range obstacle recognition task (distance threshold < 5 meters), referring to the compatibility between the laser sensor and the visual sensor in the association matrix, the laser sensor is given a higher weight (e.g., 0.6-0.7), the visual sensor a medium weight (e.g., 0.2-0.3), and the other sensors are given auxiliary weights (e.g., 0.05-0.1). This is because the laser sensor has a smaller ranging error (e.g., ±3cm) in the near-range, while the semantic recognition of the visual sensor is easily affected by occlusion. For the road environment semantic recognition task, the visual sensor is given a higher weight (e.g., 0.5-0.6), and the laser sensor an auxiliary weight (e.g., 0.3-0.4), because the visual sensor can capture richer semantic information such as road markings and traffic signs. For the long-range semantic recognition task (distance threshold > 15 meters), the visual sensor is given a higher weight (e.g., 0.6-0.7), and the laser sensor a lower weight (e.g., 0.2-0.3). Due to the significant advantage of high resolution of visual sensors at long distances, the reduced density of laser point clouds leads to a decrease in recognition capability. For interactive intent parsing tasks, a core weight (e.g., 0.7-0.8) is assigned to the voice sensor, and an auxiliary weight (e.g., 0.1-0.2) to the visual sensor. This is because the voice sensor can directly capture human interaction commands, while the visual sensor can assist in verification through lip reading, gestures, and other information. For abnormal sound localization tasks, a core weight (e.g., 0.7-0.8) is assigned to the audio sensor, and an auxiliary weight (e.g., 0.05-0.1) to other sensors. This is because the audio sensor can achieve sound localization through sound pressure changes and azimuth angle calculations. For assisting vehicle posture perception tasks, a core weight (e.g., 0.8-0.9) is assigned to the motion state sensor, and a supplementary weight (e.g., 0.02-0.05) to other sensors. This is because the motion state sensor can directly collect posture parameters such as vehicle speed and steering angle. The above weight allocation rules are structured and integrated according to sub-scenes, perception tasks, and sensor types to form the park scene sensor weight configuration rules. These rules include three core parameters: a base weight value, a scene adjustment coefficient, and a task correction coefficient.
[0020] Preferably, in the specific technical implementation of step 122, based on the park scene sensing weight configuration rules and the multimodal initial calibration data tensor, hierarchical weight labeling processing is performed on each modal feature to generate a modal feature weight label set; specifically, firstly, the modal dimension and feature dimension of the multimodal initial calibration data tensor are analyzed to clarify the specific features contained in each modality. For example, the features of road image data include pixel grayscale features, edge features, texture features, semantic segmentation features, etc.; the features of laser point cloud data include three-dimensional coordinate features, reflection intensity features, point cloud clustering features, etc.; the features of audio data include sound pressure level features, frequency spectrum features, azimuth features, etc.; the features of voice interaction data include spectral features, keyword features, semantic vector features, etc.; and the features of vehicle action status data include vehicle speed. Features include steering angle features, acceleration features, etc. For each feature of each modality, based on its corresponding perception task type, the basic weight value in the park scene sensing weight configuration rules is queried. Then, combined with the scene adjustment coefficient of the current park sub-scene (e.g., the adjustment coefficient of the office area is 1.0, and the adjustment coefficient of the densely populated teaching area is 1.2) and the feature importance task correction coefficient (the correction coefficient of the core feature is 1.1, and the correction coefficient of the auxiliary feature is 0.9), the final weight value of the feature is obtained through weighted calculation (final weight value = basic weight value × scene adjustment coefficient × task correction coefficient). A weight label is added to each feature. The label contains information such as modality identifier, feature name, weight value, scene adjustment coefficient, and task correction coefficient. All labels are integrated to form a modality feature weight label set. Preferably, in a further specific implementation of step 122, consistency verification and anomaly correction are performed on the modal feature weight label set to generate standardized modal weight annotation data. Specifically, weight consistency verification rules are established, including that the sum of the weight values of all features under the same modality is 1, the feature weight value corresponding to the core perception task is not lower than a preset threshold (e.g., 0.5), and the auxiliary feature weight value is not higher than a preset threshold (e.g., 0.2). Each label in the modal feature weight label set is verified according to the rules. If there is a case where the sum of feature weights under the same modality is not 1, normalization processing is used to adjust the weight values proportionally to ensure that the sum is 1. If the core feature weight value is lower than the preset threshold, the calculation process of the scene adjustment coefficient and the task correction coefficient is traced back, the parameters are corrected, and the weight value is recalculated. The verified and corrected modal feature weight label set is stored in a structured manner, indexed and sorted by modal dimension and time dimension to generate standardized modal weight annotation data. This data corresponds one-to-one with the dimensional structure of the multimodal initial calibration data tensor, which facilitates subsequent data fusion processing.
[0021] Preferably, in a scenario, when step 123 is specifically implemented, based on the standardized modal weight annotation data and the sub-scene division results of the park, weight regionalization binding and dynamic adaptation processing are performed to generate multimodal weight adaptation data; specifically, through the park global coordinate system, the spatial boundary coordinates of each sub-scene are associated with the standardized modal weight annotation data to form a region-weight mapping table, which records the weight values of each modal feature within each spatial coordinate range; a dynamic weight adaptation engine is designed, which includes a scene recognition module and a weight adjustment module. The scene recognition module determines the current sub-scene location of the vehicle in real time through environmental features (such as personnel density features and illumination features) in the multimodal initial calibration data tensor. In the scenario, the weight adjustment module calls the corresponding weight value according to the region-weight mapping table and fine-tunes the weight value in combination with real-time environmental parameters (such as sudden abnormal sounds or sudden changes in light intensity). For example, when the audio sensor detects an abnormal sound, the weight adjustment coefficient of the audio sensor is temporarily increased (e.g., from 1.0 to 1.5). The weight data after region binding is integrated with the fine-tuned weight parameters and filled according to the temporal dimension and feature dimension of the multimodal initial calibration data tensor to generate multimodal weight adaptation data. Each feature dimension of this data carries the corresponding regionalized weight information, and the weight value is dynamically adjusted with the scene changes, providing accurate weight basis for conflict resolution in the subsequent step 13. Preferably, in a further specific implementation of step 123, the multimodal weight adaptation data undergoes validity verification and storage optimization to ensure data availability and retrieval efficiency. Specifically, a weight validity verification index system is constructed, including the matching degree between weights and scenarios (calculated by comparing the differences between actual scenario features and the scenario features corresponding to the weights, with a threshold set to ≥0.8), the real-time performance of weight adjustment (adjustment delay ≤10 milliseconds), and the rationality of weight values (core feature weight ≥0.5, auxiliary feature weight ≤0.2), etc. The generated multimodal weight adaptation data is verified frame by frame according to the above indicators, invalid data is removed and regenerated and replaced. A hierarchical storage structure is adopted to store the valid data, with the time-series dimension indexed by time slices, the modal dimension classified by sensor type, and the feature dimension sorted by weight value. At the same time, a fast retrieval index is established to ensure that the retrieval efficiency of weight data during subsequent conflict resolution is ≤5 milliseconds / time, ultimately forming multimodal weight adaptation data that can be directly used for conflict resolution.
[0022] Optionally, step 2 includes: step 21, performing semantic element segmentation on the panoramic park environment perception tensor, identifying four core semantic elements: road areas, pedestrians, park operation vehicles, and fixed obstacles, and generating basic semantic segmentation results for the park; step 22, based on the park's subject behavior database, performing dynamic subject movement trend prediction on pedestrians and park operation vehicles, and generating dynamic subject movement trend prediction results; step 23, fusing the basic semantic segmentation results and the dynamic subject movement trend prediction results to generate a park environment semantic parsing topology with dynamic movement trends. This topology associates static semantic regions with dynamic subject movement trajectories through topological edges and provides constraints for path planning.
[0023] Optionally, step 21 includes: Step 211, using a dedicated semantic segmentation model for the park scene to extract features from the road pixels, obstacle point clouds, audio anomaly regions, road pixels, speech semantic regions, and vehicle movement state regions in the panoramic park environment perception tensor, to obtain multimodal perception feature extraction results; Step 212, classifying the multimodal perception feature extraction results to distinguish between road areas, pedestrians, park operation vehicles, and fixed obstacles, and generating semantic category labeling data; Step 213, generating basic semantic segmentation results for the park based on the semantic category labeling data, wherein the semantic boundary of the basic semantic segmentation results for the park has an error threshold ≤ a preset accuracy threshold with respect to the actual physical scene of the park, and provides a static foundation for semantic topology construction.
[0024] Preferably, the specific implementation process of step 211 is as follows: A layered architecture for a dedicated semantic segmentation model for park scenes is constructed. Modality-specific feature extraction processing is performed on the panoramic park environment perception tensor to generate a single-modality core feature set. Specifically, the dedicated semantic segmentation model for park scenes includes a modality feature extraction layer, a cross-modality fusion layer, and a semantic output layer. The modality feature extraction layer designs a dedicated sub-network for different modal data. For road pixel data, a multi-scale convolutional sub-network is used for feature extraction. This sub-network contains three convolutional blocks, each consisting of a convolutional layer, a batch normalization layer, and an activation function. The convolutional kernel sizes are 3×3, 5×5, and 3×3, with a stride of 1. By gradually increasing the receptive field, fine-grained semantic features such as road markings and lane boundaries, as well as large-scale road area features, are captured to generate a road visual feature set. For obstacle laser point cloud data, a point cloud feature extraction sub-network improved with PointNet++ is used, employing a sampling layer, a grouping layer, and... The feature propagation layer transforms sparse point clouds into dense feature maps, focusing on extracting the 3D contour features, distance features, and reflection intensity features of obstacles to generate an obstacle point cloud feature set. For audio anomaly region data, a structure combining a Mel-frequency cepstral coefficient (MFCC) extraction subnetwork and a convolutional neural network is used. The audio signal is first converted into a Mel spectrogram, and then the frequency features and spatial orientation features of the abnormal sound are extracted through a 2D convolutional layer to generate an audio anomaly feature set. For speech semantic region data, a pre-trained speech feature extraction subnetwork is used to convert the speech signal into a fixed-dimensional semantic vector, focusing on capturing semantic information such as interactive commands and shouts to generate a speech semantic feature set. For vehicle movement state region data, a temporal convolutional network is used to extract the changing trend features of parameters such as vehicle speed and steering angle to generate a vehicle state feature set. The above modal feature sets together constitute a single-modal core feature set, and each feature set contains a feature vector, feature dimension identifier, and credibility score.Preferably, in a further specific implementation of step 211, based on the single-modal core feature set, feature association and complementary integration processing are performed through a cross-modal fusion layer to generate multimodal perception feature extraction results. Specifically, the cross-modal fusion layer adopts a combination of attention mechanism and feature concatenation. First, the weight coefficient of each modal core feature set is calculated. This coefficient is determined based on the feature's credibility score and scene adaptability. For example, when identifying road areas, the weight coefficient of the road visual feature set is increased (e.g., 0.6-0.7), and when identifying abnormal risks, the weight coefficient of the audio abnormal feature set is increased (e.g., 0). (5-0.6) The expression of high-weight features is enhanced through an attention mechanism to suppress the interference of redundant features. Then, the attention-weighted features of each modality are concatenated according to the channel dimension. The concatenated features are input into a fusion convolutional layer, which contains two convolutional blocks, each with a kernel size of 3×3 and a stride of 1, to integrate the correlation of cross-modal features and generate a multimodal fusion feature map. The multimodal fusion feature map is processed by global average pooling and fully connected processing to output a fixed-dimensional multimodal perception feature vector. This vector contains complementary information of all modalities and constitutes the multimodal perception feature extraction result.
[0025] Preferably, in the specific implementation of step 212, the multimodal perception feature extraction results are subjected to category probability calculation and threshold screening to generate preliminary category labeling data. Specifically, the multimodal perception feature extraction results are input into a classification head, which consists of two fully connected layers and a Softmax activation function. The number of neurons in the fully connected layers are 512 and 4 respectively (corresponding to four core semantic elements). The Softmax function converts the feature vectors into probability values for each category (with a value range of 0-1 and a sum of 1). A category determination threshold is set (e.g., 0.5-0.6). When the probability value of a certain category is higher than the threshold, the region corresponding to the feature is determined to belong to that category. For features with probability values all lower than the threshold, they are marked as regions with undetermined categories. Based on the above determination results, a category label (road area, pedestrian, park operation vehicle, fixed obstacle, undetermined) is assigned to each feature vector in the multimodal perception feature extraction results. At the same time, the probability values of each category and the spatial coordinates corresponding to the features are recorded to generate preliminary category labeling data. Preferably, in a further specific implementation of step 212, the preliminary category labeling data is processed for undetermined region completion and category consistency verification to generate semantic category labeling data. Specifically, for undetermined category regions, a region growing algorithm is used for completion. Using pixels or point clouds of known categories around the region as seed points, the region is expanded and grown based on feature similarity (e.g., Euclidean distance threshold of 0.3-0.5) to classify the undetermined region into the most similar category. Category consistency verification is performed, and spatial neighborhood verification rules are constructed. That is, the category labels of adjacent spatial regions should be consistent. If there is a category conflict between adjacent regions (e.g., a road region and a fixed obstacle region are directly adjacent without transition), the multimodal perception feature probability value of the conflicting region is recalculated, and the label is corrected in combination with the category distribution of the surrounding region. At the same time, referring to the prior knowledge of the park scene (e.g., road regions should be continuously distributed, and fixed obstacle regions should not cover a large area of road regions), category labels that obviously do not conform to the scene logic are corrected. After completion and verification, semantic category labeling data is obtained, which includes the category label, probability value, spatial coordinate range, and verification pass mark of each spatial region.
[0026] Preferably, in a scenario, when step 213 is specifically implemented, semantic boundary optimization and region normalization are performed based on semantic category labeling data to generate basic semantic segmentation results for the park. Specifically, morphological processing and boundary smoothing algorithms are used to optimize the semantic boundaries. First, dilation is used to fill the small holes inside the category regions, and then erosion is used to shrink the boundaries to remove isolated noise points. The kernel size for both dilation and erosion is 3×3. A polynomial curve fitting algorithm is used to smooth the semantic boundaries, reducing jagged fluctuations and ensuring the continuity and smoothness of the semantic boundaries. Each category region is normalized, and areas smaller than a preset threshold (e.g., 1 square meter, corresponding to a certain number of pixels or point clouds) are deleted. The number of small regions (determined by sensor resolution) is merged, and adjacent, scattered regions of the same category are merged to ensure the integrity of the category regions. Based on the optimized category regions and semantic boundaries, the basic semantic segmentation results of the park are generated. The results are presented in the form of a semantic segmentation map and a region attribute table. Each pixel or point cloud in the semantic segmentation map corresponds to a unique category label, and the region attribute table records the spatial coordinate range, area, center position, and other information of each category region. The basic semantic segmentation results of the park are checked for errors. The distance error between the semantic boundary and the actual physical scene of the park is calculated to ensure that the error threshold is ≤ the preset accuracy threshold (e.g., 0.1-0.2 meters) to meet the static basic requirements for semantic topology construction. Preferably, in a further specific implementation of step 213, the basic semantic segmentation results of the park are validated and standardized in format to ensure data usability. Specifically, a validity validation index system is constructed, including category recognition accuracy (calculated by comparison with manually labeled samples, with a threshold of ≥0.9), semantic boundary smoothness (calculated by the standard deviation of boundary curvature, with a threshold of ≤0.3), and regional integrity (calculated by the category region coverage, with a threshold of ≥0.95). The generated basic semantic segmentation results of the park are validated according to the above indicators, invalid data is removed, and the segmentation process is re-executed. The validated results are converted into a standardized format, with the semantic segmentation map stored in PNG format and the regional attribute table stored in JSON format. At the same time, metadata such as data identifiers, generation time, and sensor parameters are added to ensure that the results can be directly used for the dynamic subject movement trend prediction in step 22 and the semantic topology construction in step 23.
[0027] Optionally, step 22 includes: step 221, retrieving the park's main behavior database, which pre-stores benchmark behavior data such as typical trajectories of pedestrians crossing roads and typical turning operation patterns of vehicles in the park related to high-frequency dynamic conflicts, forming a benchmark behavior dataset for the park's main entities; step 222, inputting the real-time position and movement speed of the dynamic entity into the behavior matching model, performing feature matching and probability calculation based on the benchmark behavior dataset for the park's main entities, and obtaining a dynamic entity behavior trend index matrix; step 223, generating a dynamic entity movement trend prediction result based on the dynamic entity behavior trend index matrix and the predicted trajectory of the dynamic entity derived from the benchmark behavior dataset for the park's main entities, the prediction time domain of this result is adapted to the response time requirements of low-speed autonomous driving, and provides dynamic constraints for semantic topology.
[0028] Preferably, the specific implementation process of step 221 is as follows: Classify and organize the dynamic conflict scenarios in the park and collect and process benchmark behavioral data to generate an original behavioral dataset; specifically, divide the dynamic conflict scenarios in the park into sub-scenarios such as pedestrians crossing roads, vehicles turning and merging, pedestrians and vehicles traveling in parallel, and multiple pedestrians gathering and crossing, according to the conflict type. For each sub-scenarios, select typical areas (such as road intersections, canteen entrances, and teaching building entrances) for data collection; the collection equipment includes high-definition cameras, LiDAR, and audio sensors, and the collection time covers different time periods in the morning, noon, and evening, with a cumulative collection time of no less than 1000 hours; the collected benchmark behavioral data includes pedestrian position coordinate sequences, movement speed change curves, posture characteristics (such as whether they are looking down or carrying items), the starting and ending positions of crossing roads, the driving trajectory of park operation vehicles, changes in turning angles, vehicle speed curves, and dwell time in the operation area, etc.; denoise the collected raw data, remove abnormal data points caused by sensor errors (such as data with speed changes exceeding 5 m / s), and obtain the original behavioral dataset. Preferably, in a further specific implementation of step 221, feature extraction and structured annotation are performed on the original behavior dataset to form a benchmark behavior dataset of the park's main body, and a database of the park's main body behavior is constructed. Specifically, for pedestrian behavior data, key features such as the curvature features of the movement trajectory, the mean and variance of speed, the rate of change of turning angle, and the time window for crossing the road are extracted and classified according to the probability of crossing (high, medium, low). For example, trajectories that complete crossing within 10 seconds and have a stable speed of 1-1.5 m / s are labeled as high-probability crossing behaviors. For park operation vehicle behavior data... For the data, features of turning intention (such as the percentage of speed reduction before turning and the duration of turn signal activation) and work mode features (such as whether the vehicle is driving in a fixed area and the distribution of dwell time) are extracted and categorized and labeled according to work type (such as material transportation and environmental cleaning) and turning intention (left turn, right turn, straight go). The labeled feature data is then stored in a structured manner according to scene type and subject type to build a database of park subject behavior. This database includes data index, subject type identifier, behavior feature vector, behavior category label and scene adaptation parameters, and supports fast retrieval by scene and subject type.
[0029] Preferably, in the specific technical implementation of step 222, feature extraction and standardization processing are performed on the real-time motion data of the dynamic subject to generate a real-time behavior feature vector. Specifically, the real-time position coordinates of the dynamic subject (pedestrians, park operation vehicles) are extracted from the basic semantic segmentation results of the park (based on the park's global coordinate system), and the real-time motion speed (including magnitude and direction), acceleration, and steering angle are calculated through continuous multi-frame data. For pedestrians, real-time posture features are additionally extracted (determined by the positional relationship of key limb points captured by visual sensors, such as head orientation and torso tilt angle). For park operation vehicles, real-time vehicle light status (such as whether the turn signals are on) and operation equipment status (such as whether the sweeping device is started) are additionally extracted. These real-time data are standardized according to the feature dimensions of the baseline behavior data to unify the data range and dimensions, and generate a fixed-length real-time behavior feature vector, which includes sub-vectors such as position, speed, acceleration, and posture / equipment status. Preferably, in a further specific implementation of step 222, the real-time behavior feature vector is input into the behavior matching model, and similarity calculation and probabilistic inference are performed with the park's main body benchmark behavior dataset to obtain a dynamic subject behavior trend index matrix. Specifically, the behavior matching model adopts a structure combining an improved K-nearest neighbor algorithm and a Bayesian network. First, it retrieves the K benchmark behavior samples (K values are 10-20) with the highest similarity to the real-time behavior feature vector from the park's main body benchmark behavior dataset. The similarity is calculated using Euclidean distance, with smaller distances indicating higher similarity. Based on the retrieved K samples, various target behaviors (such as actions) are statistically analyzed. The frequency of occurrences (people crossing the road, vehicles turning left) is used to calculate the initial probability. This initial probability is then input into a Bayesian network. The input nodes of this network are real-time behavioral features (such as speed and posture), and the output nodes are the target behavior categories. The network parameters are obtained through training on a benchmark behavioral dataset and are used to correct the initial probability, taking into account the correlation between features (e.g., the probability of a pedestrian crossing the road is higher when their head is facing the road). The corrected output probabilities constitute a dynamic subject behavior trend index matrix. The rows of this matrix represent dynamic subject individuals, the columns represent target behavior categories, and the matrix elements represent the probability value (range 0-1) of a subject performing a certain target behavior.
[0030] Preferably, in a scenario, step 223 is specifically implemented by: based on the dynamic subject behavior trend index matrix and combined with the trajectory features of the park subject baseline behavior dataset, deducing the future motion trajectory of the dynamic subject and generating a dynamic subject prediction trajectory set; specifically, selecting target behaviors with probability values higher than a preset threshold (e.g., 0.5) in the behavior trend index matrix as candidate behaviors; for each candidate behavior, selecting corresponding baseline trajectory samples from the park subject baseline behavior dataset and extracting the time-position mapping relationship of the trajectory; based on the real-time position, velocity, and acceleration of the dynamic subject, using a quadratic polynomial fitting algorithm, adapting the baseline trajectory samples to the current motion state, and deducing the position coordinate sequence within a preset time domain (adapting to the response time of low-speed autonomous driving, e.g., 3-10 seconds); generating multiple prediction trajectories for each candidate behavior, considering uncertainties such as speed fluctuations and direction deviations, with each trajectory carrying a probability weight (consistent with the probability value in the behavior trend index matrix); the prediction trajectories of all candidate behaviors together constitute the dynamic subject prediction trajectory set, which includes trajectory coordinate sequences, timestamps, behavior category labels, and probability weights. Preferably, in a further specific implementation of step 223, the dynamic subject prediction trajectory set is subjected to effectiveness screening and fusion processing to generate dynamic subject motion trend prediction results. Specifically, trajectory effectiveness evaluation indicators are constructed, including the adaptability of the trajectory to the park scene (such as whether it meets road boundary constraints), the smoothness of the trajectory (evaluated by the rate of curvature change), and the level of probability weights. Predicted trajectories with poor adaptability, insufficient smoothness, or probability weights lower than a minor threshold (such as 0.3) are eliminated. The remaining effective trajectories are weighted and fused according to probability weights to generate the final prediction trajectory (each timestamp corresponds to a weighted average position coordinate). At the same time, key parameters of the fused trajectory are extracted, including the predicted behavior category, the probability of occurrence, the expected time to reach the key location (such as the road centerline), and the spatial range covered by the trajectory. This information is integrated with the final prediction trajectory to generate dynamic subject motion trend prediction results. The prediction time domain of this result matches the response time of low-speed autonomous driving and can provide clear dynamic constraints for subsequent semantic topology construction, such as prohibiting vehicles from traveling along the original path within the area covered by the predicted trajectory.
[0031] Optionally, step 3 includes: step 31, performing autonomous driving feasible area and dynamic constraint boundary identification on the semantic parsing topology of the park environment with dynamic motion trends to generate an initial driving path; step 32, performing advance obstacle avoidance planning on the initial driving path, pre-setting deceleration and detour trajectories for high-probability dynamic conflict scenarios to obtain an intermediate path with obstacle avoidance strategy; step 33, performing low-speed stability optimization on the intermediate path with obstacle avoidance strategy and adding dynamic safety redundancy, converting the optimized path into steering, acceleration, and deceleration control commands to generate a park autonomous driving control command set with dynamic safety redundancy. This command set will link control parameters including steering angle, acceleration, and braking intensity with the dynamic subject motion trend prediction results and directly drive the vehicle to perform driving actions. Optionally, step 31 includes: step 311, extracting the boundary coordinates of the feasible area of the park environment with dynamic motion trend and marking the dynamic subject motion trajectory constraints to obtain the semantic topology constraint extraction result; step 312, generating an initial version of the conflict-free driving path based on the semantic topology constraint extraction result and the preset driving path; step 313, verifying the feasibility of the initial version of the conflict-free driving path, eliminating the path segments that conflict with the dynamic constraints, and generating the final initial driving path. The final initial driving path provides a basic path for obstacle avoidance planning, and the path direction is consistent with the overall direction of the preset driving path.
[0032] Preferably, the specific implementation process of step 311 is as follows: Boundary extraction and coordinate quantization are performed on the static semantic regions in the semantic analysis topology of the park environment with dynamic motion trends to generate a set of boundary coordinates for feasible road areas. Specifically, the semantic analysis topology of the park environment with dynamic motion trends includes topological association information between static semantic regions (road areas, fixed obstacle areas, etc.) and dynamic subject movement trajectories. First, the semantic labels of road areas in the topology are located, and contour extraction algorithms (such as a combination of Canny edge detection and contour tracking) are used to capture the outer and inner contours of the road areas (such as lane dividers). The extracted contour pixel coordinates are converted into three-dimensional coordinates in the global coordinate system of the park. The contour coordinates are simplified using a uniform sampling algorithm, retaining the coordinates of key turning points (such as road corners and intersection forks). The sampling interval is determined according to the road width (for example, the sampling interval is 0.5 meters when the road width is ≥5 meters and 0.3 meters when it is <5 meters). The sampled coordinates are sorted in clockwise order to form a closed set of boundary coordinates for feasible road areas. This set of coordinates includes a list of outer boundary coordinates, a list of inner boundary coordinates, and boundary type identifiers (such as road edges and lane lines). Preferably, in a further specific implementation of step 311, based on the dynamic constraint information in the semantic parsing topology of the park environment with dynamic motion trends, the motion trajectory of the dynamic subject is subjected to spatiotemporal dimension labeling processing to obtain a dynamic trajectory constraint label set, which is combined with the boundary coordinate set of the feasible road area to form a semantic topology constraint extraction result; specifically, the motion trend prediction results of each dynamic subject (pedestrian, park operation vehicle) are extracted from the semantic parsing topology, including the predicted trajectory coordinate sequence, timestamp, behavior category and probability weight; spatiotemporal constraint labels are added to each predicted trajectory, and the spatial constraint label is the spatial area covered by the trajectory (centered on the trajectory coordinates, according to the size of the dynamic subject (pedestrian)). The rectangular area is expanded to 0.5m × 0.3m (and 2m × 1.5m for operating vehicles). The time constraint is marked as the time window in which the dynamic subject is located in the spatial area (e.g., the 2nd to 5th second). The spatiotemporal constraint marks are classified according to the conflict risk level (high, medium, low). The risk level is determined by the behavior probability weight and the distance from the vehicle (probability ≥ 0.6 and distance < 10 meters is high risk, probability 0.3-0.6 and distance 10-20 meters is medium risk, and the rest are low risk). The dynamic trajectory constraint mark set is integrated with the road feasible area boundary coordinate set to form the semantic topological constraint extraction result, which contains two major categories of information: static boundary constraints and dynamic spatiotemporal constraints.
[0033] Preferably, in the specific technical implementation of step 312, the preset driving path is subjected to coordinate discretization and path segmentation processing to generate a set of discretized preset path segments. Specifically, the preset driving path is a sequence of global path points determined by GPS+RTK positioning. The path point sequence is interpolated and supplemented at fixed distance intervals (e.g., 1 meter) to obtain continuous discretized path points. The discretized path points are divided into several path segments according to the driving direction. Each path segment includes the coordinates of the starting point, the coordinates of the ending point, the path length (e.g., 5-10 meters), and the heading angle (the angle with the x-axis of the global coordinate system of the park). Static constraint verification is performed on each path segment, and path segments that conflict with the boundary coordinate set of the road feasible area (i.e., the path segment partially or completely exceeds the road feasible area) are removed to generate a set of discretized preset path segments after preliminary screening. Preferably, in a further specific implementation of step 312, based on the dynamic trajectory constraint marker set in the semantic topological constraint extraction results, dynamic conflict prediction and path segment reorganization are performed on the initially screened discretized preset path segment set to generate an initial version of conflict-free driving path; specifically, a path-dynamic constraint spatiotemporal conflict judgment model is constructed, inputting the start / end time of the path segment (calculated from the preset driving speed, for example, at a preset speed of 10km / h, the driving time of a 5-meter-long path segment is 1.8 seconds) and the spatiotemporal constraint information of the dynamic trajectory constraint marker set, to determine whether the path segment has spatiotemporal overlap with high- and medium-risk dynamic trajectory constraints; if there is overlap... The path segment is then adjusted using a path offset algorithm, prioritizing the offset direction away from the dynamic subject. The offset amount is 1.2-1.5 times the size of the dynamic subject (e.g., 0.6-0.75 meters for pedestrians and 2.4-3 meters for work vehicles). The adjusted path segment is then smoothly connected to adjacent path segments, and a Bezier curve fitting algorithm is used to optimize the heading angle change rate at the path segment connection (ensuring ≤5° / meter). All conflict-free path segments are then spliced together in the driving sequence to form a complete initial conflict-free driving path. This path includes a discrete path point coordinate sequence, the estimated driving time for each path point, and the heading angle.
[0034] Preferably, in one scenario, when step 313 is specifically implemented, the feasibility of the initial conflict-free driving path is verified, including vehicle dynamics constraint verification and scenario adaptability verification, to generate a path feasibility verification result; specifically, the vehicle dynamics constraint verification verifies the curvature, slope, and acceleration of the initial conflict-free driving path, and the curvature must meet the minimum turning radius requirement of the vehicle (for example, the minimum turning radius of the park shuttle bus is ≥5 meters, corresponding to a path curvature ≤0.2m). -¹), the slope must be ≤ the maximum slope of the park road (e.g., ≤15°), and the acceleration must be within the vehicle's power output range (e.g., -2m / s² to 1.5m / s²); the scene adaptability verification combines the characteristics of the park's sub-scenes, for example, in the teaching area, the path must be far away from the edge of the sidewalk (distance ≥0.8 meters), and at road intersections, the path must avoid the turning blind spot (the blind spot range is calculated based on the vehicle size and turning angle); path segments that do not meet the constraints are marked as infeasible path segments, and the conflict type (dynamic conflict or scene adaptability conflict) and conflict location coordinates are recorded to form the path feasibility verification result. Preferably, in a further specific implementation of step 313, based on the path feasibility verification results, the initial conflict-free driving path is subjected to infeasible path segment replacement and overall path smoothing optimization to generate the final initial driving path; specifically, for the marked infeasible path segments, suitable alternative path segments are retrieved from the preset path candidate library (the candidate library contains commonly used detour path segments in different areas of the park, stored according to sub-scenario categories), and the alternative path segments must meet the requirements of having consistent start / end point coordinates with the original path segments, smooth transition of heading angles, and no static or dynamic conflicts; a path fusion algorithm is used to combine the alternative path segments with the original feasible path segments. The path is spliced together, and the coordinates of the spliced path points are adjusted using a global optimization algorithm (such as a genetic algorithm) to ensure the curvature continuity and consistency of the overall path with the driving direction. The optimized path is then subjected to a final conflict check to confirm that there are no static boundary conflicts, no high / medium risk dynamic spatiotemporal conflicts, and that the path meets the requirements of vehicle dynamics and scene adaptation. The generated final initial driving path contains a complete sequence of discrete path point coordinates, the estimated driving time of each path point, the heading angle, and the constraint satisfaction indicator. Its path direction deviates from the overall direction of the preset driving path by ≤10°, providing a stable basic path for the advance obstacle avoidance planning in the subsequent step 32.
[0035] Optionally, step 32 includes: step 321, performing risk conflict scenario screening on the dynamic subject motion trend prediction results to form a list of risk dynamic conflict scenarios; step 322, planning an advance deceleration section and detour trajectory for the list of risk dynamic conflict scenarios, reserving sufficient avoidance time domain, to obtain an advance obstacle avoidance trajectory scheme; step 323, smoothly connecting the advance obstacle avoidance trajectory scheme with the initial driving path to generate an intermediate path with obstacle avoidance strategy. The intermediate path with obstacle avoidance strategy takes into account both traffic efficiency and obstacle avoidance safety, and provides a basic trajectory for stability optimization.
[0036] Preferably, the specific implementation process of step 321 is as follows: A multi-dimensional risk conflict assessment index system is constructed to quantitatively assess the dynamic subject movement trend prediction results, thereby generating a risk assessment matrix. Specifically, the risk conflict assessment index system includes four core indicators: behavioral probability weight, spatiotemporal overlap, dynamic subject type, and distance threshold. The behavioral probability weight directly adopts the probability value (0-1) from the dynamic subject movement trend prediction results. The spatiotemporal overlap is the proportion of overlap between the predicted trajectory of the dynamic subject and the initial driving path in the spatiotemporal dimension (0-1). The dynamic subject type is assigned a value according to the risk level (pedestrians are assigned a value). 1.0, park operation vehicles are assigned a value of 0.8, and other dynamic obstacles are assigned a value of 0.5. The distance threshold is the normalized value of the initial distance between the vehicle and the dynamic subject (the closer the distance, the larger the value, ranging from 0 to 1). The four indicators are weighted and summed according to preset weights (behavioral probability weight 0.4, spatiotemporal overlap 0.3, dynamic subject type 0.2, and distance threshold 0.1) to obtain the comprehensive risk conflict score for each dynamic subject (ranging from 0 to 1). A risk assessment matrix is constructed based on the comprehensive risk conflict score. The rows of the matrix represent individual dynamic subjects, the columns represent assessment indicators and comprehensive scores, and the matrix elements are the quantitative values of the corresponding indicators. Preferably, in a further specific implementation of step 321, a risk screening threshold is set based on the risk assessment matrix, and dynamic subjects are classified for risk and conflict scenarios are extracted to form a list of dynamic risk conflict scenarios. Specifically, a high-risk screening threshold (e.g., 0.6-0.7) and a medium-risk screening threshold (e.g., 0.4-0.5) are set. Dynamic subjects with a comprehensive risk conflict score ≥ the high-risk threshold are classified as high-risk conflict sources, those with scores between the medium and high-risk thresholds are classified as medium-risk conflict sources, and the rest are low-risk (not included in the list for the time being). For each high- and medium-risk conflict source, the corresponding conflict scenario information is extracted, including dynamic subject type, behavior category, predicted trajectory coordinate sequence, spatiotemporal overlapping area, comprehensive risk score, and expected conflict occurrence time. Conflict scenarios are sorted from high to low according to the comprehensive risk score, and the scenario attributes (e.g., road intersections, surrounding teaching areas) and environmental parameters (lighting, noise) of the conflict scenarios are supplemented to form a structured list of dynamic risk conflict scenarios. This list supports rapid retrieval by conflict occurrence time and risk level.
[0037] Preferably, in the specific technical implementation of step 322, for high-risk conflict scenarios in the list of dynamic conflict scenarios, based on the initial driving path and the predicted trajectory of the dynamic subject, parameters for the early deceleration segment are planned, and a deceleration segment configuration set is generated; specifically, firstly, the estimated time (T1) for the vehicle to arrive at the spatiotemporal overlap area and the estimated time (T2) for the dynamic subject to enter the area are calculated, and the deceleration start timing is determined according to the time difference between the two (ΔT=T1-T2). If ΔT≥3 seconds, deceleration is started at a preset distance (e.g., 20-30 meters) from the overlap area; if ΔT<3 seconds, deceleration is started immediately; based on The vehicle's current speed (V0), preset safe passage speed (V1, V1≤8km / h for pedestrian conflicts, V1≤10km / h for vehicle conflicts), and deceleration start distance (S) are used to calculate the deceleration acceleration (a) using a uniform deceleration motion model. This ensures that the vehicle's speed drops to V1 when it reaches the overlapping area, and the absolute value of the acceleration during deceleration is ≤1.5m / s² (to meet the requirements for low-speed driving stability). The deceleration segment configuration set includes the coordinates of the deceleration start point, the coordinates of the deceleration end point, the initial speed, the target speed, the deceleration acceleration, and the deceleration duration. Each deceleration segment is bound to the corresponding high-risk conflict scenario. Preferably, in a further specific implementation of step 322, detour trajectory parameters are planned by combining the deceleration segment configuration set and the list of dynamic risk conflict scenarios, and integrated with the deceleration segment configuration set to obtain an advance obstacle avoidance trajectory scheme; specifically, the planning of the detour trajectory aims to avoid the spatiotemporal overlap area of the predicted trajectory of the dynamic subject, and the detour direction is preferentially selected to be away from the dynamic subject and not deviating from the overall direction of the initial driving path (e.g., if the dynamic subject is on the right side of the road, detour to the left; if it is on the left side, detour to the right); the offset of the detour trajectory is determined according to the size of the dynamic subject and the safety redundancy requirements (the offset is determined when there is a pedestrian conflict). For collisions involving vehicles operating within the park (with a displacement ≥ 0.8 meters, and ≥ 1.5 meters in case of conflict), the curvature of the detour trajectory is fitted using a cubic spline curve to ensure a curvature change rate ≤ 5° / meter, avoiding sharp turns. The detour trajectory includes the coordinate sequence of the starting point (coinciding with the deceleration section's end point), the passing points, and the ending point (connecting with the initial driving path), as well as the expected driving speed and heading angle for each point. The detour trajectory is integrated with the corresponding deceleration section configuration set to generate independent obstacle avoidance sub-schemes for each high- and medium-risk conflict scenario. All obstacle avoidance sub-schemes are arranged in driving order to form an advance obstacle avoidance trajectory scheme.
[0038] Preferably, in a scenario, when step 323 is specifically implemented, the coordinates of the connection point between the advance obstacle avoidance trajectory scheme and the initial driving path are determined, and the connection area is processed to perform path smoothing transition to generate a transition path segment. Specifically, the connection point is divided into a connection point before the deceleration phase starts (point A) and a connection point after the detour trajectory ends (point B). Point A is the deceleration start point on the initial driving path, and point B is the path point on the initial driving path closest to the detour trajectory end point (distance ≤ 5 meters). A transition path from 5 meters before point A to the deceleration phase start point is generated using a Bezier curve to ensure a continuous transition in speed and heading angle between the initial driving path and the deceleration phase. The control point of the curve is determined based on the tangent direction of the initial driving path and the starting direction of the deceleration phase. Similarly, a transition path from the detour trajectory end point to 5 meters after point B is generated. By adjusting the control point, the detour trajectory and the initial driving path are smoothly connected. The heading angle change rate of the connection area is ≤ 3° / meter, and the speed change rate is ≤ 0.5m / s². Preferably, in a further specific implementation of step 323, the transition path segment, the advance obstacle avoidance trajectory scheme, and the non-conflict segment of the initial driving path are integrated and spliced together to generate an intermediate path with an obstacle avoidance strategy; specifically, the non-conflict segment of the initial driving path (from the starting point to 5 meters before point A), the transition path segment at point A, the deceleration segment, the detour trajectory, the transition path segment at point B, and the non-conflict segment of the initial driving path (5 meters after point B to the end point) are spliced in sequence according to the driving order; after splicing, the overall path is globally smoothed and optimized, and the path point coordinates are corrected using a moving average filtering algorithm to eliminate local inflection points. To ensure the continuity of path curvature, the efficiency indicators (total travel time, average speed) and safety indicators (minimum distance to the dynamic subject, obstacle avoidance redundancy time) of the intermediate path with obstacle avoidance strategy are calculated. The indicators are verified to meet the preset requirements (average speed ≥ 6 km / h, minimum distance ≥ 0.8 meters, obstacle avoidance redundancy time ≥ 1.5 seconds). If the indicators do not meet the requirements, the deceleration section and detour trajectory parameters are readjusted until the requirements are met. The final generated intermediate path with obstacle avoidance strategy takes into account both traffic efficiency and obstacle avoidance safety, providing a basic trajectory for the stability optimization in the subsequent step 33.
[0039] Optionally, step 33 includes: step 331, performing curvature smoothing optimization on the intermediate path with obstacle avoidance strategy to ensure that the path curvature adapts to the low-speed driving stability requirements, and obtaining a smoothed obstacle avoidance path; step 332, adding dynamic safety redundancy to the control parameters corresponding to the smoothed obstacle avoidance path, including extending the braking response time and increasing the obstacle avoidance lateral distance, and generating a control parameter configuration table with safety redundancy; step 333, converting the control parameter configuration table with safety redundancy into steering, acceleration, and deceleration control commands, generating a park autonomous driving control command set with dynamic safety redundancy, the parameter range of which matches the low-speed park driving safety specifications.
[0040] Preferably, the specific implementation process of step 331 is as follows: Curvature calculation and abnormal inflection point detection are performed on the intermediate path with obstacle avoidance strategy to generate a path curvature distribution table and an inflection point coordinate set. Specifically, the intermediate path with obstacle avoidance strategy consists of a sequence of discrete coordinate points. A numerical differential algorithm is used to calculate the curvature of the arc formed by three adjacent coordinate points. The curvature calculation interval is consistent with the path point sampling interval (e.g., 0.3-0.5 meters). The calculated curvature values are arranged in the order of path points to form a path curvature distribution table, which includes path point coordinates, corresponding curvature values, and curvature change rate. A curvature threshold and a curvature change rate threshold are set, and path points with curvature exceeding the threshold or curvature change rate exceeding the threshold are selected as abnormal inflection points. Their coordinates and curvature anomaly types (single-point abrupt change, continuous fluctuation) are recorded to form an inflection point coordinate set. Preferably, in a further specific implementation of step 331, based on the path curvature distribution table and the inflection point coordinate set, an adaptive smoothing algorithm is used to correct the path in the abnormal inflection point region to obtain a smoothed obstacle avoidance path. Specifically, the designed adaptive smoothing algorithm includes two stages: local correction and global optimization. In the local correction stage, for single-point abrupt inflection points, a quadratic Bézier curve is used to replace the path segment formed by the inflection point and the two path points before and after it. The curve control points are determined based on the tangent directions of the paths on both sides of the inflection point to ensure a continuous transition of curvature in the corrected path. For continuously fluctuating inflection point regions, an adaptive smoothing algorithm is used to correct the path in the abnormal inflection point region. The path points within the region are refitted using the moving least squares method. The fitting window size is dynamically adjusted according to the fluctuation range (the larger the fluctuation range, the larger the window; for example, when the fluctuation range is 5 meters, the window size is 5 path points). After the local correction is completed, the global optimization stage begins. The overall curvature distribution of the corrected path is calculated, and the path point coordinates are fine-tuned using a global curvature equalization algorithm to ensure that the maximum curvature of the entire path does not exceed the threshold and the rate of change of curvature is lower than the threshold. The optimized path is the smoothed obstacle avoidance path. The curvature continuity of this path meets the requirements for low-speed driving stability, and the steering operation is smooth.
[0041] Preferably, in the specific technical implementation of step 332, based on the driving scenario of the smoothed obstacle avoidance path and the dynamic subject's movement trend, the basic values of the dynamic safety redundancy parameters are determined, and a set of safety redundancy basic parameters is generated. Specifically, the dynamic safety redundancy parameters include three categories: braking response time redundancy, obstacle avoidance lateral distance redundancy, and speed reserve redundancy. The basic value of braking response time redundancy is determined according to the risk level of the dynamic subject (1.5-2.0 seconds for high-risk conflict scenarios, 1.0-1.5 seconds for medium-risk scenarios, and 0.5-1.0 seconds for no-risk scenarios). The basic value of obstacle avoidance lateral distance redundancy is determined according to the type of dynamic subject (pedestrians ≥ 0.8 meters, park operation vehicles ≥ 1.5 meters, fixed obstacles ≥ 0.5 meters). The basic value of speed reserve redundancy is 10%-20% of the preset speed of the current road segment (for example, reserving 1-2 km / h when the preset speed is 10 km / h). The three types of parameters are classified according to the path segments and bound to each path segment of the smoothed obstacle avoidance path to form a set of safety redundancy basic parameters. This parameter set includes path segment identifiers, basic values of the three types of redundancy parameters, and descriptions of applicable scenarios. Preferably, in a further specific implementation of step 332, the safety redundancy basic parameter set is dynamically adjusted by combining real-time environmental parameters and vehicle status parameters to generate a control parameter configuration table with safety redundancy. Specifically, the real-time environmental parameters include light intensity, road friction coefficient (estimated through vehicle driving status data), and visibility; the vehicle status parameters include current vehicle speed and load weight. If the light intensity is below a threshold (e.g., <500 lux) or the road friction coefficient is below a threshold (e.g., <0.6), the braking response time redundancy is extended by 20%-30%; if the visibility is below a threshold (e.g., <50 meters), the braking response time redundancy is extended by 20%-30%. The lateral distance redundancy for obstacle avoidance is increased by 15%-25%; if the vehicle load weight exceeds 80% of the rated load, the speed reserve redundancy is increased to 30%; the adjusted redundancy parameters are associated with the control parameters (steering angle, acceleration, braking intensity) corresponding to the smoothed obstacle avoidance path. The control parameters are calculated based on the path curvature and driving speed (steering angle is positively correlated with curvature, and acceleration is determined based on path length and driving time); the control parameters and corresponding dynamic safety redundancy parameters are organized in the order of path segments to form a control parameter configuration table with safety redundancy. This table contains path segment information, control parameter values, redundancy parameter values, and parameter activation conditions.
[0042] Preferably, in one scenario, when step 333 is specifically implemented, the parameters in the control parameter configuration table with safety redundancy are processed by range conversion and instruction encoding to generate a raw set of control instructions. Specifically, the control parameters in the control parameter configuration table with safety redundancy are physical quantity values (such as steering angle ±30°, acceleration -2-1.5m / s², braking intensity 0-100%), which need to be converted into electrical signal ranges that the vehicle control system can recognize. The steering angle is converted according to the range ratio of the vehicle steering system (for example, ±30° corresponds to an electrical signal of 0-5V), and the acceleration and braking intensity are converted according to the control range of the throttle and brake actuator (for example, acceleration 1.5m / s² corresponds to throttle opening of 80%, and braking intensity 100% corresponds to braking pressure of 10MPa). The converted electrical signal values are encoded using an instruction encoding protocol. The encoded content includes parameter type identifier, numerical information, check code, and effective timestamp. Each path segment corresponds to a set of encoded instructions, forming a raw set of control instructions. Preferably, in a further specific implementation of step 333, the original set of control commands undergoes compliance verification and timing synchronization processing to generate a set of autonomous driving control commands for the park with dynamic safety redundancy. Specifically, the compliance verification is based on low-speed park driving safety regulations, checking whether the parameter range of the control commands is within the safe range (e.g., the steering angle does not exceed the vehicle's maximum steering angle, and the braking intensity does not exceed the system's maximum allowable value), eliminating non-compliant commands and supplementing them based on interpolation of adjacent commands. The timing synchronization processing aligns the control commands with the driving time axis of the smoothed obstacle avoidance path, calculates the execution time of each command based on the path segment length and the preset driving speed, and ensures that the command execution order is consistent with the vehicle's driving progress. After synchronization, a dynamic safety redundancy identifier and priority label are added to each command (obstacle avoidance related commands have higher priority than normal driving commands), and they are arranged in the execution order to form a set of autonomous driving control commands for the park with dynamic safety redundancy. This set of commands can directly drive the vehicle to perform steering, acceleration, and deceleration actions, and the parameter range complies with safety regulations.
[0043] Optionally, step 4 includes: Step 41, collecting vehicle driving trajectory data, obstacle avoidance effect evaluation data, sensor data matching degree data, multimodal data fusion matching degree data, and dynamic prediction accuracy data, and generating a standardized driving state feedback tensor after normalization processing; Step 42, inputting the standardized driving state feedback tensor into the multimodal fusion-path planning closed-loop collaborative optimization model, adjusting the sensor weights of the multimodal spatiotemporal calibration-weight conflict resolution fusion architecture and the sensor weights of the vision-speech-action multimodal spatiotemporal calibration-weight conflict resolution fusion architecture, and obtaining a sensor weight adjustment parameter configuration table; Step 43, correcting the dynamic trend prediction logic and optimizing the safety redundancy parameters based on the standardized driving state feedback tensor, fusing the sensor weight adjustment parameter configuration table with the above correction and optimization results, generating a multimodal-path planning collaborative optimization parameter tensor, and updating the autonomous driving execution logic based on this parameter tensor. Optionally, step 41 includes: step 411, performing deviation analysis on the driving trajectory data and calculating trajectory deviation quantification data, which includes the offset between the actual trajectory and the planned trajectory. The trajectory deviation quantification data is used to quantify the accuracy of path planning; step 412, performing quantitative evaluation of obstacle avoidance effect and statistically analyzing obstacle avoidance effect quantification data, which includes the number of successful avoidances and the number of risk encounters. The obstacle avoidance effect quantification data is used to quantify the effectiveness of the obstacle avoidance strategy; step 413, based on the trajectory deviation quantification data, obstacle avoidance effect quantification data, sensor data matching degree data, multimodal data fusion matching degree data, and dynamic prediction accuracy data, performing tensor dimension normalization processing to generate a standardized driving state feedback tensor. The evaluation dimension of this tensor corresponds one-to-one with the autonomous driving performance indicators.
[0044] Preferably, the specific implementation process of step 411 is as follows: The actual driving trajectory data and the planned trajectory data of the vehicle are spatiotemporally aligned to generate a trajectory alignment dataset. Specifically, the actual driving trajectory data is collected by a GNSS+IMU integrated navigation system, including continuous timestamps, three-dimensional coordinates in the global coordinate system of the park, driving speed, and heading angle, with a collection frequency of 50-100Hz. The planned trajectory data consists of a sequence of coordinate points and timestamps corresponding to a smoothed obstacle avoidance path. A timestamp matching algorithm is used to align the actual trajectory and the planned trajectory with the same timestamp. For coordinate points in the actual trajectory that lack corresponding timestamps, linear interpolation is used to supplement them. For redundant timestamps in the planned trajectory, they are filtered according to the actual trajectory timestamps, ultimately forming a trajectory alignment dataset. This dataset contains the actual coordinates and planned coordinates corresponding to each timestamp. Preferably, in a further specific implementation of step 411, trajectory deviation parameters are calculated based on the trajectory alignment dataset to generate trajectory deviation quantification data. Specifically, the trajectory deviation parameters include four categories: lateral deviation, longitudinal deviation, heading angle deviation, and cumulative deviation. Lateral deviation is the distance difference between the actual coordinates and the planned coordinates perpendicular to the driving direction (positive for the left and negative for the right). Longitudinal deviation is the distance difference along the driving direction (positive for leading and negative for lagging). Heading angle deviation is the difference between the actual heading angle and the planned heading angle. The three types of deviation values are calculated time-stamped using numerical calculation methods, and deviation thresholds are set (lateral deviation ≤ 0.3 meters, longitudinal deviation ≤ 0.5 meters, heading angle deviation ≤ 5°). The number of deviations exceeding the threshold, the duration, and the maximum deviation value are counted. The cumulative deviation is the root mean square value of all types of deviations throughout the entire journey, used to quantify the overall trajectory deviation degree. The above deviation parameters are classified and organized according to timestamps and journey segments to form trajectory deviation quantification data, which directly reflects the accuracy of path planning.
[0045] Preferably, in the specific technical implementation of step 412, obstacle avoidance scenario matching and effect judgment processing are performed based on the dynamic subject motion trend prediction results and actual driving records to generate an obstacle avoidance scenario matching table. Specifically, the obstacle avoidance action information (deceleration start time, deceleration magnitude, detour direction, detour distance) of the vehicle is extracted from the driving records and matched with the actual motion trajectory of the dynamic subject. The spatiotemporal range of the obstacle avoidance action is matched with the scenario information in the list of risk dynamic conflict scenarios to determine the predicted scenario corresponding to each obstacle avoidance action. The matching results (successful matching, unmatched) and the fit between the obstacle avoidance action and the predicted scenario (such as whether the detour direction is consistent with the predicted trajectory avoidance direction) are recorded to form an obstacle avoidance scenario matching table. This table includes obstacle avoidance event identifiers, matching scenario information, and action fit score. Preferably, in a further specific implementation of step 412, based on the obstacle avoidance scenario matching table and the safety distance judgment standard, quantitative indicators of obstacle avoidance effect are statistically analyzed to generate quantitative data of obstacle avoidance effect. Specifically, the statistical standard for the number of successful avoidances is: the minimum distance between the vehicle and the dynamic subject after the obstacle avoidance action is ≥ the obstacle avoidance lateral distance redundancy value, and there is no risk of collision; the statistical standard for the number of risk contacts is: the minimum distance is < 80% of the obstacle avoidance lateral distance redundancy value, or there is a situation where emergency braking (braking intensity ≥ 80%) avoids a collision; additionally, the smoothness index of the obstacle avoidance action (such as the acceleration change rate during deceleration ≤ 1.0 m / s³) and the traffic efficiency impact index (such as the travel delay time caused by obstacle avoidance) are statistically analyzed; the number of successful avoidances, the number of risk contacts, the smoothness index and the efficiency impact index are classified and summarized according to the obstacle avoidance event type (pedestrians, work vehicles, other obstacles) to form quantitative data of obstacle avoidance effect, which directly reflects the effectiveness of the obstacle avoidance strategy.
[0046] Preferably, the specific implementation process of step 413 is as follows: collect sensor data matching degree data, multimodal data fusion matching degree data, and dynamic prediction accuracy data, perform unified data dimension processing, and generate a multi-dimensional evaluation dataset; specifically, the sensor data matching degree data is the consistency score between the original data of each sensor and the fusion result (value 0-1, calculated based on feature similarity), the multimodal data fusion matching degree data is the consistency score between the fused features and the actual scene (value 0-1, verified by manually labeled samples), and the dynamic prediction accuracy data is the consistency between the dynamic subject movement trend prediction result and the actual behavior (such as the degree of consistency between the predicted value of pedestrian crossing probability and whether or not they actually cross, value 0-1); associate the three types of data with timestamps and travel segments, supplement the scene parameters (such as illumination and noise) at the time of data collection, unify the data format to a fixed-dimensional vector, and generate a multi-dimensional evaluation dataset. Preferably, in a further specific implementation of step 413, the trajectory deviation quantification data, obstacle avoidance effect quantification data, and multi-dimensional evaluation dataset are subjected to tensor dimension normalization processing to generate a standardized driving state feedback tensor. Specifically, the designed normalization process includes two stages: numerical standardization and dimension construction. In the numerical standardization stage, the min-max normalization algorithm is used to map the numerical range of various data to the 0-1 interval. For example, the maximum lateral deviation in trajectory deviation is mapped to 1, and no deviation is mapped to 0. In the dimension construction stage, the core dimensions of the tensor are determined as follows: time dimension (divided according to travel segments, such as every 500 meters as a time slice), evaluation index dimension (including five sub-dimensions: trajectory deviation, obstacle avoidance effect, sensor matching degree, fusion matching degree, and dynamic prediction accuracy), and scene attribute dimension (including sub-dimensions such as illumination, noise, and dynamic subject type). The standardized values are filled according to the above dimensions to form a standardized driving state feedback tensor. Each element of this tensor corresponds to the normalized value of a certain time slice and a certain evaluation index in a specific scenario. The evaluation dimensions correspond one-to-one with the autonomous driving performance indicators and can be directly used for subsequent optimization parameter calculations.
[0047] Optionally, step 42 includes: Step 421, adjusting the weight allocation of laser, vision, audio, speech, and motion state sensors based on the sensor data matching degree data and multimodal data fusion matching degree data in the standardized driving state feedback tensor, and adjusting the weight ratio of high matching degree modes to obtain sensor weight allocation adjustment results; Step 422, optimizing the temporal alignment accuracy of multimodal spatiotemporal calibration based on the dynamic prediction accuracy data in the standardized driving state feedback tensor to obtain spatiotemporal calibration accuracy optimization parameters; Step 423, generating a sensor weight adjustment parameter configuration table based on the sensor weight allocation adjustment results and spatiotemporal calibration accuracy optimization parameters. This parameter configuration table provides a direct basis for parameter updates of the multimodal spatiotemporal calibration-weight conflict resolution fusion architecture and the vision-speech-motion multimodal spatiotemporal calibration-weight conflict resolution fusion architecture.
[0048] Preferably, the specific implementation process of step 421 is as follows: Scene classification statistical processing is performed on the sensor data matching degree data and multimodal data fusion matching degree data in the standardized driving state feedback tensor to generate a modal matching degree scene distribution matrix. Specifically, the standardized driving state feedback tensor includes three core dimensions: time, evaluation index, and scene attributes. Scenes are classified according to the lighting conditions (strong light, weak light, normal), dynamic subject type (pedestrian, park operation vehicle, others), and distance range (near distance < 5 meters, medium distance 5-15 meters, far distance > 15 meters) in the scene attribute dimension. For each classified scene, the average sensor data matching degree of five types of sensors (laser, vision, audio, voice, and action status) and the average multimodal data fusion matching degree are statistically analyzed to form a modal matching degree scene distribution matrix. The rows of this matrix represent classified scenes, the columns represent sensor types and fusion results, and the matrix elements are the average matching degree values (0-1) for the corresponding scene. Higher values indicate stronger adaptability of the modality in that scene. Preferably, in a further specific implementation of step 421, based on the modal matching degree scene distribution matrix, a dynamic weight adjustment algorithm is used to adjust the weight ratio of various sensors to obtain the sensor weight allocation adjustment result. Specifically, the designed dynamic weight adjustment algorithm uses the average matching degree as the core basis and combines it with the scene priority coefficient (e.g., densely populated scenes have higher priority than open scenes) to calculate the weight adjustment coefficient of various sensors. For sensors with an average matching degree higher than a preset threshold (e.g., 0.8) in a certain scene, their weight ratio is increased (adjustment coefficient is 1.1-1.3); for sensors with an average matching degree lower than the threshold (e.g., 0.5), their weight ratio is decreased (adjustment coefficient is 0.7-0.9). After the weight adjustment, it must be ensured that the sum of the weights of all sensors is 1, and the weight ratio of core sensors (e.g., vision, laser) is not lower than a preset lower limit (e.g., 0.2). A sensor weight adjustment table is generated according to the classification scene. The table contains the scene type, the adjusted weight values of various sensors, the adjustment basis and the effective conditions. The weight adjustment tables of all scenes together constitute the sensor weight allocation adjustment result.
[0049] Preferably, in the specific technical implementation of step 422, error source analysis is performed on the dynamic prediction accuracy data in the standardized driving state feedback tensor to generate a time-series alignment error attribution report. Specifically, the dynamic prediction accuracy data is the degree of agreement between the predicted dynamic subject movement trend and the actual behavior under each time slice. An accuracy threshold (e.g., 0.7) is set, and time slices with an accuracy lower than the threshold are selected as low-accuracy segments. The spatiotemporal alignment of multimodal data corresponding to the low-accuracy segments is analyzed. By comparing the timestamp deviation and spatial coordinate mapping error of each modality, the type of time-series alignment error (time synchronization deviation, spatial registration deviation, and mixed deviation) is determined. The frequency of occurrence, scope of influence (number of dynamic subjects involved), and magnitude of error (time deviation in milliseconds and spatial deviation in centimeters) of different error types are statistically analyzed to form a time-series alignment error attribution report. This report clarifies the core causes and degree of influence of various time-series alignment errors. Preferably, in a further specific implementation of step 422, based on the timing alignment error attribution report, the key parameters of multimodal spatiotemporal calibration are optimized to obtain spatiotemporal calibration accuracy optimization parameters. Specifically, for time synchronization deviation, the compensation coefficient of sensor timestamp synchronization is adjusted. For example, when the average timestamp deviation of a certain type of sensor is 10 milliseconds, its compensation coefficient is adjusted to -10 milliseconds to ensure accurate alignment of timestamps of each modality. For spatial registration deviation, the sensor extrinsic parameter matrix (rotation matrix and translation vector) is corrected. For example, when the spatial registration deviation between the lidar and the vision sensor is 5 centimeters, the translation vector parameter in the extrinsic parameter matrix is adjusted to reduce spatial mapping error. The designed parameter optimization rules include parameter adjustment range limits (a single adjustment shall not exceed 20% of the original value) and verification mechanisms. The optimized parameters need to be verified through simulation testing to ensure that the timing alignment error is reduced to a preset range (time deviation ≤ 5 milliseconds, spatial deviation ≤ 3 centimeters). All optimized spatiotemporal calibration parameters are organized by type to form a spatiotemporal calibration accuracy optimization parameter set, which includes parameter name, optimized value, applicable scenario and error improvement effect.
[0050] Preferably, the specific implementation process of step 423 is as follows: A correlation analysis is performed on the sensor weight allocation adjustment results and the spatiotemporal calibration accuracy optimization parameters to generate a parameter correlation mapping table; specifically, the mutual influence between sensor weight adjustment and spatiotemporal calibration parameter optimization is analyzed. For example, after increasing the weight of the visual sensor, the spatiotemporal calibration parameters of the visual sensor and other modalities need to be optimized accordingly to ensure that the sensor data after weight increase can be more accurately fused with other modalities; the correlation between sensor weight and spatiotemporal calibration parameters is established according to the classification scenarios, clarifying the spatiotemporal calibration parameter optimization requirements corresponding to the weight adjustment of a certain type of sensor, and forming a parameter correlation mapping table. This table includes scenario type, sensor weight adjustment items, associated spatiotemporal calibration parameter optimization items, and collaborative effectiveness conditions. Preferably, in a further specific implementation of step 423, based on the parameter association mapping table, the sensor weight allocation adjustment results and spatiotemporal calibration accuracy optimization parameters are integrated to generate a sensor weight adjustment parameter configuration table. Specifically, according to the classification scenario and effective conditions, the sensor adjusted weight values, spatiotemporal calibration optimization parameters, parameter association relationships, and verification results are integrated into a structured configuration table. The core fields of the configuration table include scenario identifier, sensor type, adjusted weight, spatiotemporal calibration parameters (time compensation coefficient, extrinsic parameter matrix parameters), associated parameter ID, effective time, and failure conditions. The configuration table adopts a hierarchical index structure, supporting quick retrieval of corresponding parameter configurations by scenario type and sensor type. The generated sensor weight adjustment parameter configuration table can be directly used for parameter updates of the multimodal spatiotemporal calibration-weight conflict resolution fusion architecture and the vision-speech-action multimodal spatiotemporal calibration-weight conflict resolution fusion architecture, ensuring that the perception performance of the architecture is dynamically optimized with driving status feedback.
[0051] Optionally, step 43 includes: step 431, correcting the behavior matching weights of the park's main behavior database based on the dynamic prediction accuracy data in the standardized driving state feedback tensor, and obtaining behavior database matching weight correction parameters; step 432, optimizing the safety redundancy parameters of the low-speed obstacle avoidance path based on the obstacle avoidance effect quantification data in the standardized driving state feedback tensor, and obtaining safety redundancy parameter optimization results; step 433, fusing the behavior database matching weight correction parameters, safety redundancy parameter optimization results, and sensor weight adjustment parameter configuration table to generate a multimodal-path planning collaborative optimization parameter tensor, which prioritizes the optimization parameters of each module according to their dimensions, and the priority ranking is strongly correlated with autonomous driving safety indicators.
[0052] Preferably, the specific implementation process of step 431 is as follows: Based on the dynamic prediction accuracy data in the standardized driving state feedback tensor, the association mapping process between behavior categories and scene attributes is performed to generate a behavior prediction scene adaptation matrix; specifically, the dynamic prediction accuracy data is cross-classified in three dimensions according to dynamic subject type (pedestrians, park operation vehicles), behavior category (crossing roads, walking along the edge, turning operations, etc.) and park scene attributes (lighting conditions, road type, personnel density), and the average prediction accuracy, misjudgment rate and misjudgment type under each combination scenario are statistically analyzed; the rows of the matrix represent "behavior category". The "Different-Dynamic Subject Type" combination represents the combination of park scene attributes. The matrix element is the product of the prediction accuracy and scene adaptation coefficient for the corresponding combination scene. The scene adaptation coefficient is set according to the scene complexity (0.8-0.9 for complex scenes such as intersections, and 1.0-1.1 for simple scenes such as open roads). For example, the prediction accuracy of the combination "pedestrian crossing-backlight-intersection" is 0.65, the scene adaptation coefficient is 0.85, and the corresponding matrix element value is 0.65×0.85=0.5525. This matrix clearly presents the difference in prediction performance of various behaviors under different scenes. Preferably, in a further specific implementation of step 431, based on the behavior prediction scenario adaptation matrix, the behavior matching weights of the park's main behavior database are corrected to obtain behavior database matching weight correction parameters; specifically, each type of benchmark behavior data in the park's main behavior database is configured with initial matching weights. This application introduces a scenario adaptation gain factor and a prediction error compensation factor to construct a weight correction model; the scenario adaptation gain factor is determined by the corresponding element value in the behavior prediction scenario adaptation matrix (values range from 0.7 to 1.2), and the prediction error compensation factor is calculated by the inverse function of the prediction accuracy (the lower the prediction accuracy, the larger the compensation factor, value ranges from 0.05 to 0.2); specifically, the matching weights are iteratively updated using a Bayesian inference algorithm, with "minimizing the variance of behavior prediction accuracy under the same scenario" as the objective function, and a Gaussian process model is used to capture... To capture the nonlinear relationship between weights and prediction performance, the prior distribution parameters are dynamically adjusted during the iteration process (the initial prior distribution is a normal distribution with a mean of 0.5 and a variance of 0.1). The iteration convergence threshold is set to 0.001 (the number of iterations is usually 15-25). For scene combinations with a misclassification rate higher than 0.3 (such as pedestrians crossing at night), multi-scale feature enhancement terms are added to the corresponding baseline behavior data to expand the feature dimensions (from the original 25 dimensions to 40 dimensions, adding features such as nighttime contours and motion blur compensation). At the same time, the initial matching weights are reduced by 10%-15% to avoid mismatches. The corrected weights are organized according to the combination of "behavior category-subject type-scene attribute" to form the behavior database matching weight correction parameters, which include combination identifiers, corrected weights, scene adaptation gain factors, prediction error compensation factors, and feature enhancement identifiers.
[0053] Preferably, in the specific technical implementation of step 432, based on the obstacle avoidance effect quantification data in the standardized driving state feedback tensor, a correlation analysis between risk level and obstacle avoidance parameters is performed to generate an obstacle avoidance parameter risk response matrix. Specifically, the obstacle avoidance effect quantification data includes successful avoidance rate, risk contact frequency, obstacle avoidance acceleration change rate, and traffic efficiency loss rate, and is classified and statistically analyzed according to the dynamic subject risk level (high, medium, low) and obstacle avoidance scenario type (close-range sudden, long-range prediction, multiple obstacle superposition). The rows of the matrix represent the "risk level-scenario type" combination, and the columns represent the obstacle avoidance parameters. The obstacle safety redundancy parameter types (braking response time redundancy, lateral distance redundancy, speed reserve redundancy) are defined. Matrix elements represent the correlation between the current value of the corresponding parameter and the obstacle avoidance performance index in a given scenario (values range from 0 to 1; higher correlation indicates better parameter adaptability). For example, in the "high-risk - close-range sudden" combination, the correlation between braking response time redundancy and successful obstacle avoidance rate is 0.82, and the correlation between lateral distance redundancy and successful obstacle avoidance rate is 0.75, corresponding to matrix elements of 0.82 and 0.75 respectively. This matrix provides a correlation basis for optimizing safety redundancy parameters. Preferably, in a further specific implementation of step 432, based on the obstacle avoidance parameter risk response matrix, the safety redundancy parameters of the low-speed obstacle avoidance path are optimized to obtain the optimized safety redundancy parameter results. Specifically, Model Predictive Control (MMCC) is used. A multi-objective optimization model is constructed using the Control-MPC (Multi-Purpose Control, Multi-Purpose Control, Multi-Purpose Control) framework. The optimization objectives are to maximize the successful obstacle avoidance rate, optimize obstacle avoidance smoothness (minimum rate of acceleration change), and minimize the traffic efficiency loss rate. Safety redundancy parameters are used as optimization variables. Constraints include vehicle dynamics limitations (maximum braking acceleration ≤ 2.0 m / s², minimum turning radius ≥ 5 m), park road boundary constraints (lateral redundancy does not exceed the feasible road area), and parameter adjustment range limitations (single adjustment does not exceed 30%). The prediction time domain of the MPC framework is set to 3-5 seconds (adapting to low-speed autonomous driving response time), and the control time domain is set to 2-3 seconds. Optimal parameters are solved using a rolling time-domain optimization algorithm. Specifically, risk field theory is introduced during the optimization process. The risk level of a dynamic subject is transformed into a spatial risk potential energy distribution. The optimization direction of safety redundancy parameters is tilted towards high-risk potential energy regions. For example, in the "high-risk - close-range sudden" scenario, the braking response time redundancy is extended from 1.5 seconds to 2.2 seconds, and the lateral distance redundancy is expanded from 0.8 meters to 1.3 meters. At the same time, the Pareto optimization algorithm is used to handle multi-objective conflicts and screen the compromise solutions in the non-dominated solution set (allocated according to safety priority weights: successful avoidance rate 0.5, obstacle avoidance smoothness 0.3, and passage efficiency 0.2). The optimized parameters are classified and organized according to scenario type and risk level to form the safety redundancy parameter optimization results, including parameter type, optimized value, applicable scenario combination, optimization objective satisfaction, and constraint satisfaction.
[0054] Preferably, the specific implementation process of step 433 is as follows: Establish a priority quantification system for autonomous driving safety indicators; perform priority assignment processing on the behavior database matching weight correction parameters, safety redundancy parameter optimization results, and sensor weight adjustment parameter configuration table to generate a parameter priority quantification table; specifically, the safety indicator priority quantification system adopts the Analytic Hierarchy Process (AHP). The system is constructed using an Autonomous Driving Process (AHP). The target layer is "Comprehensive Safety of Autonomous Driving in Low-Speed Parks". The criterion layer includes four categories of indicators: collision risk avoidance, obstacle avoidance stability, perception reliability, and traffic efficiency. The solution layer consists of three categories of parameters to be optimized. A judgment matrix is constructed using the pairwise comparison method to calculate the weights of the criterion layer (collision risk avoidance 0.45, obstacle avoidance stability 0.25, perception reliability 0.2, and traffic efficiency 0.1). Then, the weights of the solution layer are calculated based on the correlation strength between the parameters and the criterion layer. For example, the correlation strength between the safety redundancy parameter and collision risk avoidance is 0.9, corresponding to a weight value of 0.45 × 0.9 = 0.405. The correlation strength between the sensor weight adjustment parameter and perception reliability is 0.85, corresponding to a weight value of 0.2 × 0.85 = 0.17. The parameter priority quantification table clearly presents the priority ranking of various parameters and parameter items (value range 0-1, higher values indicate higher priority).Preferably, in a further specific implementation of step 433, based on the parameter priority quantization table, tensor fusion processing of multi-dimensional parameters is performed to generate a multimodal-path planning collaborative optimization parameter tensor; specifically, firstly, the three types of parameters are constructed as independent three-dimensional tensors: the behavior database matching weight correction parameter is constructed as a four-dimensional tensor of "behavior category-subject type-scene attribute-weight value", the safety redundancy parameter optimization result is constructed as a four-dimensional tensor of "risk level-scene type-parameter type-parameter value", and the sensor weight adjustment parameter configuration table is constructed as a four-dimensional tensor of "sensor type-scene attribute-perception task-weight value"; Tensor Train Decomposition is adopted. The Decomposition (TTD) algorithm extracts features from three independent tensors. During decomposition, the goal is to minimize the decomposition error (with an error threshold of 0.01). It solves the core tensor and factor matrix using alternating least squares. The core tensor represents the intrinsic correlation features of the parameters, while the factor matrix represents the feature mapping relationships of each dimension. Based on the weight values of the parameter priority quantization table, the core tensor is reconstructed using weighted methods. During reconstruction, the feature representation of high-priority parameters is strengthened (the reconstruction coefficient for high-priority parameters is 0.8-1.0, and for low-priority parameters it is 0.5-0.7). Finally, a six-dimensional collaborative optimization parameter tensor is generated, consisting of "parameter module - safety criteria - scenario type - risk level - priority - parameter value". Each element of this tensor corresponds to an optimized parameter value under a specific combination of dimensions. The priority ranking is strongly correlated with autonomous driving safety indicators and can be directly used for parameter updates in multimodal spatiotemporal calibration-weight conflict resolution fusion architectures and vision-speech-action multimodal spatiotemporal calibration-weight conflict resolution fusion architectures, ensuring the global synergy of the optimized parameters of each module.
[0055] The above are merely preferred embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A low-speed autonomous driving method for campuses based on VLA, characterized in that, include: Step 1: Collect road image data, obstacle laser point cloud data, environmental audio data, park voice interaction data, and vehicle movement status data of the park environment to generate a panoramic park environment perception tensor. Step 2: Perform semantic element recognition and dynamic subject motion trend prediction on the panoramic park environment perception tensor to generate a semantic parsing topology of the park environment with dynamic motion trends; Step 3: Perform advance obstacle avoidance path planning and stability path optimization on the semantic parsing topology of the park environment with dynamic motion trends and the preset driving path to generate a set of park autonomous driving control instructions with dynamic safety redundancy. Step 4: Execute the park autonomous driving control instruction set with dynamic safety redundancy and collect vehicle driving status feedback data to generate a multimodal-path planning collaborative optimization parameter tensor. Continuously adjust the autonomous driving execution logic based on the multimodal-path planning collaborative optimization parameter tensor.
2. The low-speed campus autonomous driving method based on VLA according to claim 1, characterized in that, Step 1 includes: Step 11: Align the road image data, obstacle laser point cloud data, environmental audio data, park voice interaction data, and vehicle movement status data collected by the VLA multimodal heterogeneous sensing fusion unit with the spatiotemporal dimension benchmark to generate a multimodal initial calibration data tensor. Step 12: Perform multi-sensor accuracy weight assignment and scene adaptability calibration on the multimodal initial calibration data tensor to generate multimodal weight adaptation data; Step 13: Perform multi-source conflict detection and weighting resolution on the multimodal weighted adaptation data, and fuse them to generate a panoramic park environment perception tensor. The panoramic park environment perception tensor links and binds the road marking features, obstacle 3D features, abnormal sound spatial features, road marking visual features, voice interaction semantic features, and vehicle action status features in tensor dimensions, and the feature dimensions correspond one-to-one with the park's autonomous driving perception requirements.
3. The low-speed campus autonomous driving method based on VLA according to claim 2, characterized in that, Step 13 includes: Step 131: Perform feature conflict detection on the multimodal weighted adaptation data, identify conflict items such as visual-laser semantic contradiction, audio-visual spatial misalignment, visual-speech semantic contradiction, and action-visual state misalignment, and form a list of multimodal feature conflict items; Step 132: Perform weighted resolution on the list of conflicting multimodal features based on preset weights, retain the feature results of high-weight modes, and obtain the single-modal feature array after weight resolution; Step 133: After fusing and weighting the single-modal feature array, a panoramic park environment perception tensor is generated. The feature credibility of the panoramic park environment perception tensor is calculated by weighting the modal weights, and the three types of environmental features are unified based on a single dimension.
4. The low-speed campus autonomous driving method based on VLA according to claim 1, characterized in that, Step 2 includes: Step 21: Perform semantic element segmentation on the panoramic park environment perception tensor, identify four types of core semantic elements: road area, pedestrians, park operation vehicles, and fixed obstacles, and generate basic semantic segmentation results for the park. Step 22: Based on the park's main behavior database, perform dynamic motion trend prediction on pedestrians and park operation vehicles, and generate dynamic main motion trend prediction results; Step 23: Integrate the basic semantic segmentation results of the park with the dynamic subject movement trend prediction results to generate a semantic parsing topology of the park environment with dynamic movement trends. This topology associates the static semantic region with the dynamic subject movement trajectory through topological edges and provides a constraint basis for path planning.
5. The low-speed campus autonomous driving method based on VLA according to claim 4, characterized in that, Step 22 includes: Step 221: Retrieve the park's main behavior database. This database pre-stores benchmark behavior data such as typical trajectories of pedestrians crossing roads and typical turning operation patterns of vehicles in the park, which are related to high-frequency dynamic conflicts in the park, to form a benchmark behavior dataset of the park's main body. Step 222: Input the real-time location and movement speed of the dynamic subject into the behavior matching model, perform feature matching and probability calculation based on the benchmark behavior dataset of the park's subjects, and obtain the dynamic subject behavior trend index matrix; Step 223: Based on the dynamic subject behavior trend index matrix and the dynamic subject prediction trajectory derived from the park subject benchmark behavior dataset, generate the dynamic subject motion trend prediction result. The prediction time domain of this result is adapted to the response time requirements of low-speed autonomous driving and provides dynamic constraints for semantic topology.
6. The low-speed campus autonomous driving method based on VLA according to claim 1, characterized in that, Step 3 includes: Step 31: Perform autonomous driving feasible region and dynamic constraint boundary identification on the semantic parsing topology of the park environment with dynamic motion trend, and generate an initial driving path; Step 32: Perform advance obstacle avoidance planning on the initial driving path, and preset deceleration and detour trajectories for high-probability dynamic conflict scenarios to obtain an intermediate path with obstacle avoidance strategy; Step 33: Perform low-speed stability optimization on the intermediate path with obstacle avoidance strategy and add dynamic safety redundancy. Convert the optimized path into steering, acceleration and deceleration control commands to generate a park autonomous driving control command set with dynamic safety redundancy. This command set will link the control parameters including steering angle, acceleration and braking intensity with the dynamic subject motion trend prediction results and directly drive the vehicle to perform driving actions.
7. The low-speed campus autonomous driving method based on VLA according to claim 6, characterized in that, Step 33 includes: Step 331: Perform curvature smoothing optimization on the intermediate path with obstacle avoidance strategy to ensure that the path curvature adapts to the stability requirements of low-speed driving and obtain a smoothed obstacle avoidance path. Step 332: Add dynamic safety redundancy to the control parameters corresponding to the smoothed obstacle avoidance path, including extending the braking response time and increasing the lateral distance of obstacle avoidance, and generate a control parameter configuration table with safety redundancy; Step 333: Convert the control parameter configuration table with safety redundancy into steering, acceleration, and deceleration control commands to generate a park autonomous driving control command set with dynamic safety redundancy. The parameter range of this command set matches the low-speed park driving safety specifications.
8. The low-speed campus autonomous driving method based on VLA according to claim 1, characterized in that, Step 4 includes: Step 41: Collect vehicle driving trajectory data, obstacle avoidance effect evaluation data, sensor data matching degree data, multimodal data fusion matching degree data, and dynamic prediction accuracy data, and generate a standardized driving state feedback tensor after normalization processing. Step 42: Input the standardized driving state feedback tensor into the multimodal fusion-path planning closed-loop collaborative optimization model, adjust the sensor weights of the vision-speech-action multimodal spatiotemporal calibration-weight conflict resolution fusion architecture, and obtain the sensor weight adjustment parameter configuration table. Step 43: Based on the standardized driving state feedback tensor, correct the dynamic trend prediction logic and optimize the safety redundancy parameters. Integrate the sensor weight adjustment parameter configuration table with the above correction and optimization results to generate a multimodal-path planning collaborative optimization parameter tensor. Update the autonomous driving execution logic based on this parameter tensor.
9. The low-speed campus autonomous driving method based on VLA according to claim 8, characterized in that, Step 42 includes: Step 421: Based on the sensor data matching degree data and multimodal data fusion matching degree data in the standardized driving state feedback tensor, adjust the weight allocation of laser, vision, audio vision, voice and action state sensors, adjust the weight ratio of high matching degree modes, and obtain the sensor weight allocation adjustment result. Step 422: Based on the dynamic prediction accuracy data in the standardized driving state feedback tensor, optimize the temporal alignment accuracy of multimodal spatiotemporal calibration to obtain the spatiotemporal calibration accuracy optimization parameters; Step 423: Based on the sensor weight allocation adjustment results and the spatiotemporal calibration accuracy optimization parameters, generate a sensor weight adjustment parameter configuration table. This parameter configuration table provides a direct basis for the parameter update of the multimodal spatiotemporal calibration-weight conflict resolution fusion architecture vision-speech-action multimodal spatiotemporal calibration-weight conflict resolution fusion architecture.
10. The low-speed campus autonomous driving method based on VLA according to claim 8, characterized in that, Step 43 includes: Step 431: Based on the dynamic prediction accuracy data in the standardized driving state feedback tensor, correct the behavior matching weights of the park's main behavior database to obtain the behavior database matching weight correction parameters. Step 432: Based on the obstacle avoidance effect quantification data in the standardized driving state feedback tensor, optimize the safety redundancy parameters of the low-speed obstacle avoidance path and obtain the safety redundancy parameter optimization results; Step 433: Integrate the behavior database to match the weight correction parameters, safety redundancy parameter optimization results, and sensor weight adjustment parameter configuration table to generate a multimodal-path planning collaborative optimization parameter tensor. This tensor will prioritize the optimization parameters of each module according to their dimensions, and the priority order is strongly correlated with autonomous driving safety indicators.