Global path planning method based on large visual models

CN122041885BActive Publication Date: 2026-09-01BEIJING DECK SMART TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610151696.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-03
Publication Date
2026-09-01
Estimated Expiration
2046-02-03

AI Technical Summary

Technical Problem

[0003]现有一种传统的全局路径规划方案,其先基于预设的静态地图(如栅格地图)构建道路网络拓扑,再采用 A、Dijkstra 等经典算法,以路径长度最短为主要优化目标生成规划路径;在避障处理上,仅通过传感器实时检测当前障碍物位置,对路径进行局部调整以规避碰撞;路径评估仅关注路径长度与是否可通行,未考虑环境动态变化对通行效率的影响

Benefits of technology

在环境建模层面,通过获取目标场景全局环境图像与目标位置信息,结合视觉大模型的多尺度语义融合能力,生成带动态通行优先级的实时语义全局地图。该地图不仅能精准识别道路、障碍物、通行限制区域等语义信息,还能基于时序环境数据动态调整各区域通行优先级,使地图能实时匹配环境变化,从根源上解决了传统静态地图与实际通行需求脱节的问题,为路径规划提供了更贴合实际的环境基础。在路径生成层面,而本方法基于动态语义全局地图,生成带避障预测的候选全局路径特征张量。通过关联动态障碍物的视觉语义特征与历史运动数据,提前预判障碍物未来预设时段的运动轨迹及动态影响区域,使候选路径自带避障预测属性,相比传统被动避障方式,能减少路径重规划频率,提升通行的流畅性与安全性,有效缓解了传统方案中频繁启停、绕行的问题。在路径评估层面,本方法构建五维度路径评估体系,综合考量通行时间、避障难度、路径平滑度、通行优先级、路径 -环境适配度,且各维度权重基于实时环境数据波动幅度自适应调整。这种多维度自适应评估方式,能根据不同环境场景与任务需求,筛选出更合理的最优路径,相比传统单一维度评估,路径规划的综合性与适配性更高,能更好地满足复杂动态环境下的导航需求。在闭环优化层面,而本方法将机器人行驶的全流程数据映射至实时语义全局地图,进行路径规划适配。通过计算路径空间执行偏差率、节点通行时效偏差率等量化指标,反向修正地图通行优先级系数、规划目标权重配比及轨迹预测参数,使系统具备持续学习能力,随着运行时间积累,路径规划的合理性与准确性逐步提升,解决了传统方案长期运行下鲁棒性下降的问题。与传统全局路径规划方案相比,本方法通过 “动态建模 - 预测生成 - 多维评估 - 闭环适配” 的全流程设计,实现了从静态规划到动态适配、从被动避障到主动预判、从固定评估到自适应优化的转变,显著提升了机器人在复杂动态环境中的通行效率与任务完成质量,为自主移动机器人在各类园区场景的大规模应用提供了可靠技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122041885B_ABST
    Figure CN122041885B_ABST
Patent Text Reader

Abstract

This application discloses a global path planning method based on a large visual model. The method includes: acquiring a global environment image of the target scene and target location information to generate a real-time semantic global map with dynamic passage priorities; generating a candidate global path feature tensor with obstacle avoidance prediction based on the real-time semantic global map with dynamic passage priorities; constructing a five-dimensional path evaluation system including travel time, obstacle avoidance difficulty, path smoothness, passage priority, and path-environment adaptability, to generate a dynamically environment-adaptive global path planning topology based on the candidate global path feature tensor with obstacle avoidance prediction; and mapping the entire process data of the robot traveling along the dynamically environment-adaptive global path planning topology to the real-time semantic global map with dynamic passage priorities for global path planning adaptation. This improves the robot's travel efficiency and task completion quality in complex dynamic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot autonomous navigation technology, and more specifically, to a global path planning method based on a large visual model. Background Technology

[0002] In closed or semi-closed environments such as industrial parks, science parks, and campuses, the application of autonomous mobile robots (such as unmanned delivery vehicles and inspection robots) is becoming increasingly widespread. These robots need to achieve efficient and safe autonomous navigation in dynamically changing environments. Global path planning, as the core of the navigation system, directly determines the robot's operating efficiency and task completion quality.

[0003] A traditional global path planning scheme first constructs a road network topology based on a pre-set static map (such as a raster map), and then uses A... Classic algorithms such as Dijkstra's algorithm generate planned paths with the shortest path length as the main optimization objective. In obstacle avoidance, they only use sensors to detect the current obstacle position in real time and make local adjustments to the path to avoid collisions. Path evaluation only focuses on path length and whether it is passable, without considering the impact of dynamic environmental changes on passage efficiency.

[0004] This traditional approach has significant technical drawbacks: due to its reliance on static maps, it cannot respond promptly to dynamic changes in the environment (such as temporary construction or regional congestion), resulting in planned routes often being out of sync with actual traffic demands; it lacks forward-looking prediction of dynamic obstacles, and can only passively avoid them, easily leading to frequent starts, stops, or detours; and its assessment dimensions are limited, failing to comprehensively balance multiple requirements such as traffic efficiency, obstacle avoidance difficulty, and path adaptability, resulting in low rationality and environmental adaptability of the route planning. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a global path planning method based on a large visual model, which can at least alleviate the aforementioned technical problems.

[0006] The technical solutions provided in this application are as follows:

[0007] A global path planning method based on a large visual model includes: acquiring a global environment image of the target scene and target location information to generate a real-time semantic global map with dynamic passage priority; generating a candidate global path feature tensor with obstacle avoidance prediction based on the real-time semantic global map with dynamic passage priority; constructing a five-dimensional path evaluation system including passage time, obstacle avoidance difficulty, path smoothness, passage priority, and path-environment adaptability to generate a dynamic environment-adaptive global path planning topology based on the candidate global path feature tensor with obstacle avoidance prediction; and mapping the entire process data of the robot traveling along the dynamic environment-adaptive global path planning topology to the real-time semantic global map with dynamic passage priority for global path planning adaptation.

[0008] The technical solution of this application has the following technical advantages: At the environmental modeling level, by acquiring global environmental images and target location information of the target scene, and combining the multi-scale semantic fusion capabilities of the large visual model, a real-time semantic global map with dynamic traffic priority is generated. This map can not only accurately identify semantic information such as roads, obstacles, and traffic restriction areas, but also dynamically adjust the traffic priority of each area based on temporal environmental data, enabling the map to match environmental changes in real time. This fundamentally solves the problem of traditional static maps being out of touch with actual traffic needs, providing a more realistic environmental foundation for path planning. At the path generation level, this method generates candidate global path feature tensors with obstacle avoidance prediction based on the dynamic semantic global map. By associating the visual semantic features of dynamic obstacles with historical motion data, the movement trajectory and dynamic influence area of ​​obstacles in the future for a preset period are predicted in advance, giving the candidate paths built-in obstacle avoidance prediction attributes. Compared with traditional passive obstacle avoidance methods, this reduces the frequency of path replanning, improves the smoothness and safety of traffic, and effectively alleviates the problems of frequent starts, stops, and detours in traditional solutions. At the path evaluation level, this method constructs a five-dimensional path evaluation system, comprehensively considering travel time, obstacle avoidance difficulty, path smoothness, travel priority, and path-environment adaptability. The weights of each dimension are adaptively adjusted based on real-time environmental data fluctuations. This multi-dimensional adaptive evaluation method can select more reasonable optimal paths according to different environmental scenarios and task requirements. Compared with traditional single-dimensional evaluation, path planning is more comprehensive and adaptable, better meeting navigation needs in complex dynamic environments. At the closed-loop optimization level, this method maps the robot's entire driving process data to a real-time semantic global map for path planning adaptation. By calculating quantitative indicators such as path space execution deviation rate and node travel time deviation rate, the map travel priority coefficient, planning target weight ratio, and trajectory prediction parameters are corrected in reverse, enabling the system to have continuous learning capabilities. As running time accumulates, the rationality and accuracy of path planning gradually improve, solving the problem of decreased robustness of traditional solutions under long-term operation. Compared with traditional global path planning schemes, this method achieves a transformation from static planning to dynamic adaptation, from passive obstacle avoidance to active prediction, and from fixed evaluation to adaptive optimization through a full-process design of "dynamic modeling - prediction generation - multi-dimensional evaluation - closed-loop adaptation". This significantly improves the robot's mobility and task completion quality in complex dynamic environments, and provides reliable technical support for the large-scale application of autonomous mobile robots in various park scenarios. Attached Figure Description

[0009] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0010] Figure 1 This is a flowchart of a global path planning method based on a large visual model according to an embodiment of the present invention. Detailed Implementation

[0011] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. Various aspects are provided by way of explanation and not limitation of the invention. Indeed, those skilled in the art will recognize that modifications and variations can be made to the invention without departing from its scope or spirit. For example, a feature represented or described as part of one embodiment may be used in another embodiment to produce yet another embodiment. Therefore, it is desirable that the invention encompass such modifications and variations falling within the scope of the appended claims and their equivalents.

[0012] In the description of this invention, the terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," and "bottom," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and do not require the invention to be constructed and operated in a specific orientation; therefore, they should not be construed as limitations on the invention. The terms "connected," "linked," and "set up" used in this invention should be interpreted broadly. For example, they can refer to a fixed connection or a detachable connection; they can refer to a direct connection or an indirect connection through intermediate components. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.

[0013] like Figure 1 As shown, this invention provides a global path planning method based on a large visual model, comprising: Step 1, acquiring a global environment image of the target scene and target location information to generate a real-time semantic global map with dynamic passage priority; Step 2, generating a candidate global path feature tensor with obstacle avoidance prediction based on the real-time semantic global map with dynamic passage priority; Step 3, constructing a five-dimensional path evaluation system including passage time, obstacle avoidance difficulty, path smoothness, passage priority, and path-environment adaptability, to generate a dynamic environment-adaptive global path planning topology based on the candidate global path feature tensor with obstacle avoidance prediction; Step 4, mapping the entire process data of the robot traveling along the dynamic environment-adaptive global path planning topology to the real-time semantic global map with dynamic passage priority for global path planning adaptation.

[0014] Optionally, step 1 includes: Step 11, inputting the global environmental image of the target scene into the multi-scale semantic fusion branch of the visual large model, extracting the shallow spatial contour features and deep category semantic features of the image respectively, performing cross-scale fusion on the two types of features, and outputting the semantic labels, spatial coordinates and category confidence of roads, obstacles, and traffic restriction areas to generate an initial environmental semantic feature map; Step 12, collecting the temporal environmental data of the target scene, including regional congestion duration sequence, temporary construction range change data, and traffic bandwidth fluctuation data, constructing a regional trafficability assessment model with a time decay factor, inputting the temporal environmental data of the target scene into the model to obtain the trafficability quantification value of each region, configuring the traffic priority coefficient for each region in the initial environmental semantic feature map based on the quantification value, and generating a traffic priority feature layer; Step 13, binding the geographic grid of the initial environmental semantic feature map to the coefficient grid of the traffic priority feature layer one by one, performing spatial topology verification on the bound grid, and generating a real-time semantic global map with dynamic traffic priority.

[0015] Optionally, step 11 includes: step 111, performing block segmentation and scale normalization preprocessing on the global environment image of the target scene to generate a multi-scale image block sequence; step 112, inputting the multi-scale image block sequence into the multi-scale semantic fusion branch of the visual large model, extracting spatial contour features from each block through shallow convolution, extracting category semantic features through a deep Transformer module, and completing cross-scale feature association through a feature fusion layer; step 113, integrating all block fusion features corresponding to the multi-scale image block sequence, eliminating semantic breaks at block boundaries, and generating an initial environment semantic feature map.

[0016] Preferably, the specific implementation process of step 111 includes: performing adaptive block size division and scale normalization processing on the global environment image of the target scene to generate a multi-scale image block sequence; specifically, firstly, resolution analysis and semantic region prediction are performed on the global environment image of the target scene. The semantic region prediction uses a lightweight object detection algorithm to quickly identify the distribution range and size features of core semantics such as roads, obstacles, and traffic restriction areas in the image. If a large continuous road area is detected (accounting for more than 35% of the total image area), a larger block size (e.g., 512 pixels × 512 pixels, 1024 pixels × 1024 pixels) is adopted to reduce the number of blocks and improve processing efficiency; if densely distributed small dynamic obstacles (such as pedestrians, delivery vehicles, etc., with a single obstacle pixel area of ​​less than 8000 pixels) are detected, a smaller block size (e.g., 256 pixels × 256 pixels, 384 pixels × 384 pixels) is adopted to ensure the integrity of the obstacle semantic information. The block size is set to an integer power of 2 to accommodate the computational needs of convolutional layers and Transformer modules in large visual models, avoiding size mismatch issues during feature processing. Subsequently, scale normalization is performed on each block using an adaptive Z-Score normalization method. The mean μ and standard deviation σ of all pixel values ​​within the block are calculated, and each pixel value x is converted to (x-μ) / σ. This ensures that the normalized pixel values ​​conform to a standard normal distribution (mean 0, standard deviation 1), eliminating the impact of pixel value fluctuations caused by different lighting conditions (e.g., strong sunlight on sunny days, weak light on cloudy days, and nighttime lighting) and differences in imaging equipment. To avoid semantic information loss at block boundaries, overlapping regions are set between adjacent blocks. The overlap ratio is dynamically adjusted according to the block size; the smaller the block size, the higher the overlap ratio (e.g., 20% overlap for a 256-pixel block and 10% overlap for a 1024-pixel block). Pixel values ​​in the overlapping region belong to both adjacent blocks simultaneously, ensuring the continuity of semantics at the boundaries. Finally, based on the spatial location information of the blocks in the original global environment image, a location index label is added to each block to form a multi-scale image block sequence containing different scales, no missing edge information, and location indexes.

[0017] Preferably, in the specific technical implementation of step 112, based on the multi-scale image block sequence, hierarchical feature extraction and cross-scale association are performed through the multi-scale semantic fusion branch of the visual large model to generate block fusion features; specifically, the multi-scale semantic fusion branch of the visual large model adopts a three-level architecture of "shallow convolutional feature extraction sub-network - deep Transformer semantic modeling sub-network - cross-scale attention fusion sub-layer", and each level is customized for the semantic characteristics of the park scene to ensure the pertinence and effectiveness of feature extraction. The shallow convolutional feature extraction subnetwork consists of 5 convolutional layers with an increasing kernel size configuration: 3×3 kernels for the first and second layers, 5×5 kernels for the third and fourth layers, and 7×7 kernels for the fifth layer. The stride is set to 2 for all layers, and SAME padding is used. By progressively increasing the kernel size, the receptive field is expanded. Simultaneously, downsampling with a stride of 2 gradually reduces the size of the image blocks (e.g., a 256-pixel block is reduced to 8 pixels after 5 convolutional layers). This effectively extracts spatial contour features such as texture gradients, edge contours, and corner distribution from the blocks. These features accurately represent the geometric shape of road boundaries, the outline of obstacles, and the boundaries of restricted areas, generating a block spatial contour feature map. Each pixel in this block spatial contour feature map corresponds to a local region in the original block, and the pixel value is a spatial feature vector (512-dimensional) for that region. Each element in the vector corresponds to a specific spatial feature attribute (such as edge strength, texture direction, curvature). The deep Transformer semantic modeling subnetwork comprises a 12-layer Transformer encoder, with each layer having 16 attention heads, each with a 64-dimensional dimension. A multi-head attention mechanism is used to calculate the association weights of feature vectors at different locations in the block spatial contour feature map, capturing semantic association information globally, such as the topological association between roads and adjacent sidewalks, the exclusion relationship between obstacles and passageways, and the correspondence between traffic signs and road functions. Simultaneously, a park-specific semantic dictionary is introduced, containing 28 core semantic labels covering high-frequency semantic categories in park scenarios such as main roads, sidewalks, pedestrians, motor vehicles, non-motor vehicles, construction barriers, traffic signs, restricted areas, and logistics loading / unloading areas. Each semantic label is converted into a semantic embedding vector (512-dimensional) through an embedding layer. After dimensional alignment with the block spatial contour feature map, this vector is input into the Transformer encoder. The encoder uses a self-attention mechanism to fuse spatial features and semantic embedding vectors, extracting the category semantic features corresponding to each semantic label and generating a block category semantic feature vector (512-dimensional). Each element in the vector corresponds to a confidence score for a semantic category. The higher the score, the greater the probability that the block belongs to the corresponding semantic category.The cross-scale attention fusion sublayer employs a semantic confidence-weighted fusion mechanism. First, it unifies the dimensions of the block spatial contour feature map and the block category semantic feature vector. The block category semantic feature vector is then converted to the same size as the spatial contour feature map (e.g., 8 pixels × 8 pixels × 512 dimensions) through a fully connected layer. Next, it calculates the mean semantic confidence at each location. Locations with higher mean semantic confidence indicate a clearer semantic category, and thus receive a higher fusion weight for the category semantic feature. Conversely, locations with lower mean semantic confidence indicate a more ambiguous semantic category, and thus receive a higher fusion weight for the spatial contour feature. The fusion weights are dynamically generated using an adaptive function to ensure the complementarity of spatial and semantic features. A pixel-by-pixel weighted sum is performed on the two types of features to generate a block-level fusion feature for each block. This block-level fusion feature retains precise spatial location information while containing clear category semantic attributes, achieving deep coupling between spatial and semantic information.

[0018] Preferably, in a scenario, step 113 involves performing boundary semantic compensation and global integration processing on all block fusion features to generate an initial environmental semantic feature map. Specifically, firstly, for block fusion features with overlapping regions in the multi-scale image block sequence, boundary semantic consistency verification and compensation are performed. An overlapping region feature difference matrix is ​​constructed, where the rows and columns correspond to the width and height directions of the overlapping region, respectively. Each element in the matrix is ​​the cosine similarity of the block fusion feature vectors of two adjacent blocks at that position. The lower the cosine similarity, the greater the feature difference and the higher the risk of semantic discontinuity. When the cosine similarity is lower than a preset threshold (e.g., 0.75), the boundary semantic compensation algorithm is activated. Based on the feature distribution pattern of non-overlapping regions of adjacent blocks, interpolation is used to adjust the feature vectors of the difference regions, making the feature vectors of the overlapping regions transition smoothly and eliminating semantic discontinuity at the block boundaries. Subsequently, based on the location index labels of the multi-scale image block sequence, a block spatial mapping matrix is ​​constructed. This matrix is ​​a two-dimensional matrix, with rows and columns corresponding to the width and height directions of the original global environment image, respectively. Each element in the matrix stores the storage address and size information of the corresponding block fusion feature. This matrix allows for rapid location of the coordinate range of each block fusion feature in the original image. Based on this location mapping matrix, all block fusion features that have undergone boundary compensation are stitched together according to their original spatial positions. During the stitching process, the size of the block fusion features is fine-tuned using a bilinear interpolation algorithm to ensure seamless connection of feature maps from adjacent blocks, forming a global feature matrix. The size of this global feature matrix is ​​consistent with the size of the original global environment image, and the feature vector of each pixel has a dimension of 512, containing both spatial contour and category semantic information. Finally, the global feature matrix is ​​subjected to semantic dimension enhancement and dimensionality reduction. The feature vector dimension is mapped from 512 dimensions to 128 dimensions through a 1×1 convolutional layer. Each dimension corresponds to a semantic category specific to the park scene. The bias term of the convolutional layer is initialized according to the frequency and importance of each semantic category in the park scene (the higher the frequency and the stronger the importance of the semantic category, the larger the corresponding bias term value). After convolution processing, each pixel of the global feature matrix corresponds to a 128-dimensional semantic feature vector. The value of each element in the vector represents the confidence of the pixel belonging to the corresponding semantic category, generating an initial environmental semantic feature map containing global spatial semantic information and without boundary faults.

[0019] Preferably, the training process of the large visual model is specifically implemented as follows: adopting a scene-adaptive training strategy, a training dataset covering various closed or semi-closed scenes such as industrial parks, science and technology parks, campuses, and hospitals is constructed. The images in the dataset include different environmental conditions (such as sunny days, rainy days, snowy days, and foggy days), different time periods (such as morning, noon, evening, and night), and different dynamic scenes (such as peak pedestrian flow, vehicle traffic, temporary construction, and material loading and unloading). The image resolution covers multiple specifications such as 2K (2048×1080 pixels), 4K (4096×2160 pixels), and 8K (7680×4320 pixels), and the total number of samples is not less than 150,000. Data augmentation is performed on the training data, including random cropping (cropping ratio range of 0.6-1.0), random rotation (rotation angle range of -18° to 18°), brightness adjustment (brightness gain range of 0.7-1.3), contrast adjustment (contrast gain range of 0.7-1.3), horizontal flipping (flipping probability of 50%), and Gaussian noise addition (noise variance range of 0.001-0.01). Data augmentation expands the diversity of training data and improves the generalization ability of the model. A multi-task joint loss function is designed, comprising three parts: semantic segmentation loss, feature alignment loss, and boundary consistency loss. The semantic segmentation loss uses a weighted cross-entropy loss function, assigning higher weights to key semantic categories such as roads and obstacles (e.g., roads with a weight of 2.0, obstacles with a weight of 1.8, and other categories with a weight of 1.0), calculating the difference between the semantic labels output by the model and the true semantic labels. The feature alignment loss uses a mean squared error loss function, calculating the fusion error between the block spatial contour features and the block category semantic features, ensuring semantic consistency between the two types of features. The boundary consistency loss uses an L1 loss function, calculating the difference in feature vectors in overlapping block regions, ensuring the semantic continuity of block boundaries. The total loss of the multi-task joint loss function is the weighted sum of the three losses, with the weight ratios dynamically adjusted according to the training stage. Initially, the semantic segmentation loss weight is 0.5, the feature alignment loss weight is 0.3, and the boundary consistency loss weight is 0.2. As the number of training iterations increases, the weights are gradually adjusted to 0.4 for semantic segmentation loss, 0.3 for feature alignment loss, and 0.3 for boundary consistency loss. The AdamW optimizer was used during training. The initial learning rate was set to 1.5e-4, the weight decay coefficient was 1e-5, and a cosine annealing learning rate adjustment strategy was adopted. Every 8 epochs, the learning rate was reduced to 0.85 times the current value. The total number of training iterations was 100 epochs. An early stopping strategy was adopted during training. Training was stopped when the semantic recognition accuracy on the validation set did not improve for 6 consecutive epochs.After training, the model achieves an overall semantic recognition accuracy of no less than 94% on the validation set, with an accuracy of no less than 96% for recognizing roads, obstacles, and restricted areas, and no less than 92% for recognizing dynamic obstacles (pedestrians and vehicles). The model's shallow convolutional sub-network can stably extract spatial contour features in different scenarios, the deep Transformer sub-network can accurately identify various environmental semantic categories, and the cross-scale feature fusion sub-layer can effectively associate two types of features. Preferably, the effectiveness verification of the block-fused features is specifically implemented by adopting a triple verification mechanism of "feature response value analysis + semantic recognition accuracy verification + spatial positioning accuracy verification" to ensure that the block-fused features can accurately represent the semantic and spatial information of the park scene. Feature response value analysis calculates the response value intensity of each semantic category in the block-fused feature. For each semantic category, the feature response value (the value of the corresponding dimension in the feature vector) of all pixels in the block-fused feature map corresponding to that category is calculated. The proportion of feature points with response value intensity higher than a preset threshold (e.g., 0.7) is counted. If the proportion exceeds 82%, the block-fused feature is considered effective in representing that semantic category. For blocks containing multiple semantic categories, it is necessary to ensure that the proportion of effective feature points for each major semantic category (the category occupying more than 12% of the block area) exceeds 78%. Semantic recognition accuracy verification is achieved by building an independent semantic decoding network. The semantic decoding network consists of 3 fully connected layers and 1 softmax layer. The block-fused feature is input into the semantic decoding network, and the predicted probability of each semantic category is output. The matching degree between the semantic category with the highest predicted probability and the real semantic category is calculated. If the matching degree is higher than 88%, the block-fused feature is considered effective. Spatial positioning accuracy verification is achieved by calculating the Intersection over Union (IoU) ratio between the semantic region corresponding to the fused block feature and the true semantic region. If the IoU is higher than 0.8, the spatial positioning attribute of the fused block feature is considered valid. For fused block features that fail verification, optimization is performed by tracing back to the previous processing steps: if the failure is due to inappropriate block size, the block size of the region is adjusted and reprocessed; if the failure is due to unreasonable feature fusion weights, the weight allocation rules of the cross-scale feature fusion sub-layer are dynamically adjusted; if the failure is due to insufficient model training, training samples for the corresponding scene are added and the model is retrained. Through this triple verification mechanism, it is ensured that all fused block features can provide reliable support for the generation of the subsequent initial environmental semantic feature map, improving the accuracy and stability of the initial environmental semantic feature map.

[0020] Optionally, step 12 includes: Step 121, constructing a regional trafficability assessment model, with the model input being the time-series environmental data of the target scenario, and the assessment dimensions including the regional congestion index, the proportion of construction impact area, and the traffic bandwidth redundancy, with each dimension configured with a time decay factor to weaken the weight of historical data; Step 122, collecting real-time congestion data, construction area range change data, and road cross-section traffic bandwidth data from the time-series environmental data of the target scenario, and completing data quantification and normalization according to the assessment dimensions; Step 123, inputting the quantified data into the regional trafficability assessment model, outputting the trafficability quantification value of each region, assigning traffic priority coefficients to each region in the initial environmental semantic feature map according to the quantification value range, and adjusting the coefficients of congested areas and construction areas according to the rules corresponding to the time decay factor to generate a traffic priority feature layer.

[0021] Preferably, the specific implementation process of step 121 includes: based on the environmental characteristics and traffic requirements of the target scenario, constructing a regional trafficability assessment model containing a multi-dimensional evaluation module and a time decay mechanism to achieve accurate analysis of time-series environmental data and quantification of trafficability; specifically, the regional trafficability assessment model adopts a four-level architecture of "input layer - feature processing layer - multi-dimensional evaluation layer - fusion output layer", with each level customized for the dynamic characteristics of the park scenario. The input layer is responsible for receiving time-series environmental data of the target scenario, including regional congestion duration sequence, temporary construction area change data, and traffic bandwidth fluctuation data. The data sampling frequency is set according to the dynamic change frequency of the scenario (e.g., 1Hz to 10Hz) to ensure timely capture of environmental changes. The feature processing layer smooths and denoises the input time-series data and extracts trends. A sliding window averaging method is used to smooth the data, with the window size dynamically adjusted based on the data type (5 to 10 sampling points for congestion data, and 10 to 20 sampling points for construction data) to eliminate random noise during data acquisition. A linear regression algorithm is used to extract trends in the time-series data, identifying trends such as congestion worsening / easing, construction area expansion / contraction, and traffic bandwidth increase / decrease, generating hierarchical time-series feature vectors. The multi-dimensional assessment layer contains three parallel assessment submodules, corresponding to regional congestion index assessment, construction impact area percentage assessment, and traffic bandwidth redundancy assessment, respectively. Each submodule has independent calculation logic and parameters. The regional congestion index assessment submodule analyzes congestion duration sequences and trend characteristics to calculate the density change rate and average dwell time of obstacles (pedestrians and vehicles) within a unit of time, generating a regional congestion index. The construction impact range proportion assessment submodule calculates the proportion and rate of change of the construction area to the total area of ​​the region based on construction range change data and trend characteristics, generating a construction impact range proportion. The traffic bandwidth redundancy assessment submodule calculates the difference and changes between the actual usable width of the road cross-section and the required passage width for robots based on traffic bandwidth fluctuation data and trend characteristics, generating traffic bandwidth redundancy. The fusion output layer adopts a dynamic weighted fusion mechanism, configuring initial weights for the three evaluation dimensions (e.g., regional congestion index weight 0.4, construction impact area ratio weight 0.3, and traffic bandwidth redundancy weight 0.3), and introducing a time decay factor to weaken the weight of historical data. The time decay factor is dynamically generated based on the difference between the data collection time and the current time. The larger the difference, the smaller the decay factor (e.g., the decay factor is 0.9 when the difference is 10 seconds, 0.7 when the difference is 30 seconds, and 0.5 when the difference is 60 seconds), ensuring that the model focuses more on recent data for evaluation, and finally outputs the trafficability quantification value of each region.Preferably, in a further specific implementation of step 121, a factor-weight linkage mechanism is constructed to improve the scenario adaptability of the model evaluation, in order to dynamically adapt the time decay factor and adjust the multi-dimensional weights in a scenario-based manner. Specifically, the time decay factor is not a fixed mapping, but rather the decay rate is dynamically adjusted in combination with the data type and scenario. For rapidly changing time-series data such as congestion data, an exponential decay mode is adopted, and the decay rate coefficient is set to a high value (e.g., 0.15) to ensure that outdated data from the distant past is quickly forgotten. For relatively slow-changing data such as construction data, a linear decay mode is adopted, and the decay rate coefficient is set to a low value (e.g., 0.05) to avoid evaluation distortion due to untimely data updates. Simultaneously, a multi-dimensional weighted, scenario-based adjustment rule base is constructed. This rule base includes weight configuration schemes for different scenarios (such as weekday peak hours, weekend off-peak hours, emergency task periods, and intensive construction periods). For example, during weekday peak hours, the regional congestion index weight is increased to 0.5, and the traffic bandwidth redundancy weight is decreased to 0.2; during emergency task periods, the traffic bandwidth redundancy weight is increased to 0.4, and the construction impact area percentage weight is decreased to 0.2; during intensive construction periods, the construction impact area percentage weight is increased to 0.4, while maintaining the regional congestion index weight at 0.4. Scenario recognition is automatically triggered by analyzing the statistical characteristics of time-series environmental data (such as peak values ​​of congestion data, coverage of construction data, and fluctuations in traffic bandwidth data), without requiring additional triggering commands. This allows the model to adapt to the evaluation needs of different scenarios, further improving the accuracy of traffic quantification values.

[0022] Preferably, in the specific technical implementation of step 122, based on the input requirements of the regional trafficability assessment model, the time-series environmental data of the target scenario is specifically collected, quantified, and normalized to generate standardized time-series environmental data. Specifically, a multi-source data acquisition system is first constructed, integrating multi-source data acquisition devices such as visual sensors, park management system API interfaces, and LiDAR to achieve comprehensive collection of time-series environmental data. Real-time congestion data is obtained by processing image data collected by visual sensors through crowd counting and vehicle detection algorithms. The crowd counting algorithm uses a density map estimation method, generating a density map by detecting pedestrian head features in the image, integrating to obtain the number of pedestrians in the area, and calculating the pedestrian density by combining the area area. The vehicle detection algorithm uses the YOLO series target detection algorithm to identify the number and type of vehicles in the area, calculating the vehicle density by combining the area area, and generating real-time congestion data by combining pedestrian density and vehicle density. The data on changes in the construction area are obtained through the park management system's API interface, which provides construction reporting information (including construction start time, estimated end time, and initial construction scope). Combined with image data collected by visual sensors, the actual construction scope is identified using a semantic segmentation algorithm. The amount and rate of change of the construction scope are calculated to generate the data on changes in the construction area's scope. The traffic bandwidth data for the road cross-section is obtained by scanning the road cross-section with LiDAR, acquiring the three-dimensional coordinates of the road's two sides, calculating the actual width of the road cross-section, and subtracting the width occupied by obstacles on both sides of the road (such as guardrails and green belts) to obtain the actual passable width, i.e., the traffic bandwidth data of the road cross-section. Subsequently, the three types of data are quantified: regional congestion data is quantified as a regional congestion index between 0 and 1 (0 indicates no congestion, 1 indicates severe congestion); the data on changes in the construction area's scope is quantified as the percentage of construction impact area between 0 and 1 (0 indicates no construction impact, 1 indicates the entire area is occupied by construction); and the traffic bandwidth data for the road cross-section is quantified as traffic bandwidth redundancy between 0 and 1 (0 indicates no traffic redundancy, 1 indicates sufficient traffic bandwidth). Finally, the Min-Max normalization method is used to normalize the quantized data, mapping the data to a unified range of 0 to 1, eliminating the dimensional differences between different data types, generating standardized time-series environmental data, and ensuring that the data can adapt to the input requirements of the regional accessibility assessment model. Preferably, in the specific implementation of step 122, a synchronization-verification mechanism is constructed to ensure data quality for data synchronization and outlier handling during multi-source data acquisition. Specifically, multi-source data synchronization adopts a timestamp alignment strategy, configuring a high-precision clock module (clock error ≤ 1ms) for each acquisition device. All data acquisitions are appended with precise timestamps. After the data is transmitted to the processing module, the timestamps of the visual sensor data are used as a benchmark to calibrate the timestamps of the park management system API interface data and LiDAR data. The calibrated timestamp deviation is controlled within 5ms to avoid distortion of quantization results due to data asynchrony.Simultaneously, a data outlier detection and correction mechanism is constructed. The 3σ criterion is used to identify outliers, which involves calculating the mean μ and standard deviation σ of various data types and classifying data exceeding the range [μ-3σ, μ+3σ] as outliers. For outlier congestion data, a weighted average of adjacent timestamp data is used for correction, with the weight decreasing as the time interval increases. For outlier construction area data, construction reporting information from the park management system's API interface is used for correction; if the actual identified construction area differs from the reported area by more than 30%, the reported area is used as the benchmark for adjustment. For outlier bandwidth data, the average of multiple LiDAR scans is used for correction, with the number of scans dynamically adjusted according to the degree of anomaly (3 scans for mild anomalies, 5 scans for severe anomalies) to ensure the accuracy and reliability of the standardized time-series environmental data after quantification and normalization.

[0023] Preferably, in one scenario, when step 123 is specifically implemented, standardized time-series environmental data is input into the regional accessibility assessment model to generate accessibility quantification values ​​for each region. Based on the quantification values ​​and dynamic adjustment rules, accessibility priority coefficients are allocated to generate a accessibility priority feature layer. Specifically, firstly, standardized time-series environmental data is input into the input layer of the regional accessibility assessment model. After processing by the feature processing layer, a hierarchical time-series feature vector is generated and passed into the multi-dimensional assessment layer for evaluation of each dimension. Evaluation results of three dimensions are generated: regional congestion index, construction impact range ratio, and access bandwidth redundancy. The fusion output layer combines the time decay factor and dynamic weights to weight and fuse the evaluation results of the three dimensions, and outputs the accessibility quantification values ​​for each region (within the range of 0 to 1, where a larger value indicates better accessibility). Subsequently, a rule base for allocating access priority coefficients is constructed, dividing the accessibility quantification values ​​into multiple intervals (e.g., 0.8 to 1.0 is the first interval, 0.6 to 0.8 is the second interval, 0.4 to 0.6 is the third interval, 0.2 to 0.4 is the fourth interval, and 0 to 0.2 is the fifth interval). Each interval corresponds to a basic access priority coefficient (e.g., the first interval corresponds to 1.0, the second interval to 0.8, the third interval to 0.6, the fourth interval to 0.4, and the fifth interval to 0.2). For congested and construction areas, the basic traffic priority coefficient is lowered according to the rules corresponding to the time decay factor. If the area is congested, the reduction is positively correlated with the congestion index (reduced by 0.1 when the congestion index is 0.5, and by 0.2 when the congestion index is 0.8). If the area is under construction, the reduction is positively correlated with the proportion of the construction impact area (reduced by 0.1 when the proportion of the construction impact area is 0.3, by 0.2 when the proportion of the construction impact area is 0.6, and by 0.3 when the proportion of the construction impact area is 0.9). At the same time, the reduction is adjusted in conjunction with the time decay factor, with the reduction corresponding to recent data being higher than that corresponding to older data (for example, a congested area determined based on data from 10 seconds ago is reduced by 0.1, while a congested area determined based on data from 30 seconds ago is reduced by 0.08). Finally, based on the geographic grid division of the initial environmental semantic feature map, the access priority coefficients of each region are mapped to the corresponding grids, generating a access priority feature layer that corresponds one-to-one with the geographic grids of the initial environmental semantic feature map. Each grid element of this feature layer is the access priority coefficient of the corresponding region, providing priority data support for the subsequent generation of a real-time semantic global map with dynamic access priorities.Preferably, in the specific implementation of step 123, a coordination-update mechanism is constructed for the conflict coordination and dynamic updating of the traffic priority coefficient to ensure the stability and timeliness of the traffic priority feature layer. Specifically, when a certain area is simultaneously identified as a congested area and a construction area, a superimposed reduction rule is adopted. The total reduction is the weighted sum of the individual reductions of the two types of areas. The weights are determined according to the time decay factor (the area type corresponding to recent data has a higher weight). For example, if a certain area is identified as a congested area based on data from 5 seconds ago (reduction of 0.15) and as a construction area based on data from 8 seconds ago (reduction of 0.2), then the total reduction = 0.15 × 0.55 + 0.2 × 0.45 = 0.1725, to avoid repeated reductions that would cause the coefficient to be too low. Simultaneously, a dynamic update mechanism for traffic priority coefficients is constructed. The update cycle is dynamically adjusted according to the rate of change of the regional environment. By calculating the absolute value of the first derivative (rate of change) of standardized time-series environmental data, if the rate of change exceeds a preset threshold (e.g., 0.08 / second), the update cycle is shortened to 1 second; if the rate of change is lower than the preset threshold (e.g., 0.02 / second), the update cycle is extended to 5 seconds, ensuring timeliness while reducing computational overhead. In addition, to avoid frequent fluctuations in coefficients, an update lag threshold is set. When the difference between two adjacent calculated traffic priority coefficients is less than 0.05, no update operation is performed, ensuring the stability of the traffic priority feature layer and providing a stable priority reference for subsequent path planning.

[0024] Preferably, the training process of the regional accessibility assessment model specifically includes: adopting a scenario-adaptive training strategy to construct a training dataset covering various closed or semi-closed scenarios such as industrial parks, science parks, and campuses. The dataset contains time-series environmental data and corresponding manually labeled accessibility quantification values ​​under different time periods (e.g., morning and evening rush hours, off-peak hours), different weather conditions (e.g., sunny days, rainy days), and different dynamic scenarios (e.g., dense pedestrian traffic, construction encroachment, vehicle congestion), with a total sample size of no less than 50,000. Data augmentation processing is performed on the training data, including time-scale stretching (stretching ratio of 0.8 to 1.2), noise addition (Gaussian noise variance of 0.01 to 0.05), and data splicing (sponging time-series data from different scenarios according to reasonable logic) to expand the diversity of the training data and improve the model's generalization ability. A multi-dimensional joint loss function is designed, comprising four parts: regional congestion index loss, construction impact area ratio loss, traffic bandwidth redundancy loss, and fusion loss. The regional congestion index loss, construction impact area ratio loss, and traffic bandwidth redundancy loss all employ the mean squared error loss function to calculate the difference between the model's predicted value and the manually labeled value for each dimension. The fusion loss employs the cross-entropy loss function to calculate the difference between the final traffic quantification value output by the model and the manually labeled value. The total loss of the multi-dimensional joint loss function is the weighted sum of the four losses, with the weight ratios dynamically adjusted according to the training phase (initially, the weight of each dimension loss is 0.25; as the number of training iterations increases, the weights are gradually adjusted to: regional congestion index loss 0.3, construction impact area ratio loss 0.3, traffic bandwidth redundancy loss 0.2, and fusion loss 0.2). The training process employs the Adam optimizer with an initial learning rate of 1e-4 and a weight decay coefficient of 1e-5. A learning rate decay strategy is used, reducing the learning rate to 0.9 times its current value every 10 epochs. The total training iterations are 50 epochs. An early stopping strategy is employed during training, stopping when the prediction error of the mobility quantization value on the validation set fails to decrease for five consecutive epochs. After training, the model's mobility quantization value prediction error (MAE) on the validation set is below 0.05, and the prediction accuracy of each dimension is above 90%, ensuring the model can accurately handle temporal environment data in different scenarios. Preferably, the verification and optimization mechanism of the mobility priority feature layer specifically includes constructing a three-level verification system of "numerical verification - logical verification - actual adaptation verification" to ensure the rationality and practicality of the mobility priority feature layer. Numerical verification checks whether the coefficients are within a reasonable range of 0 to 1 by traversing all grid coefficients of the mobility priority feature layer. If any coefficients exceed this range, they are corrected by tracing back to the previous quantization or model evaluation stage.Logical verification establishes verification rules based on the traffic logic of the park scenario. For example, the traffic priority coefficient for main road areas should be higher than that for pedestrian areas, the traffic priority coefficient for areas without construction or congestion should be no less than 0.7, and the traffic priority coefficient for construction areas should be lower than that for non-construction areas. If these logical rules are violated, the coefficients of the relevant areas are automatically adjusted to conform to the logic. Actual adaptation verification compares the matching degree between the traffic priority feature layer and the robot's historical traffic data, calculating the correlation coefficient between the traffic priority coefficient of each area and the actual traffic efficiency. If the correlation coefficient is lower than a preset threshold (e.g., 0.6), the model evaluation parameters corresponding to that area (such as time decay factor and dimension weight) are adjusted, and the traffic priority coefficient is regenerated to ensure that the traffic priority feature layer can truly reflect the actual traffic efficiency of the area. After three levels of verification, the rationality of the coefficients in the traffic priority feature layer is improved, providing a reliable guarantee for the generation of a real-time semantic global map with dynamic traffic priorities.

[0025] Optionally, step 2 includes: Step 21, based on the geographic topology of the real-time semantic global map with dynamic passage priority, setting multi-dimensional optimization objectives such as path length, number of obstacles to avoid, and passage efficiency, embedding the robot's motion constraints into the path generation logic, and generating multiple initial candidate global paths; Step 22, extracting the visual semantic features and historical motion data of dynamic obstacles from the real-time semantic global map with dynamic passage priority, and associating and binding the visual semantic features of dynamic obstacles with historical motion data to determine the motion trajectory of obstacles in a future preset time period and the corresponding dynamic influence area; Step 23, performing an intersection operation on the coordinate sequence of the initial candidate global paths and the spatial range of the dynamic influence area to generate a candidate global path feature tensor with obstacle avoidance prediction.

[0026] Optionally, step 22 includes: step 221, extracting the visual semantic features of dynamic obstacles and the real-time pixel coordinates, category labels, and displacement from the real-time semantic global map with dynamic traffic priority, and calculating the real-time movement speed and direction of the obstacles; step 222, constructing trajectory-semantic association prediction logic, using the visual semantic features of dynamic obstacles and the visual semantic features in the historical movement data as constraint parameters of the network input gate, and using the visual semantic features of dynamic obstacles and the historical movement trajectory data in the historical movement data as the network temporal input, so as to obtain the movement trajectory of the obstacles in the future preset time period; step 223, converting the pixel coordinates of the predicted trajectory into geographic coordinates, expanding the trajectory spatial range by a preset safety distance, and generating a dynamic influence area.

[0027] Preferably, the specific implementation process of step 221 includes: extracting the visual semantic features and historical motion data of dynamic obstacles from the real-time semantic global map with dynamic traffic priority, and obtaining the real-time motion state parameters of the obstacles through quantization calculation; specifically, firstly, based on the semantic label filtering function of the real-time semantic global map with dynamic traffic priority, identifying and extracting the visual semantic features and historical motion data of objects marked as dynamic obstacles (pedestrians, motor vehicles, non-motor vehicles, construction equipment, etc.) in the map. The visual semantic features include the category label of the dynamic obstacle (such as pedestrians, forklifts, electric vehicles) and the shape and size features (pixel quantization values ​​of length, width, and height). The system extracts several key features, including appearance attributes (such as whether pedestrians are carrying items or vehicles are loaded with goods). These features are extracted through the semantic segmentation and object detection branches of the large-scale visual model. The confidence threshold for the category labels is set to a general range (e.g., above 0.7) to ensure the reliability of feature extraction. Historical motion data includes real-time pixel coordinate sequences and displacement sequences of dynamic obstacles within a preset historical time period (e.g., the most recent 5 seconds). The real-time pixel coordinate sequences are obtained by tracking the position of dynamic obstacles in consecutive frames of images using object tracking algorithms (such as KCF and CSRT algorithms). The displacement sequences are generated by calculating the difference between pixel coordinates of adjacent timestamps. Subsequently, the extracted historical motion data is preprocessed, and a sliding window averaging method is used to eliminate random noise (the window size is 3 to 5 sampling points) to obtain smoothed pixel coordinate sequences and displacement sequences. Based on smoothed historical motion data, the real-time motion speed and direction of dynamic obstacles are calculated: the real-time motion speed is calculated by the average displacement per unit time, in pixels per second (for example, if the total displacement of a pedestrian in 5 seconds is 500 pixels, then the real-time motion speed is 100 pixels per second), and combined with the conversion ratio between map pixels and actual geographical distance (for example, 10 pixels correspond to 1 meter) to convert it into actual speed (in meters per second); the real-time motion direction is determined by calculating the vector angle between adjacent pixel coordinates, with the positive x-axis of the image coordinate system as the reference, clockwise as the positive direction, and the range from 0° to 360°, generating real-time motion state parameters of obstacles that include category labels, dimensions, real-time speed, and real-time direction. Preferably, in a further specific implementation of step 221, a state verification mechanism is constructed to improve the accuracy of motion parameters for the abnormal detection and correction of the motion state of dynamic obstacles. Specifically, based on the motion characteristics of different types of dynamic obstacles in the park scene, a motion state threshold library is constructed. The threshold library includes speed thresholds for different types of obstacles (such as the normal speed range of pedestrians being 0.5 to 2.0 m / s, and the normal speed range of forklifts being 1.0 to 3.0 m / s) and direction change rate thresholds (such as the normal direction change rate of pedestrians being 0° to 30° / s, and the normal direction change rate of vehicles being 0° to 15° / s).The calculated real-time obstacle motion state parameters are compared with the corresponding thresholds in the threshold library. If the real-time speed or the rate of change of direction exceeds the normal range, the motion state is judged as abnormal. For speed abnormalities, if the speed is below the lower threshold limit, it is judged as a stationary state (possibly a temporary pause), and the real-time speed is corrected to 0; if the speed is above the upper threshold limit, it is judged as a data abnormality, and a weighted average of historical speeds is used for correction (recent historical data has higher weight). For abnormal rate of change of direction, it is judged as an obstacle turning or changing lanes, the current direction calculation result is retained, but it is marked as a turning state to provide a state reference for subsequent trajectory prediction. At the same time, the motion state parameters are further verified by combining the shape and appearance attributes in the visual semantic features of dynamic obstacles. For example, the speed of pedestrians carrying large items should be lower than that of pedestrians without items, and the speed of forklifts loaded with goods should be lower than that of empty forklifts. If the verification finds a discrepancy, it is corrected again according to the speed range corresponding to the attribute to ensure that the real-time obstacle motion state parameters can accurately reflect its actual motion.

[0028] Preferably, in the specific technical implementation of step 222, a trajectory-semantic association prediction logic is constructed based on the real-time motion state parameters and visual semantic features of the obstacle to generate the motion trajectory of the obstacle in a future preset time period. Specifically, the trajectory-semantic association prediction logic adopts a three-level architecture of "semantic constraint layer - temporal modeling layer - trajectory generation layer". This architecture is customized for the motion characteristics and semantic constraint relationships of dynamic obstacles in the park scene. The semantic constraint layer converts the visual semantic features of the dynamic obstacle into constraint parameters for trajectory prediction. The category label is converted into a semantic embedding vector (256-dimensional) through the embedding layer. The shape size and appearance attributes are converted into constraint coefficients through the fully connected layer (e.g., the constraint coefficient for pedestrians is 0.8, and the constraint coefficient for forklifts is 1.2). The semantic embedding vector and the constraint coefficient together constitute the constraint parameters of the network input gate, which are used to limit the range and trend of trajectory prediction (e.g., the turning flexibility of pedestrian trajectory is higher than that of forklift, so the constraint coefficient is smaller). The temporal modeling layer employs a designed Transformer-GNN hybrid network. It uses the real-time velocity and direction from the obstacle's real-time motion state parameters, along with smoothed pixel coordinate sequences from historical motion data, as the network's temporal input. The Transformer encoder captures long-term dependencies in the motion data (e.g., the uniform motion trend of pedestrians), while the GNN (Graph Neural Network) module constructs an interaction graph between the dynamic obstacle and its surrounding environment (e.g., road boundaries, other obstacles). Nodes represent elements of the dynamic obstacle and its environment, and edges represent the strength of the interaction. An attention mechanism is used to enhance the influence of key interactions (e.g., the turning tendency increases when a pedestrian approaches the road boundary), generating motion feature vectors containing temporal dependencies and environmental interaction information. The trajectory generation layer uses fully connected layers and activation functions (e.g., ReLU) to map the motion feature vectors to a pixel coordinate sequence for a future preset time period (e.g., 10 to 30 seconds). The length of the preset time period can be dynamically adjusted according to task requirements (shorter preset time periods for emergency tasks, longer preset time periods for inspection tasks), generating the obstacle's future motion trajectory (in pixel coordinate form). Preferably, in a further specific implementation of step 222, a scene-adaptive training mechanism is constructed for the training and optimization of the trajectory-semantic association prediction logic to improve the accuracy of trajectory prediction. Specifically, a dynamic obstacle trajectory dataset for a park scene is constructed. The dataset contains trajectory data of various types of obstacles such as pedestrians, forklifts, electric vehicles, and construction vehicles. Each trajectory data includes visual semantic features (category labels, dimensions, appearance attributes), historical motion data (coordinate sequences of 5 to 10 seconds), and real future trajectory data (coordinate sequences of 10 to 30 seconds), with a total sample size of no less than 100,000. The dataset is preprocessed, including data cleaning (removing abnormal trajectories) and data augmentation (trajectory translation, scaling, and time stretching) to expand the diversity of the dataset.A multi-task joint loss function was designed, comprising trajectory coordinate loss and semantic consistency loss. The trajectory coordinate loss uses the mean squared error loss function to calculate the difference between the predicted and actual trajectory coordinates. The semantic consistency loss uses the cross-entropy loss function to ensure that the predicted trajectory conforms to the motion characteristics of the obstacle's semantic category (e.g., pedestrian trajectories should not exhibit high-speed straight-line motion of vehicles). The AdamW optimizer was used during training, with an initial learning rate of 1e-4 and a weight decay coefficient of 1e-5. A cosine annealing learning rate adjustment strategy was employed, reducing the learning rate to 0.9 times its current value every 10 epochs. The total number of training iterations was 50 epochs, and an early stopping strategy was used (training stopped if the validation set loss did not decrease for 5 consecutive epochs). After training, the model's trajectory prediction error (RMSE) on the validation set was less than 0.5 meters, ensuring high accuracy in generating future obstacle motion trajectories.

[0029] Preferably, the specific implementation process of step 223 is as follows: The pixel coordinates of the future movement trajectory of the obstacle are converted into geographic coordinates, and a dynamic influence area is generated after safety range expansion processing. Specifically, firstly, the geographic coordinate mapping parameters of the real-time semantic global map with dynamic passage priority are obtained. This parameter includes a transformation matrix between the image coordinate system and the geographic coordinate system (such as the WGS84 coordinate system). The rows and columns of the transformation matrix correspond to the dimensions of the image coordinates and the geographic coordinates, respectively, and the matrix elements are transformation coefficients. Based on this transformation matrix, each pixel coordinate of the future movement trajectory of the obstacle is converted into the corresponding geographic coordinate, generating a geographic coordinate sequence of the future trajectory. Subsequently, based on the visual semantic features and real-time movement state parameters of the dynamic obstacle, a preset safety distance is determined: the base value of the safety distance is set according to the obstacle category (e.g., the base safety distance for pedestrians is 0.8 to 1.2 meters, and the base safety distance for forklifts is 1.5 to 2.0 meters), and dynamically adjusted according to the real-time speed (the higher the speed, the larger the safety distance, with an adjustment coefficient of 0.1 to 0.3 meters). (seconds / meter). For example, if a forklift's real-time speed is 2.0 m / s, the basic safety distance is 1.5 meters, and the adjustment factor is 0.2, then the final safety distance is 1.5 + 2.0 × 0.2 = 1.9 meters. Based on the determined safety distance, the spatial range of the geographic coordinate sequence of the future trajectory is expanded: a circular region is constructed with each trajectory point as the center and the safety distance as the radius. The union of all circular regions constitutes the initial influence region. Considering the shape and size of the dynamic obstacle, the initial influence region is shaped and adjusted, replacing the circular regions with rectangular regions that match the shape and size of the obstacle (the length and width of the rectangle are the obstacle's shape and size plus twice the safety distance). The direction of the rectangular region is consistent with the real-time movement direction of the obstacle, generating a dynamic influence region containing the geographic coordinates of the future trajectory, the safety distance, and the expanded range. Preferably, in a further specific implementation of step 223, a region optimization mechanism is constructed to improve the reliability of obstacle avoidance reference, focusing on the timely updating and conflict detection of the dynamic influence region. Specifically, the timeliness of the dynamic influence region is consistent with the preset time period of the obstacle's future movement trajectory. Steps 221 to 223 are re-executed every preset update cycle (e.g., 1 to 3 seconds) to update the dynamic influence region based on the latest dynamic obstacle movement data, ensuring that the region range can reflect the movement changes of the obstacle in real time. Simultaneously, a conflict detection mechanism for the dynamic influence region is constructed. When the dynamic influence regions of multiple dynamic obstacles overlap, the area and duration of the overlapping region are calculated. If the overlapping area exceeds 30% of the area of ​​a single region and the overlapping duration exceeds 2 seconds, it is determined to be a conflict region. The safety distance of the conflict region is adjusted by superposition (superposition ratio of 0.5 to 0.8), expanding the scope of the conflict region and reminding subsequent path planning to prioritize avoidance. Furthermore, by combining the road network and traffic restriction areas in the real-time semantic global map with dynamic traffic priority, the dynamic impact area is cropped. If the dynamic impact area exceeds the road range or contains restricted areas, the part exceeding the road range or the part within the restricted areas is cropped to ensure that the dynamic impact area only contains the potential conflict range within the passable area, providing accurate obstacle avoidance reference for the generation of candidate global path feature tensors with obstacle avoidance prediction.

[0030] Preferably, the model optimization and scene adaptation of the trajectory-semantic association prediction logic are specifically implemented as follows: A region adaptation module is constructed to address the differences in dynamic obstacle motion characteristics across different areas of the park scenario (such as main roads, intersections, and construction areas). This module includes motion pattern features of different regions (e.g., vehicles on main roads tend to move at a more linear and uniform speed, while obstacles at intersections change direction more frequently). The region type of the dynamic obstacle is identified using a large visual model, and the region type features are converted into region embedding vectors. These vectors are then input into the temporal modeling layer of the trajectory-semantic association prediction logic and fused with motion data features, enabling the model to adapt to the motion patterns of different regions. Simultaneously, a transfer learning mechanism is introduced. The basic model is trained using existing trajectory data from the park scenario, and then fine-tuned using a small amount of labeled data from a specific park to improve the model's prediction accuracy in specific scenarios. For example, for an intersection in a technology park, after fine-tuning using 500 obstacle trajectory data points from that intersection, the trajectory prediction error (RMSE) is reduced compared to the basic model, ensuring that the trajectory-semantic association prediction logic can accurately adapt to the dynamic obstacle motion characteristics of different park scenarios.

[0031] Optionally, step 3 includes: Step 31, constructing a five-dimensional path evaluation system, wherein the travel time dimension is the product of path length and regional travel efficiency, the obstacle avoidance difficulty dimension is the intersection ratio of the path and the dynamically affected area, the path smoothness dimension is the mean of the path curvature change rate, the travel priority dimension is the mean of the priority coefficient of the area traversed by the path, and the path-environment adaptability dimension is the spatial feature matching degree between the path topology and the real-time semantic global map with dynamic travel priority. The weights of each dimension are adaptively adjusted based on the fluctuation range of the real-time environmental data; Step 32, inputting the corresponding parameters in the feature tensor of the candidate global path with obstacle avoidance prediction into the five-dimensional path evaluation system, and completing the comprehensive score calculation by weighted summation according to the adaptive weights configured according to the real-time environmental data; Step 33, sorting the optimal path from high to low according to the comprehensive score, identifying curvature change nodes and obstacle avoidance nodes in the path, configuring corresponding dynamic travel instructions such as deceleration and turning for the nodes, and integrating the path topology and dynamic travel instructions to generate a dynamic environment-adaptive global path planning topology. Optionally, step 31 includes: Step 311, quantifying the travel time dimension by multiplying the geographical length of the path by the travel efficiency coefficient of the areas it passes through, where the travel efficiency coefficient is taken from the regional parameters of the real-time semantic global map with dynamic travel priority; Step 312, quantifying the obstacle avoidance difficulty dimension by the proportion of the intersection time of the path and the dynamic influence area to the total path time; Step 313, quantifying the path smoothness dimension by the arithmetic mean of the curvature change rate of all inflection points of the path; Step 314, quantifying the travel priority dimension by the weighted mean of the travel priority coefficients of all areas the path passes through, where the priority coefficients are taken from the real-time semantic global map with dynamic travel priority; Step 315, quantifying the path-environment adaptability dimension by the matching degree between the directional features of the path topology and the directional features of the road network in the real-time semantic global map with dynamic travel priority; Step 316, configuring adaptive weights for each dimension based on the fluctuation range of real-time environmental data, with higher weight coefficients for dimensions with greater fluctuation range of real-time environmental data, and generating a five-dimensional path evaluation system based on the quantification rules and weights of each dimension.

[0032] Preferably, the specific implementation process of step 311 includes: quantifying the travel time dimension based on a real-time semantic global map with dynamic travel priority, and generating a quantified value of the path travel time; specifically, firstly, the geographical length of the candidate path is extracted, and the Euclidean distance between adjacent coordinate points is calculated and accumulated by parsing the coordinate sequence of the path to obtain the total geographical length of the path (in meters). For example, the total length of a candidate path after accumulating the coordinate sequence is 500 meters. Subsequently, the travel efficiency coefficient of all areas traversed by the path is extracted from the real-time semantic global map with dynamic travel priority. The travel efficiency coefficient is positively correlated with the regional travel priority coefficient and is generated by mapping the dynamic travel priority coefficient of each region in the map (the mapping relationship is: travel efficiency coefficient = travel priority coefficient × base speed coefficient, where the base speed coefficient is the standard driving speed of the robot under undisturbed conditions, such as 1.5 m / s). For example, if the travel priority coefficient of a certain region is 0.8, then the corresponding travel efficiency coefficient is 0.8 × 1.5 = 1.2 m / s. Based on the proportion of geographical length of each region traversed by the route, a weighted average traffic efficiency coefficient is calculated (the weight is the proportion of each region's geographical length to the total route length). For example, if the route passes through regions A and B, with region A accounting for 60% of the route length and having a traffic efficiency coefficient of 1.2 m / s, and region B accounting for 40% of the route length and having a traffic efficiency coefficient of 1.0 m / s, then the weighted average traffic efficiency coefficient = 0.6 × 1.2 + 0.4 × 1.0 = 1.12 m / s. The ratio of the total geographical length of the route to the weighted average traffic efficiency coefficient is used as the quantified travel time value (in seconds), i.e., quantified travel time value = total geographical length of the route ÷ weighted average traffic efficiency coefficient. For example, 500 meters ÷ 1.12 meters / s ≈ 446 seconds, generating the quantified travel time value for the route. Preferably, in a further specific implementation of step 311, an efficiency calibration mechanism is constructed to improve the accuracy of travel time quantification for dynamic correction and time estimation compensation of the traffic efficiency coefficient. Specifically, the traffic efficiency coefficient of road segments passing through the dynamic influence area is corrected based on the distribution of the dynamic influence area. If a road segment intersects with the dynamic influence area, the traffic efficiency coefficient is lowered according to the proportion of the intersection length to the length of the road segment. The higher the proportion, the greater the reduction (e.g., a reduction of 0.1 when the intersection length accounts for 30%, and a reduction of 0.2 when it accounts for 50%). The corrected traffic efficiency coefficient = original traffic efficiency coefficient × (1 - intersection length proportion × correction coefficient), and the correction coefficient ranges from 0.3 to 0.5. Meanwhile, considering the time spent by the robot to decelerate and accelerate during obstacle avoidance, the quantified travel time is compensated. The compensation time = number of obstacle avoidance attempts × average time per obstacle avoidance attempt (the average time per obstacle avoidance attempt is set according to the type of obstacle, such as the average time to avoid pedestrians is 2 to 3 seconds, and the average time to avoid vehicles is 3 to 5 seconds). The final quantified travel time = basic quantified travel time + compensation time, ensuring that the quantified obstacle avoidance time can truly reflect the actual travel time of the path.

[0033] Preferably, in the specific technical implementation of step 312, based on the spatiotemporal correlation between the dynamic influence area and the path, the obstacle avoidance difficulty dimension is quantitatively calculated to generate a quantitative value for the path obstacle avoidance difficulty. Specifically, firstly, the dynamic influence area corresponding to each dynamic obstacle is obtained, and the geographical coordinate range and existence duration of the dynamic influence area are extracted (existence duration is the duration of the dynamic obstacle in a future preset time period, such as 30 seconds). The coordinate sequence of the candidate path is spatiotemporally intersected with the geographical coordinate range of the dynamic influence area to determine the length of the intersection segment between the path and each dynamic influence area and the corresponding intersection duration (intersection duration = intersection segment length ÷ traffic efficiency coefficient of the segment). The intersection durations corresponding to all dynamic influence areas are summed to obtain the total intersection duration of the path. The ratio of the total intersection time of the calculated paths to the quantified path travel time is the quantified obstacle avoidance difficulty value (ranging from 0 to 1). For example, if the total intersection time is 80 seconds and the quantified travel time is 446 seconds, then the quantified obstacle avoidance difficulty value is 80 ÷ 446 ≈ 0.18. A higher ratio indicates greater obstacle avoidance difficulty. The path obstacle avoidance difficulty quantification value is generated. Preferably, in a further specific implementation of step 312, a difficulty enhancement mechanism is constructed based on the conflict level classification and obstacle avoidance difficulty weighting of the dynamic impact area to improve the accuracy of obstacle avoidance difficulty quantification. Specifically, based on the type and movement characteristics of dynamic obstacles, a conflict level classification standard is constructed, dividing the dynamic impact area into three conflict levels: low, medium, and high. Low conflict level corresponds to obstacles with slow movement speed and flexible trajectory, such as pedestrians; medium conflict level corresponds to obstacles with medium movement speed and relatively fixed trajectory, such as electric vehicles; and high conflict level corresponds to obstacles with fast movement speed and large size, such as forklifts and construction vehicles. Assign corresponding weight coefficients to different conflict levels (low conflict level weight 1.0, medium conflict level weight 1.5, high conflict level weight 2.0). Calculate the product of the intersection duration and the weight coefficient for each dynamic impact area, sum them to obtain the weighted total intersection duration, and then calculate the ratio of the weighted total intersection duration to the path travel time quantification value as the final obstacle avoidance difficulty quantification value. For example, if the low conflict level intersection duration is 50 seconds and the medium conflict level intersection duration is 30 seconds, then the weighted total intersection duration = 50 × 1.0 + 30 × 1.5 = 95 seconds, and the obstacle avoidance difficulty quantification value = 95 ÷ 446 ≈ 0.21, so that the obstacle avoidance difficulty quantification value can distinguish the difference in obstacle avoidance pressure for different types of obstacles.

[0034] Preferably, the specific implementation process of step 313 includes: extracting and analyzing the inflection point features of the candidate path, quantifying the path smoothness dimension, and generating a quantitative value for path smoothness; specifically, firstly, the Douglas-Peucker algorithm is used to simplify the coordinate sequence of the candidate path using a polyline, retaining the significant inflection points of the path. The inflection point judgment threshold is set according to the robot's kinematic constraints (e.g., the turning angle is greater than 30°). For example, after simplification, a candidate path retains 10 inflection points. The rate of change of curvature at each inflection point is calculated. The rate of change of curvature is determined by the ratio of the curvature difference of the path segment before and after the inflection point to the length of the path segment. The curvature is calculated by the ratio of the tangent angle of the path segment to the length of the path segment. For example, if the curvature of the path segment before an inflection point is 0.1 / m, the curvature of the path segment after the inflection point is 0.3 / m, and the length of the path segment is 10 meters, then the rate of change of curvature at the inflection point = (0.3-0.1)÷10 = 0.02 / m². Calculate the arithmetic mean of the curvature change rates of all inflection points. This mean is the path smoothness quantification value (ranging from 0 to +∞). The smaller the mean, the smoother the path. Preferably, in a further implementation of step 313, a smoothness verification mechanism is constructed to improve the practicality of the smoothness quantification, addressing kinematic adaptation and abnormal inflection point correction for path smoothness. Specifically, combining the robot's kinematic constraints (such as minimum turning radius and maximum steering angular velocity), the curvature change rate of each inflection point is verified. If the curvature change rate of an inflection point exceeds the robot's adaptation range (e.g., the curvature change rate threshold corresponding to the maximum steering angular velocity is 0.05 / m²), it is determined to be an abnormal inflection point, and the curvature change rate of that inflection point is corrected (to the upper limit of the threshold). Simultaneously, a tangent continuity index is introduced for the path. The change in the tangent angle between adjacent inflection points is calculated. If the change is too large (e.g., exceeding 45°), a penalty coefficient is added to the smoothness quantification value (penalty coefficient = change ÷ 90°). The final path smoothness quantification value = mean rate of change of curvature at inflection points × (1 + sum of penalty coefficients). For example, if a path has two adjacent inflection points with a tangent angle change of 60°, the sum of penalty coefficients = (60 ÷ 90) + (60 ÷ 90) = 1.33, and the mean rate of change of curvature at inflection points is 0.02 / m². Then, the final path smoothness quantification value = 0.02 × (1 + 1.33) = 0.0466 / m², so that the path smoothness quantification value can match the actual motion capability of the robot.

[0035] Preferably, in the specific technical implementation of step 314, based on a real-time semantic global map with dynamic traffic priority, the traffic priority dimension is quantitatively calculated to generate a path traffic priority quantification value. Specifically, traffic priority coefficients for all regions traversed by the candidate path are extracted from the real-time semantic global map with dynamic traffic priority. The traffic priority coefficients are dynamic evaluation parameters for each region in the map (ranging from 0 to 1). Based on the geographical length proportion of each region traversed by the path, a weighted average of the traffic priority coefficients is calculated. This weighted average is the traffic priority quantification value (ranging from 0 to 1). For example, if the path traverses three regions with length proportions of 30%, 40%, and 30%, respectively, and the corresponding traffic priority coefficients are 0.9, 0.8, and 0.7, then the traffic priority quantification value = 0.3 × 0.9 + 0.4 × 0.8 + 0.3 × 0.7 = 0.8. A higher value indicates a higher overall traffic priority for the regions traversed by the path, thus generating the path traffic priority quantification value. Preferably, in a further specific implementation of step 314, a priority optimization mechanism is constructed to enhance the guidance of traffic priority quantification by strengthening the traffic priority of key areas and penalizing low-priority areas. Specifically, key areas in the park (such as main roads, emergency passages, and areas surrounding mission objectives) are defined, and a priority strengthening coefficient (ranging from 1.2 to 1.5) is configured for key areas, while a base coefficient of 1.0 is configured for non-key areas. The weighted traffic priority coefficient of each area is calculated as: Area traffic priority coefficient × Corresponding coefficient (the strengthening coefficient for key areas and the base coefficient for non-key areas). At the same time, a penalty coefficient is configured for low-priority areas whose traffic priority coefficient is lower than a preset threshold (e.g., 0.5). The penalty coefficient is 1 + (0.5 - area traffic priority coefficient) × 2. For example, if the traffic priority coefficient of a low-priority area is 0.3, then the penalty coefficient is 1 + (0.5 - 0.3) × 2 = 1.4, and the weighted traffic priority coefficient of that area is 0.3 × 1.4 = 0.42. Based on the adjusted weighted traffic priority coefficient, the weighted average value of the areas traversed by the route is recalculated as the final traffic priority quantification value, so that the traffic priority quantification value can better highlight the importance of key areas and avoid low priority areas.

[0036] Preferably, the specific implementation process of step 315 includes: extracting the directional features of the candidate path topology and road network, quantifying the path-environment adaptability dimension, and generating a quantified value of the path-environment adaptability metric; specifically, firstly, extracting the directional feature vector of the candidate path topology, dividing the path into multiple path segments according to a preset length (e.g., 5 meters), calculating the directional vector of each path segment (based on the north direction of the geographic coordinate system, with clockwise as the positive direction, ranging from 0° to 360°), and combining the directional vectors of all path segments into a path directional feature vector (the dimension being the number of path segments). Subsequently, extracting the road network directional feature vector of the area traversed by the path from the real-time semantic global map with dynamic traffic priority, performing the same segmentation process on the road network (each segment having the same length as the path segment), calculating the directional vector of each road segment, and forming a road network directional feature vector (the dimension being the same as the path directional feature vector). Calculate the cosine similarity between two directional feature vectors. The cosine similarity ranges from -1 to 1. Map it to the interval between 0 and 1 (the mapping formula is: basic fit value = (cosine similarity + 1) ÷ 2). This mapped value is the path-environment fit quantification value. For example, if the cosine similarity is 0.8, then the basic fit value = (0.8 + 1) ÷ 2 = 0.9. The higher the value, the better the path topology matches the road network direction. Generate the path-environment fit quantification value. Preferably, in a further specific implementation of step 315, an adaptation enhancement mechanism is constructed for the topological adaptation and regional functional adaptation of the road network to improve the comprehensiveness of adaptation quantification. Specifically, the topological features of the road network (such as road intersection density and road width variation patterns) and the topological features of the path (such as the number of times the path crosses intersections and path width adaptability) are extracted, and the structural similarity between the two is calculated (within the range of 0 to 1). For example, if the road intersection density is 0.2 intersections / meter and the density corresponding to the number of times the path crosses intersections is 0.18 intersections / meter, then the structural similarity is 0.9. At the same time, based on the task type of the path (such as delivery task, inspection task) and the functional attributes of the areas traversed (such as logistics area, office area), the functional adaptability is calculated (within the range of 0 to 1). For example, the functional adaptability of a delivery task traversing a logistics area is 0.9, and the functional adaptability of a delivery task traversing an office area is 0.7. The adaptation fusion formula is constructed as follows: Path-environment adaptation quantification value = Adaptation base value × 0.5 + Structural similarity × 0.3 + Functional adaptation × 0.2. For example, if the adaptation base value is 0.9, the structural similarity is 0.9, and the functional adaptation is 0.9, then the final path-environment adaptation quantification value = 0.9 × 0.5 + 0.9 × 0.3 + 0.9 × 0.2 = 0.9, so that the path-environment adaptation quantification value can comprehensively reflect the adaptation status in terms of direction, structure, and function.

[0037] Preferably, in the specific technical implementation of step 316, adaptive weights are configured for the five dimensions based on the fluctuation characteristics of real-time environmental data to generate a complete five-dimensional path evaluation system. Specifically, firstly, a fluctuation amplitude calculation index for real-time environmental data is defined. Fluctuation coefficients are calculated for congestion data, construction data, and traffic bandwidth data respectively. The fluctuation coefficient = (maximum value - minimum value) ÷ mean value (range 0 to +∞). For example, if the maximum value of congestion data is 0.8, the minimum value is 0.2, and the mean value is 0.5, then the fluctuation coefficient = (0.8 - 0.2) ÷ 0.5 = 1.2. The mean value of the fluctuation coefficients of all environmental data is calculated as the overall environmental fluctuation amplitude index. The fluctuation level is divided according to this index (low fluctuation level: fluctuation amplitude < 0.5, medium fluctuation level: 0.5 ≤ fluctuation amplitude < 1.5, high fluctuation level: fluctuation amplitude ≥ 1.5). A five-dimensional initial weight matrix is ​​configured for each fluctuation level. At low fluctuation levels, the weights of each dimension are relatively balanced (travel time 0.25, obstacle avoidance difficulty 0.25, path smoothness 0.2, travel priority 0.2, path-environment adaptability 0.1). At medium fluctuation levels, the weights of dynamically relevant dimensions are increased (travel time 0.3, obstacle avoidance difficulty 0.3, path smoothness 0.15, travel priority 0.15, path-environment adaptability 0.1). At high fluctuation levels, the weights of dynamic dimensions are further strengthened (travel time 0.35, obstacle avoidance difficulty 0.35, path smoothness 0.1, travel priority 0.1, path-environment adaptability 0.1). Based on the fluctuation sensitivity of each dimension (dynamic dimensions are more sensitive than static dimensions), the initial weight matrix is ​​fine-tuned. Fluctuation sensitivity = the correlation coefficient between the data of that dimension and environmental fluctuations; the higher the correlation coefficient, the larger the fine-tuning amplitude. Finally, adaptive weights for each dimension are generated, which, together with the five-dimensional quantification rules, constitute a complete five-dimensional path evaluation system. Preferably, in a further specific implementation of step 316, a weight optimization mechanism is constructed for weight adaptation and verification of task types to improve the scenario adaptability of the evaluation system. Specifically, a task type-weight mapping library is constructed, containing weight adjustment rules corresponding to different task types (such as emergency delivery, routine inspection, and security patrol). For example, for emergency delivery tasks, the weight of the passage time dimension is increased (+0.1), and the weight of the path smoothness dimension is decreased (-0.05); for security patrol tasks, the weight of the obstacle avoidance difficulty dimension is increased (+0.1), and the weight of the passage priority dimension is increased (+0.05). According to the current task type of the robot, the adaptive weights are adjusted, and the total weight after adjustment remains at 1.0. At the same time, a weight verification mechanism is constructed to calculate the product of the variance of each dimension's quantified value and the weight. If the product exceeds a preset threshold (e.g., 0.3), it indicates that the weight configuration may lead to an imbalance in the evaluation results. The weights are automatically adjusted to reduce the product to below the threshold, ensuring that the five-dimensional obstacle path evaluation system can adapt to environmental fluctuations and task requirements while maintaining the balance and rationality of the evaluation.

[0038] Preferably, the specific implementation process of step 32 includes: performing parameter parsing and mapping on the candidate global path feature tensor with obstacle avoidance prediction, combining the adaptive weights of the five-dimensional path evaluation system, completing the comprehensive score calculation, and generating a path comprehensive score table; specifically, firstly, performing dimensional parsing on the candidate global path feature tensor with obstacle avoidance prediction. This feature tensor is a three-dimensional matrix, with the first dimension corresponding to the number of candidate paths, the second dimension corresponding to the sampling point sequence of the path, and the third dimension corresponding to the feature parameter type (including path length, passage priority coefficient of the route area, intersection duration with the dynamic influence area, inflection point curvature change rate, path direction vector, etc.), and the matrix elements being the quantized values ​​of the corresponding parameters. For each candidate path, the corresponding parameter data in the feature tensor are extracted, and the parameters are mapped according to the quantification rules of the five-dimensional path evaluation system: the path length and the traffic efficiency coefficient of the traversed area are mapped to the quantified value of travel time; the intersection time with the dynamic influence area and the obstacle conflict level are mapped to the quantified value of obstacle avoidance difficulty; the inflection point curvature change rate sequence is mapped to the quantified value of path smoothness; the traversed area traffic priority coefficient sequence is mapped to the quantified value of traffic priority; and the path direction vector and the road network direction vector are mapped to the quantified value of path-environment adaptation. Obtain the adaptive weights of the five-dimensional path evaluation system (e.g., travel time weight 0.35, obstacle avoidance difficulty weight 0.3, path smoothness weight 0.15, travel priority weight 0.1, path-environment adaptability weight 0.1). Calculate the weighted sum of the five-dimensional quantified values ​​for each candidate path using the following formula: Overall Score = Travel Time Quantified Value × Time Weight + Obstacle Avoidance Difficulty Quantified Value × Obstacle Avoidance Weight + Path Smoothness Quantified Value × Smoothness Weight + Travel Priority Quantified Value × Priority Weight + Path-Environment Adaptability Quantified Value × Adaptability Weight. The obstacle avoidance difficulty quantified value is a negative indicator and must be converted to a positive score before calculation (positive score = 1 - obstacle avoidance difficulty quantified value). For example, if the positive vectorized values ​​of the five dimensions of a candidate path are 0.8, 0.7, 0.9, 0.85, and 0.9, respectively, and the corresponding weights are 0.35, 0.3, 0.15, 0.1, and 0.1, then the comprehensive score = 0.8 × 0.35 + 0.7 × 0.3 + 0.9 × 0.15 + 0.85 × 0.1 + 0.9 × 0.1 = 0.79, a path comprehensive score table containing all candidate path numbers and their corresponding comprehensive scores is generated. Preferably, in a further specific implementation of step 32, a parameter calibration mechanism is constructed to improve the accuracy of the comprehensive score by checking for anomalies in the feature tensor parameters and correcting the scores. Specifically, based on the reasonable range of path parameters in the park scenario, a parameter threshold library is constructed. The threshold library contains the normal value range of each feature parameter (e.g., the normal range of the path smoothness quantification value is 0 to 0.1 / m², and the normal range of the travel time quantification value is 0.8 to 1.5 times the ratio of the corresponding path's geographical length to its maximum speed). The parsed feature parameters are compared with the threshold library. If a parameter exceeds the normal range, it is determined to be an abnormal parameter.For outlier parameters, the process is traced back to the feature tensor generation stage. If the error is due to data acquisition, the average parameter value of similar paths is used for replacement. If the error is due to the characteristics of the path itself (e.g., the path smoothness is low when traversing narrow passages), a correction coefficient is added when calculating the quantization value of the corresponding dimension (e.g., smoothness quantization value correction coefficient = 0.9) to reduce the impact of outlier parameters on the overall score. Simultaneously, a scoring consistency verification mechanism is constructed to calculate the difference in overall scores for the same path at different sampling intervals. If the difference exceeds a preset threshold (e.g., 0.05), the sampling point density is increased, and the feature parameters and overall score are recalculated to ensure the reliability and stability of the scores in the obstacle path overall scoring table.

[0039] Preferably, in the specific technical implementation of step 33, the optimal path is selected based on the comprehensive path scoring table, key nodes are identified and dynamic passage instructions are configured, and a dynamic environment-adaptive global path planning topology is generated. Specifically, the comprehensive scores in the comprehensive path scoring table are first sorted in descending order, and the path with the highest score is selected as the optimal path. If there are multiple paths with scores less than a preset threshold (e.g., 0.03), a secondary screening mechanism is initiated, prioritizing the selection of paths with lower obstacle avoidance difficulty quantification values ​​and higher path-environment adaptability quantification values ​​to ensure that the optimal path has better comprehensive performance. Key nodes are identified on the selected optimal path: the improved Douglas-Peucker algorithm is used to simplify the path coordinate sequence into a polyline, retaining curvature change nodes (turning angles greater than a preset threshold, e.g., 30°) and obstacle avoidance nodes (the closest distance to the dynamic influence area is less than a preset safety distance, e.g., 2 meters). For example, after simplification, an optimal path identifies 8 curvature change nodes and 5 obstacle avoidance nodes, generating a list of key node coordinates and node type labels. Based on node type and surrounding environment information, dynamic passage instructions are configured for each key node: For nodes with abrupt curvature changes, turning and speed limit instructions are configured according to the angle (e.g., a "slow turn" instruction is configured for angles between 30° and 60°, with a speed limit of 0.7 times the normal speed; a "sharp turn" instruction is configured for angles greater than 60°, with a speed limit of 0.5 times the normal speed); for obstacle avoidance nodes, avoidance instructions are configured according to the type and distance of dynamic obstacles (e.g., a "decelerate and avoid pedestrians, maintain a speed of 0.8 m / s" instruction is configured at a distance of 1.5 meters from the dynamic influence area of ​​pedestrians; a "stop and observe before proceeding" instruction is configured at a distance of 2 meters from the dynamic influence area of ​​forklifts). The coordinate sequence of the optimal path, key node information, and corresponding dynamic passage instructions are integrated and encapsulated in a format parsable by the robot navigation system (e.g., structured data containing path point coordinates, node numbers, instruction types, and instruction parameters) to generate a dynamic environment-adaptive global path planning topology. Preferably, in a further specific implementation of step 33, a topology optimization mechanism is constructed to improve the executability of the planned topology, targeting the optimization of instructions for key nodes and the feasibility verification of the planned topology. Specifically, the dynamic passage instructions for key nodes are optimized in conjunction with the robot's kinematic constraints (such as minimum turning radius and maximum deceleration): For nodes with sudden curvature changes, the required steering advance is calculated based on the robot's minimum turning radius, and the trigger position of the steering instruction is adjusted to a preset distance in front of the node (such as 1.5 times the minimum turning radius) to ensure that the robot has enough space to complete the turn; for obstacle avoidance nodes, the deceleration advance is calculated based on the robot's braking performance (such as the braking distance calculated based on the maximum deceleration and the current speed), and the trigger position of the deceleration instruction is set at the braking distance in front of the obstacle avoidance node to avoid the risk of collision due to untimely braking.Simultaneously, a feasibility verification mechanism for the planned topology is constructed. The robot's kinematic model is used to simulate and extrapolate the planned topology, simulating the entire process of the robot traveling along the planned topology to verify whether the path exceeds kinematic constraints or whether the command triggering timing is unreasonable. If the trajectory tracking error exceeds a preset threshold (e.g., 0.3 meters) or the command action cannot be completed during the simulation, the command parameters of the corresponding nodes are adjusted (e.g., increasing the steering lead or reducing the speed limit) or the path coordinate sequence is optimized, and the planned topology is regenerated until the simulation passes, ensuring that the obstacle dynamic environment-adaptive global path planning topology can be accurately executed by the robot.

[0040] Optionally, step 4 includes: step 41, collecting full-process data of driving along the dynamic environment-adaptive global path planning topology; step 42, normalizing the collected full-process driving data, and calculating the path space execution deviation rate, node passage timeliness deviation rate, obstacle avoidance response accuracy, and regional priority adaptation deviation value respectively; step 43, based on the visual semantic-planning parameter linkage mechanism, mapping the regional priority adaptation deviation value to the passage priority coefficient correction amount of the real-time semantic global map with dynamic passage priority, mapping the path space execution deviation rate to the weight ratio adjustment amount of the multi-dimensional optimization target, and mapping the obstacle avoidance response accuracy to the feature association parameter correction amount of the trajectory-semantic association prediction logic.

[0041] Preferably, the specific implementation process of step 41 includes: based on the execution requirements of the dynamic environment-adaptive global path planning topology, constructing a multi-source collaborative acquisition system to collect multi-dimensional data of the entire robot driving process and generate a structured driving dataset; specifically, firstly, clarifying the type and frequency of the acquired data, the data types cover four categories: position and posture data, node passage data, obstacle avoidance data, and area passage status data, and the acquisition frequency is dynamically adjusted according to the data type (the acquisition frequency of position and posture data is 10 to 20 Hz, the node passage data and obstacle avoidance data are event-triggered acquisitions, and the area passage status data is 1 to 5 Hz). Position and attitude data are collected through a positioning module that integrates laser SLAM and RTK-GPS, including parameters such as the robot's real-time geographic coordinates, heading angle, speed, and acceleration. The geographic coordinate accuracy is controlled within a general range (e.g., ±0.05 meters), providing a basis for calculating deviations in path space execution. Node passage data is collected collaboratively by the robot's odometer and positioning module, including parameters such as arrival time, passage time, and actual speed of key nodes (curvature change nodes, obstacle avoidance nodes). Each key node corresponds to a data record, labeled with the node number and node type. Obstacle avoidance data is collected collaboratively by visual sensors and LiDAR, including parameters such as the actual position of dynamic obstacles, the robot's avoidance actions (deceleration, turning, stopping), avoidance trigger time, avoidance completion time, and avoidance distance, forming obstacle avoidance event data. Area passage status data is collected by connecting to the park management system and the robot's own sensor data, including parameters such as the actual congestion density, construction status changes, and actual passage bandwidth values ​​of each area traversed by the path, which are compared with the area passage priority coefficient during planning. All collected data are aligned by timestamp and structured and encapsulated in the format of "data type-collection time-parameter value-data source" to generate a structured driving dataset containing four types of data: location, node, avoidance, and region, ensuring data integrity and traceability. Preferably, in a further implementation of step 41, a data quality assurance mechanism is constructed for the synchronous calibration and anomaly filtering of multi-source data to improve the reliability of the structured driving dataset. Specifically, a timestamp synchronization algorithm is used to calibrate the multi-source data. Using the robot's master clock (error ≤ 1ms) as a benchmark, the timestamps of data collected by each sensor are uniformly corrected to the master clock time. The timestamp deviation after calibration is controlled within 5ms to avoid subsequent calculation errors due to data asynchrony. Simultaneously, a data anomaly filtering rule base is constructed. The rule base includes the normal value range of each data type (e.g., the normal range of driving speed is 0 to 1.2 times the robot's maximum speed, and the normal range of acceleration is -2 to 2m / s²) and anomaly judgment conditions (e.g., if three consecutive sampling points exceed the normal range, it is judged as an anomaly).Anomaly detection is performed on each data point in the structured driving dataset. If anomalies are detected, different processing methods are adopted according to the data type: for anomalies in position and attitude data, linear interpolation of the preceding and following sampling points is used for replacement; for anomalies in node passage data and obstacle avoidance data, they are marked as invalid events and not included in subsequent index calculations; for anomalies in area passage status data, the statistical mean of historical data from the same period in the same area is used for replacement. In addition, a data integrity verification mechanism is constructed to check whether the structured driving dataset contains the complete process data of path execution. If data is missing (such as the passage data of a key node is missing), a supplementary data collection mechanism is triggered to supplement the data through the robot's local cache or backup data from the park management system, ensuring the integrity and accuracy of the structured driving dataset.

[0042] Preferably, in the specific technical implementation of step 42, the structured driving dataset is normalized to eliminate dimensional differences and generate standardized driving data. Specifically, firstly, differentiated normalization methods are designed based on the distribution characteristics of different types of data: For continuous data such as driving speed and acceleration in position and attitude data, the Min-Max normalization method is used to map the data to the interval between 0 and 1. The mapping formula is: standardized value = (original value - minimum value) / (maximum value - minimum value), where the maximum and minimum values ​​are taken from the robot's kinematic constraints (e.g., the maximum speed is the robot's maximum design speed, and the minimum value is 0); For time-related data in node passage data, the Z-Score normalization method is used to convert the data into values ​​that conform to a standard normal distribution (mean is 0, standard deviation is 1), eliminating the influence of differences in basic passage time between different nodes; For distance, angle, and other data in obstacle avoidance data, the interval scaling normalization method is used to map the data according to the actual environmental range of the park scene (e.g., the normal range of avoidance distance is 0 to 5 meters), ensuring that the standardized data can reflect relative size relationships. During the normalization process, normalization parameters (such as minimum, maximum, mean, and standard deviation) for each data type are recorded for subsequent result restoration and analysis after quantitative index calculation. Simultaneously, a normalization consistency verification mechanism is constructed to calculate the similarity of the distribution trends of data of the same type before and after normalization. If the similarity is lower than a preset threshold (e.g., 0.9), the normalization method is adjusted and reprocessed to ensure that the normalization process does not distort the original distribution characteristics of the data, generating standardized driving data with eliminated dimensional differences and a reasonable distribution. Preferably, in a further specific implementation of step 42, based on standardized driving data, four quantitative indicators are calculated: path spatial execution deviation rate, node passage timeliness deviation rate, obstacle avoidance response accuracy, and regional priority adaptation deviation value, to generate a path execution quantitative evaluation report. Specifically, the path spatial execution deviation rate is calculated as follows: extract the real-time geographic coordinate sequence from the standardized driving data and the preset coordinate sequence of the dynamic environment-adaptive global path planning topology, calculate the Euclidean distance (i.e., spatial deviation) of each sampling point, and smooth the spatial deviation sequence using a sliding window averaging method (window size of 5 to 10 sampling points) to obtain a smoothed spatial deviation sequence. The ratio of the mean of this sequence to the total path length is used as the path spatial execution deviation rate. For example, if the total path length is 500 meters and the mean of the smoothed spatial deviation is 1.25 meters, then the path spatial execution deviation rate = 1.25 ÷ 500 = 0.0025. The smaller the ratio, the higher the spatial accuracy of the path execution.

[0043] Optionally, step 43 includes: Step 431, based on the sign and absolute value of the regional priority adaptation deviation value calculated from the full-process driving data, performing increase or decrease correction on the traffic priority coefficient of the corresponding region in the real-time semantic global map with dynamic traffic priority; the larger the absolute value of the deviation value in the full-process driving data, the higher the correction magnitude; Step 432, based on the deviation rate and node traffic efficiency deviation rate of the path space calculated from the full-process driving data, adjusting the weight ratio of path smoothness and traffic efficiency dimensions in the multi-dimensional optimization target; the higher the deviation rate in the full-process driving data, the higher the corresponding dimension weight coefficient; Step 433, based on the obstacle avoidance response accuracy calculated from the full-process driving data, correcting the constraint weight of visual semantic features in the trajectory-semantic association prediction logic; the lower the response accuracy in the full-process driving data, the stronger the constraint of semantic features on trajectory prediction.

[0044] Preferably, the specific implementation process of step 431 includes: based on the characteristics of the regional priority adaptation deviation value, constructing a hierarchical correction rule, and adjusting the traffic priority coefficient of the real-time semantic global map with dynamic traffic priority to generate an updated dynamic semantic global map; specifically, firstly, defining the hierarchical standard of the regional priority adaptation deviation value, dividing it into three levels according to the absolute value of the deviation value: slight deviation (0 to 0.1), moderate deviation (0.1 to 0.3), and severe deviation (greater than 0.3), each level corresponding to a different correction magnitude coefficient (the correction magnitude coefficient for slight deviation is 0.1 to 0.2, for moderate deviation it is 0.2 to 0.4, and for severe deviation it is 0.4 to 0.6). If the deviation value is positive, it indicates that the actual traffic efficiency is higher than the planned traffic efficiency coefficient, and the traffic priority coefficient of the corresponding region needs to be increased, with the increase = deviation value × correction magnitude coefficient; if the deviation value is negative, it indicates that the actual traffic efficiency is lower than the planned traffic efficiency coefficient, and the traffic priority coefficient of the corresponding region needs to be decreased, with the decrease = |deviation value| × correction magnitude coefficient. For example, if the regional priority adaptation deviation of a certain area is 0.25 (moderate positive deviation) and the correction factor is 0.3, then the upward adjustment = 0.25 × 0.3 = 0.075. If the original traffic priority factor is 0.8, then the corrected value is 0.875. If the deviation of a certain area is -0.35 (severe negative deviation) and the correction factor is 0.5, then the downward adjustment = 0.35 × 0.5 = 0.175. The original factor of 0.7 is corrected to 0.525. At the same time, upper and lower limits (0 to 1) are set for the factor correction to avoid the corrected factor from exceeding a reasonable range. After the correction is completed, the corresponding regional parameters of the real-time semantic global map with dynamic traffic priority are updated to generate the updated dynamic semantic global map. Preferably, in a further specific implementation of step 431, a scenario-based optimization mechanism is constructed to coordinate the correction of regional types and deviation trends, thereby improving the adaptability of the traffic priority coefficient. Specifically, the correction magnitude coefficient is adjusted according to the regional type (main road, secondary road, construction area, densely populated area, etc.). For example, the correction magnitude coefficient for main road areas is increased by 20% based on the corresponding deviation level, while that for construction areas is decreased by 20%, ensuring that the correction is more in line with the functional characteristics of the area. At the same time, the changing trend of historical deviation data is analyzed. If a region experiences multiple consecutive deviations in the same direction (e.g., three consecutive positive deviations), a trend strengthening coefficient (0.1 to 0.2) is added to the current correction magnitude to accelerate the adjustment of the coefficient towards the adaptation direction. If the deviation direction changes alternately, a trend smoothing coefficient (0.8 to 0.9) is added to slow down the correction magnitude and avoid frequent fluctuations in the coefficient. For example, if a main road area has three consecutive positive deviations, the current correction magnitude coefficient is 0.3, and the additional trend reinforcement coefficient is 0.15, then the actual correction magnitude coefficient = 0.3 × (1 + 0.2) × (1 + 0.15) = 0.414, making the coefficient correction more in line with the long-term traffic characteristics of the area.

[0045] Preferably, in the specific technical implementation of step 432, a dynamic weight adaptation rule is constructed based on the path space execution deviation rate and the node passage timeliness deviation rate to adjust the weight ratio of multi-dimensional optimization targets and generate an updated optimization target weight configuration. Specifically, the mapping relationship between deviation rate and weight adjustment amount is first defined. The higher the path space execution deviation rate, the greater the weight adjustment amount of the path smoothness dimension (the mapping ratio is deviation rate × 0.5 to deviation rate × 0.8); the higher the node passage timeliness deviation rate, the greater the weight adjustment amount of the passage efficiency dimension (the mapping ratio is deviation rate × 0.4 to deviation rate × 0.7). The weight adjustment must follow the principle of "total conservation", that is, the increase in the weight of a certain dimension is equal to the sum of the decrease in the weight of other dimensions, and the weights of relatively secondary dimensions such as path length and obstacle avoidance difficulty are preferentially reduced. For example, if the path space execution deviation rate of a certain path is 0.04 and the node passage time deviation rate is 0.06, the path smoothness weight adjustment amount = 0.04 × 0.6 = 0.024, the passage efficiency weight adjustment amount = 0.06 × 0.5 = 0.03, and the total improvement amount = 0.054. The path length weight (originally 0.3) is reduced by 0.03, the obstacle avoidance difficulty weight (originally 0.3) is reduced by 0.024, and the adjusted path smoothness weight = 0.15 + 0.024 = 0.174, the passage efficiency weight = 0.2 + 0.03 = 0.23. The weights of other dimensions are adjusted accordingly to generate the updated optimization target weight configuration. Preferably, in a further specific implementation of step 432, a task-oriented optimization mechanism is constructed to improve the scenario adaptability of multi-objective planning by adjusting the weights for different task types. Specifically, a task type-weight adjustment rule library is constructed, including weight bias rules corresponding to different task types such as emergency delivery, routine inspection, and security patrol. For example, in an emergency delivery task, the weight adjustment amount for the traffic efficiency dimension is increased by 1.2 times, and the weight adjustment amount for the path length dimension is increased by 0.9 times. In a security patrol task, the weight adjustment amount for the obstacle avoidance difficulty dimension is increased by 1.3 times, and the weight adjustment amount for the path smoothness dimension is increased by 1.1 times. Based on the current task type of the robot, the weight configuration after the initial adjustment is corrected a second time to ensure that the weight ratio is both adapted to the deviation rate characteristics and in line with the core requirements of the task. At the same time, a weight rationality verification mechanism is constructed to calculate the variance of the weights of each dimension after adjustment. If the variance exceeds a preset threshold (e.g., 0.01), it indicates that the weight distribution is unbalanced, and the weights of each dimension are automatically fine-tuned to reduce the variance below the threshold, ensuring that the optimized target weight configuration after obstacle update is both targeted and balanced.

[0046] Preferably, the specific implementation process of step 433 includes: constructing constraint weight enhancement rules based on obstacle avoidance response accuracy, correcting the constraint weights of visual semantic features in the trajectory-semantic association prediction logic, and generating optimized trajectory prediction logic; specifically, firstly, a grading standard for obstacle avoidance response accuracy is defined, divided into three levels: high response accuracy (greater than 1.5), medium response accuracy (0.8 to 1.5), and low response accuracy (less than 0.8), each level corresponding to a different constraint weight adjustment coefficient (the adjustment coefficient for high response accuracy is 0.9 to 1.0, for medium response accuracy it is 1.0 to 1.2, and for low response accuracy it is 1.2 to 1.5). The initial value of the constraint weights of visual semantic features is set according to the obstacle type (pedestrian 0.7, vehicle 0.8, construction equipment 0.9). The corrected constraint weight = initial weight × adjustment coefficient. The lower the response accuracy, the larger the adjustment coefficient, and the stronger the constraint weight. For example, if the response accuracy of an obstacle avoidance event is 0.7 (low response accuracy), the corresponding obstacle type is a pedestrian, the initial constraint weight is 0.7, and the adjustment coefficient is 1.3, then the corrected constraint weight = 0.7 × 1.3 = 0.91, strengthening the constraint effect of visual semantic features (such as pedestrian shape and movement posture) on trajectory prediction. Simultaneously, the maximum value of the constraint weight is set to 1.0 to avoid over-constraint leading to rigid prediction, generating an optimized trajectory-semantic association prediction logic. Preferably, in a further specific implementation of step 433, the constraint weights are subdivided according to obstacle type and scene characteristics to construct a refined enhancement mechanism to improve the accuracy of trajectory prediction; specifically, visual semantic features are subdivided into three categories: shape features, behavioral features, and environmental association features, and basic constraint weights are configured for each (shape features 0.3, behavioral features 0.4, environmental association features 0.3), adjusting the constraint weights of the corresponding features according to the weakest dimension of obstacle avoidance response accuracy. For example, in a low-response accuracy event, analysis revealed insufficient prediction of a pedestrian suddenly crossing the road. Therefore, the constraint weight adjustment coefficient for behavioral features was increased (by 0.2). If the delay in avoidance was due to the failure to recognize the shape features of construction equipment, the constraint weight adjustment coefficient for shape features was increased. Simultaneously, the constraint weight of environmental features was adjusted based on scene characteristics (such as intersections and narrow passages). The more complex the scene, the higher the constraint weight of environmental features, allowing the trajectory-semantic association prediction logic to specifically strengthen weak links. The optimized trajectory prediction error is lower than before the correction.

[0047] The above are merely preferred embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A global path planning method based on a visual large model, characterized in that, include: Step 1: Obtain the global environment image and target location information of the target scene to generate a real-time semantic global map with dynamic access priority; Step 2: Based on the real-time semantic global map with dynamic passage priority, generate a candidate global path feature tensor with obstacle avoidance prediction; Step 3: Construct a five-dimensional path evaluation system including travel time, obstacle avoidance difficulty, path smoothness, travel priority, and path-environment adaptability, so as to generate a dynamic environment-adaptive global path planning topology based on the feature tensor of the candidate global path with obstacle avoidance prediction. Step 4: Map the entire process data of the robot traveling along the dynamic environment-adaptive global path planning topology to a real-time semantic global map with dynamic traffic priority for global path planning adaptation. Step 1 includes: Step 11: Input the global environment image of the target scene into the multi-scale semantic fusion branch of the visual large model, extract the shallow spatial contour features and deep category semantic features of the image respectively, perform cross-scale fusion on the two types of features, and output the semantic labels, spatial coordinates and category confidence of roads, obstacles and traffic restriction areas to generate the initial environment semantic feature map. Step 12: Collect time-series environmental data of the target scenario, including regional congestion duration sequence, temporary construction scope change data, and traffic bandwidth fluctuation data. Construct a regional trafficability assessment model with a time decay factor. Input the time-series environmental data of the target scenario into the model to obtain the trafficability quantification value of each region. Based on the quantification value, configure the traffic priority coefficient for each region in the initial environmental semantic feature map to generate a traffic priority feature layer. Step 13: Bind the geographic grid of the initial environmental semantic feature map to the coefficient grid of the access priority feature layer one by one, perform spatial topology verification on the bound grid, and generate a real-time semantic global map with dynamic access priority.

2. The visual large model based global path planning method according to claim 1, wherein, Step 11 includes: Step 111: Perform block segmentation and scale normalization preprocessing on the global environment image of the target scene to generate a multi-scale image block sequence; Step 112: Input the multi-scale image block sequence into the multi-scale semantic fusion branch of the large visual model. Each block extracts spatial contour features through shallow convolution, extracts category semantic features through deep Transformer module, and completes cross-scale feature association through feature fusion layer. Step 113: Integrate all block fusion features corresponding to the multi-scale image block sequence, eliminate semantic faults at block boundaries, and generate an initial environmental semantic feature map.

3. The global path planning method based on a large visual model according to claim 1, characterized in that, Step 12 includes: Step 121: Construct a regional trafficability assessment model. The model input is the time-series environmental data of the target scenario. The assessment dimensions include regional congestion index, construction impact range ratio, and traffic bandwidth redundancy. Each dimension is configured with a time decay factor to weaken the weight of historical data. Step 122: Collect real-time congestion data, construction area range change data, and road cross-section bandwidth data from the time-series environmental data of the target scenario, and complete data quantification and normalization according to the evaluation dimensions. Step 123: Input the quantified data into the regional accessibility assessment model, output the accessibility quantification value of each region, assign the accessibility priority coefficient to each region in the initial environmental semantic feature map according to the quantification value range, and adjust the coefficients of congested areas and construction areas according to the rules corresponding to the time decay factor to generate the accessibility priority feature layer.

4. The global path planning method based on a large visual model according to claim 1, characterized in that, Step 2 includes: Step 21: Based on the geographic topology of the real-time semantic global map with dynamic passage priority, set multi-dimensional optimization objectives such as path length, number of obstacles to avoid, and passage efficiency, embed the robot's motion constraints into the path generation logic, and generate multiple initial candidate global paths. Step 22: Extract the visual semantic features and historical motion data of dynamic obstacles from the real-time semantic global map with dynamic passage priority, and associate and bind the visual semantic features and historical motion data of dynamic obstacles to determine the motion trajectory of the obstacles in the future preset time period and the dynamic influence area corresponding to the trajectory. Step 23: Perform an intersection operation between the coordinate sequence of the initial candidate global path and the spatial range of the dynamic influence area to generate a candidate global path feature tensor with obstacle avoidance prediction.

5. The global path planning method based on a large visual model according to claim 4, characterized in that, Step 22 includes: Step 221: Extract the visual semantic features of dynamic obstacles from the real-time semantic global map with dynamic passage priority and the real-time pixel coordinates, category labels and displacement from the historical motion data, and calculate the real-time movement speed and direction of the obstacles. Step 222: Construct trajectory-semantic association prediction logic, using the visual semantic features of dynamic obstacles and the visual semantic features in historical motion data as constraint parameters of the network input gate, and using the visual semantic features of dynamic obstacles and the historical motion trajectory data in historical motion data as the network temporal input, so as to obtain the motion trajectory of the obstacle in the future preset time period. Step 223: Convert the pixel coordinates of the predicted trajectory into geographic coordinates, expand the spatial range of the trajectory by a preset safety distance, and generate a dynamic influence area.

6. The global path planning method based on a large visual model according to claim 1, characterized in that, Step 3 includes: Step 31: Construct a five-dimensional path evaluation system. The travel time dimension is the product of path length and regional travel efficiency; the obstacle avoidance difficulty dimension is the intersection ratio of the path and the dynamically affected area; the path smoothness dimension is the average of the path curvature change rate; the travel priority dimension is the average priority coefficient of the areas traversed by the path; and the path-environment adaptability dimension is the spatial feature matching degree between the path topology and the real-time semantic global map with dynamic travel priority. The weights of each dimension are adaptively adjusted based on the fluctuation range of real-time environmental data. Step 32: Input the corresponding parameters in the candidate global path feature tensor with obstacle avoidance prediction into the five-dimensional path evaluation system, and complete the comprehensive score calculation by weighted summation according to the adaptive weights configured by real-time environmental data. Step 33: Sort the optimal path from high to low according to the comprehensive score, identify the curvature change nodes and obstacle avoidance nodes in the path, configure the corresponding deceleration and turning dynamic passage instructions for the nodes, and integrate the path topology and dynamic passage instructions to generate a dynamic environment-adaptive global path planning topology.

7. The global path planning method based on a large visual model according to claim 6, characterized in that, Step 31 includes: Step 311: For the travel time dimension, the product of the geographical length of the path and the travel efficiency coefficient of the area it passes through is used for quantification. The travel efficiency coefficient is taken from the regional parameters of the real-time semantic global map with dynamic travel priority. Step 312: For the obstacle avoidance difficulty dimension, the ratio of the intersection time of the path and the dynamic influence area to the total path time is used for quantification. Step 313: For the path smoothness dimension, quantify it using the arithmetic mean of the rate of curvature change of all inflection points of the path; Step 314: For the passage priority dimension, the weighted average of the passage priority coefficients of all areas traversed by the path is used for quantification. The priority coefficients are taken from the real-time semantic global map with dynamic passage priority. Step 315: For the path-environment adaptability dimension, the matching degree between the directional features of the path topology and the directional features of the road network in the real-time semantic global map with dynamic traffic priority is quantified. Step 316: Configure adaptive weights for each dimension based on the fluctuation range of real-time environmental data. The greater the fluctuation range of real-time environmental data, the higher the weight coefficient of the dimension. Generate a five-dimensional path evaluation system based on the quantification rules and weights of each dimension.

8. The global path planning method based on a large visual model according to claim 1, characterized in that, Step 4 includes: Step 41: Collect full-process data of driving along the dynamic environment-adaptive global path planning topology; Step 42: Normalize the collected full-process driving data and calculate the path space execution deviation rate, node passage time deviation rate, obstacle avoidance response accuracy, and area priority adaptation deviation value respectively. Step 43: Based on the visual semantic-planning parameter linkage mechanism, map the regional priority adaptation deviation value to the traffic priority coefficient correction amount of the real-time semantic global map with dynamic traffic priority, map the path space execution deviation rate to the weight ratio adjustment amount of the multi-dimensional optimization target, and map the obstacle avoidance response accuracy to the feature association parameter correction amount of the trajectory-semantic association prediction logic.

9. The global path planning method based on a large visual model according to claim 8, characterized in that, Step 43 includes: Step 431: Based on the sign and absolute value of the regional priority adaptation deviation value calculated from the full-process driving data, perform increase or decrease correction on the traffic priority coefficient of the corresponding region in the real-time semantic global map with dynamic traffic priority. The larger the absolute value of the deviation value in the full-process driving data, the higher the correction magnitude. Step 432: Based on the path space execution deviation rate and node passage time deviation rate calculated from the full-process driving data, adjust the weight ratio of path smoothness and passage efficiency dimensions in the multi-dimensional optimization objectives. The higher the deviation rate in the full-process driving data, the higher the corresponding dimension weight coefficient. Step 433: Based on the obstacle avoidance response accuracy calculated from the full-process driving data, correct the constraint weight of visual semantic features in the trajectory-semantic association prediction logic. The lower the response accuracy in the full-process driving data, the stronger the constraint of semantic features on trajectory prediction.

Citation Information

Patent Citations

  • Obstacle avoidance and planning cooperative path generation method in urban complex environment

    CN120972975A