A full-path guidance system for container trucks outside automated container terminals
Through multimodal image acquisition and deep learning feature extraction, multi-sensor data fusion and intelligent path optimization algorithm, the positioning accuracy and path planning problems of the container truck guidance system outside the automated container terminal are solved, and efficient and safe path guidance is achieved.
Patent Information
- Application Number
- CN202510838215.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing container truck guidance system outside the automated container terminal has insufficient multi-sensor fusion positioning accuracy, lacks the ability to adapt to dynamic environments and intelligent path optimization mechanism, and cannot meet the needs of centimeter-level precise path tracking and real-time traffic adjustment.
An on-board multimodal image acquisition unit is used to obtain lane line and interactive work area information, combined with a ResNet-50 convolutional neural network for feature extraction, and position fusion is performed through an adaptive Kalman filter. A Markov chain model is used to predict traffic flow, and an improved ant colony algorithm is used for path replanning to generate a detour node sequence to avoid congestion.
It achieves centimeter-level positioning accuracy and dynamic path optimization in complex dock environments, improves dock traffic efficiency and safety, and adapts to the complex and dynamic environment of the dock.
Smart Images

Figure CN120355056B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a full-path guidance system for container trucks outside an automated container terminal. Background Art
[0002] Automated container terminals, a crucial component of modern port logistics, utilize automated equipment and intelligent management and control systems to achieve efficient container loading, unloading, and transportation. Existing technologies primarily rely on GPS positioning systems and preset static routes for guidance after entering automated terminals. Basic traffic scheduling and operational arrangements are performed through the Terminal Operating System (TOS). Current guidance technologies for container trucks typically utilize magnetic pin-embedded positioning methods or GPS navigation alone, combined with fixed traffic signals and manual control.
[0003] However, existing container truck guidance technology suffers from significant technical drawbacks. Traditional magnetic positioning systems are expensive to build and difficult to maintain. Positioning accuracy is susceptible to interference in complex terminal environments. GPS single-point positioning is also inaccurate in complex port environments, failing to meet the requirements for centimeter-level precise path tracking. Furthermore, existing technology lacks the ability to adapt to dynamic environments, relying primarily on pre-set static path planning. This makes it impossible to dynamically adjust to real-time traffic flow and congestion conditions, forcing container trucks to wait for manual intervention when encountering temporary obstacles or traffic jams.
[0004] Existing guidance systems lack the real-time positioning capabilities of multi-sensor fusion and dynamic decision-making mechanisms based on artificial intelligence. Traditional methods fail to organically integrate visual perception, location fusion, traffic flow prediction, and path optimization, lacking the intelligent perception and adaptive response capabilities to the complex and dynamic terminal environment. In particular, existing technologies struggle to achieve full-path intelligent guidance when handling deviation correction for external container trucks, real-time obstacle avoidance, and multi-vehicle collaborative optimization, impacting the overall operational efficiency and safety of automated terminals. Summary of the Invention
[0005] The present application provides a full-path guidance system for container trucks outside container automated terminals, which is used to solve the technical problems of insufficient multi-sensor fusion positioning accuracy, lack of dynamic environment adaptability and intelligent path optimization mechanism in existing container truck guidance systems outside container automated terminals.
[0006] The present application provides a full-path guidance system for external container trucks at an automated container terminal, which includes: an acquisition module for performing real-time image acquisition processing on the driving path of the external container truck terminal through a vehicle-mounted multimodal image acquisition unit to generate lane line coordinate data and interactive operation area identification position information; an extraction module for inputting the lane line coordinate data into a ResNet-50 convolutional neural network for feature extraction processing, outputting a guide marking feature vector and identifying the coordinates of a path reference point; a fusion module for performing external container truck position fusion calculation processing through an adaptive Kalman filter based on the path reference point coordinates and the guide marking feature vector, and calculating the lateral deviation amount and the heading angle deviation degree; a prediction module for inputting the historical passage record of the external container truck in combination with the interactive operation area identification position information into a Markov chain model for arrival time prediction processing, predicting the transfer probability of each functional area and calculating the future queue length; a planning module for performing path replanning processing through an improved ant colony algorithm based on the transfer probability and the lateral deviation amount, generating a detour node sequence for avoiding congestion based on the heading angle deviation degree and outputting a steering angle correction instruction.
[0007] In the technical solution provided by the present application, real-time image acquisition and processing of the driving path of the external container truck terminal is realized through the on-board multimodal image acquisition unit, and lane line coordinate data and interactive work area identification position information are generated. Compared with the traditional single sensor positioning method, the multimodal image fusion technology significantly improves the perception accuracy and environmental adaptability in the complex environment of the terminal. By inputting the lane line coordinate data into the ResNet-50 convolutional neural network for feature extraction processing, outputting the guide marking feature vector and identifying the coordinates of the path reference point, the deep learning feature extraction algorithm can effectively handle the lighting changes, weather interference and occlusion problems in the terminal environment, and has stronger robustness and accuracy than the traditional template matching method. According to the path reference point coordinates and the guide marking feature vector, the external container truck position fusion calculation and processing are performed through adaptive Kalman filtering to calculate the lateral deviation and heading angle offset. The multi-sensor data fusion technology achieves centimeter-level positioning accuracy, solving the problem of insufficient accuracy of single GPS positioning in the port environment. The Markov chain model combines historical traffic records of external container trucks with the location information of interactive operation area markers to predict arrival times, predict the transition probabilities of each functional area, and calculate future queue lengths. This probabilistic and statistically based traffic flow prediction model accurately predicts terminal traffic conditions, providing a scientific basis for dynamic route planning. An improved ant colony algorithm is used to replan routes based on transition probabilities and lateral deviations. A sequence of detour nodes is generated based on heading angle deviations, and steering angle correction instructions are output. The intelligent route optimization algorithm enables dynamic congestion avoidance and real-time route adjustment for external container trucks, significantly improving terminal traffic efficiency and safety.
[0008] In the specific application field of full-path guidance for container trucks at automated container terminals, the technical solution of this application fully leverages the domain-specific advantages of each core algorithm. In the lane line recognition task in a terminal environment, the residual learning mechanism of the ResNet-50 convolutional neural network can effectively extract the deep semantic features of the terminal road markings, significantly improving the recognition accuracy under complex lighting and occlusion conditions compared to traditional image processing methods. The adaptive Kalman filter algorithm optimizes parameters based on the motion characteristics of large container trucks. When integrating GPS and visual positioning data, it can dynamically adjust the fusion weights according to the terminal environment to achieve the optimal position estimation effect. The Markov chain model makes full use of the regular characteristics of the passage of container trucks at the terminal. Through three-state transition probability modeling, it accurately captures the flow pattern of container trucks between the gate area, buffer zone and interactive operation area. The prediction accuracy is significantly improved compared to simple statistical methods. The improved ant colony algorithm combines the topological characteristics of the terminal road network and the motion constraints of container trucks. Through dynamic adjustment of pheromones and multi-objective optimization, it realizes the rapid search of congestion-avoiding paths. The algorithm convergence speed and solution quality are both adapted to the needs of real-time terminal scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are some embodiments of the present invention. Those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0010] Figure 1 This is a schematic diagram of an embodiment of a full-path guidance system for container trucks outside an automated container terminal in an embodiment of the present application. DETAILED DESCRIPTION
[0011] An embodiment of the present application provides a full-path guidance system for container trucks outside a container terminal automation terminal. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0012] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, an embodiment of the full-path guidance system for container trucks outside the container automated terminal includes:
[0013] The acquisition module 101 is used to acquire and process real-time images of the driving path of the external container terminal through the vehicle-mounted multimodal image acquisition unit to generate lane line coordinate data and interactive operation area identification position information;
[0014] Extraction module 102, configured to input the lane line coordinate data into a ResNet-50 convolutional neural network for feature extraction, output guide line feature vectors, and identify path reference point coordinates;
[0015] A fusion module 103 is configured to perform a fusion calculation process on the position of the external container truck based on the coordinates of the path reference point and the guide line feature vector through an adaptive Kalman filter to calculate the lateral deviation and the heading angle offset;
[0016] Prediction module 104, for inputting the historical passage records of external container trucks into the Markov chain model in combination with the interactive operation area identification location information to perform arrival time prediction processing, predict the transfer probability of each functional area and calculate the future queue length;
[0017] The planning module 105 is configured to perform path replanning processing using an improved ant colony algorithm according to the transition probability and the lateral deviation, generate a detour node sequence based on the heading angle offset, and output a steering angle correction instruction.
[0018] It is understandable that the execution subject of this application can be the container terminal external container truck full path guidance system, or it can be a terminal or a server, and the specific details are not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0019] Specifically, acquisition module 101 uses an onboard multimodal image acquisition unit to achieve real-time visual perception of the external container terminal path. This module activates a CCD camera to continuously scan the terminal pavement, identifying lane boundaries based on pixel-level grayscale value differences. A Canny edge detection operator calculates the gradient magnitude and direction at each pixel, marking it as a boundary point when the gradient magnitude exceeds a set threshold. This generates a sequence of lane boundary pixel coordinates. A Hough transform algorithm converts these discrete pixels into line parameters. The slope and intercept of the lane are determined by searching for peaks in the parameter space, generating standardized lane coordinate data. Simultaneously, an infrared thermal imager captures the temperature distribution of the interactive work area. The area is identified using the thermal radiation differences between different materials and equipment. A temperature threshold of 30-40 degrees Celsius is set to segment the concrete floor and metal equipment areas, identifying the boundary contour coordinates of the interactive work area. A morphological closing operation, using dilation followed by erosion, fills in holes and breaks in the contour to ensure a complete and continuous region boundary and generate the interactive work area marker location information.
[0020] Extraction module 102 inputs lane line coordinate data into a ResNet-50 convolutional neural network for deep feature learning. The input coordinate data is normalized, scaled to a range of 0-1, and rearranged into a standard 224×224 pixel format. The first convolutional block extracts features from the input data using a 7×7 convolution kernel. The convolution kernel slides across the image by two pixels at a time, calculating the weighted sum of the local region as the output feature value, generating a 112×112 pixel primary feature map. The batch normalization layer calculates the mean and variance of the current batch of data, normalizing each feature value by subtracting the mean and dividing by the standard deviation to prevent vanishing and exploding gradients. Four residual blocks process the feature data sequentially. Each residual block consists of three convolutional layers and a skip connection. The skip connection directly adds the input features to the convolution output, preserving the original information while learning residual features. Global average pooling compresses feature maps of different scales into a fixed-length 2048-dimensional vector. This dimensionality is reduced by averaging all pixels in each feature map. The fully connected layer performs a linear transformation on this 2048-dimensional feature vector. The output of each neuron is equal to the dot product of the input vector and the weight vector plus the bias term. The output is converted into a probability distribution through the softmax activation function. The position with the highest probability value corresponds to the coordinates of the path reference point.
[0021] Fusion module 103 achieves optimal fusion of multi-sensor data through adaptive Kalman filtering, constructing a six-dimensional state vector containing the external truck's position coordinates, velocity components, and heading angle. This state is then combined with GPS positioning data to establish an initial state. The state prediction phase predicts the next moment's position and attitude based on the vehicle's motion model. The prediction equation considers the external truck's turning radius and speed constraints, and outputs the predicted state vector and corresponding covariance matrix. The observation update phase converts the guideline feature vector into observation data, calculating the residual between the predicted and actual values. The Kalman gain determines the fusion weight based on the ratio of the prediction error covariance to the observation noise covariance. An adaptive mechanism dynamically adjusts the noise covariance parameter based on GPS signal strength and the confidence level of visual observations, increasing the weight of visual observations when the GPS signal is weak and increasing the weight of GPS data when the signal is weak. The state update process fuses the predicted and observed values using a weighted average to obtain the optimal position estimate. The deviation calculation performs geometric operations on the fused precise position and the guidance path centerline. The lateral deviation is equal to the perpendicular distance between the external truck's current position and the path centerline, and the heading offset is equal to the angle between the vehicle's heading and the path tangent.
[0022] Prediction module 104 uses a Markov chain model to predict and analyze the arrival times and traffic flows of external container trucks. The terminal is divided into three state spaces: the gate area, the buffer zone, and the interactive operation area. Historical traffic records show the dwell time and number of transfers of external container trucks in each area. State transition probabilities are calculated using maximum likelihood estimation. The probability of a transition from the gate area to the buffer zone is equal to the number of such transfers divided by the total number of departures from the gate area. Transition probabilities between other states are calculated similarly, forming a 3×3 transition probability matrix. Poisson distribution modeling models the arrival time intervals of external container trucks as a random process. The arrival rate parameter is estimated based on historical arrival frequencies during different time periods. The arrival rate is higher during the morning rush hour and lower during the night. The Markov chain state equation predicts the state distribution at each future moment through matrix multiplication. The current state vector is multiplied by the power of the transition probability matrix to obtain the future state probability distribution. Queue length prediction calculates the average queue length based on the vehicle arrival rate and service rate in each area. When the arrival rate approaches the upper limit of the service rate, the queue length increases sharply, triggering a congestion warning.
[0023] Planning module 105 uses an improved ant colony algorithm for dynamic path optimization. The port road network is modeled as a weighted directed graph, where nodes represent intersections and key locations, and edges represent road segments. Edge weights take into account distance, travel time, and congestion levels. The weight calculation formula takes the weighted sum of distance cost, time cost, and congestion cost. The congestion cost is dynamically adjusted based on the predicted transition probability, with higher weights corresponding to higher congestion probabilities. During the initialization phase of the ant colony algorithm, pheromones are evenly distributed across all road segments. The pheromone concentration is adjusted based on the amount of lateral deviation. This concentration is reduced for segments with significant deviations, guiding the virtual ants to avoid paths with significant deviations. Fifty virtual ants simultaneously search the road network for paths. Each ant selects the next node based on pheromone concentration and heuristic information, with the selection probability proportional to the pheromone concentration and path quality. During the path evaluation phase, a total cost function is calculated for each candidate path, which incorporates path length, expected travel time, and number of turns. An additional penalty term is added to segments with heading angle deviations exceeding a threshold. Pheromone updating adjusts the pheromone concentration of each path segment based on path quality, increasing pheromone concentration on high-quality paths and evaporating pheromone concentration on low-quality paths. After multiple rounds of iteration, the system converges to the optimal solution. Steering control is based on the node sequence of the optimal path. It uses geometric methods to calculate the azimuth angle difference between adjacent nodes, determines the steering angle and turning radius based on the vehicle's kinematic constraints, and generates precise steering angle correction instructions to guide the truck along the optimized path.
[0024] In a specific embodiment, the acquisition module 101 is configured to:
[0025] The CCD camera is used to continuously scan and photograph the dock pavement, and edge detection is performed based on pixel grayscale value analysis to extract the lane boundary pixel coordinate sequence;
[0026] Inputting the lane line boundary pixel coordinate sequence into the Hough transform algorithm for straight line fitting processing, calculating the lane line slope and intercept parameters and generating lane line coordinate data;
[0027] Use an infrared thermal imager to scan the temperature distribution of the interactive operation area, perform area segmentation based on the temperature difference threshold, and identify the boundary contour coordinates of the interactive operation area;
[0028] Performing morphological closing operation filtering on the boundary contour coordinates of the interactive operation area to eliminate noise interference and extract the complete area contour to generate interactive operation area identification position information;
[0029] The lane line coordinate data and the interactive operation area identification position information are subjected to coordinate system unification processing, a terminal global coordinate reference system is established, and path environment data in a unified coordinate format is output.
[0030] Specifically, acquisition module 101 uses a CCD camera to continuously capture images of the terminal pavement. The camera scans the road ahead of the external container truck at a frequency of 30 frames per second. The resulting raw image contains pavement texture, lane markings, and various interference information. Edge detection uses the Sobel operator to calculate the gradient of each pixel. The operator calculates the grayscale difference in the horizontal and vertical directions. When the grayscale difference between a pixel and its adjacent pixels exceeds a set threshold, it is marked as an edge point. Lane markings, as high-contrast linear features on the road surface, are prioritized for identification. During pixel grayscale value analysis, the grayscale values of white lane markings typically range from 200-255, while those of dark pavement range from 50-100. This significant grayscale difference enables the edge detection algorithm to accurately locate lane boundaries. The extracted lane boundary pixel coordinate sequence contains the coordinate information of all edge points on the left and right lane lines, with each coordinate point recording its horizontal and vertical coordinate position in the image coordinate system. The Hough transform algorithm converts these discrete pixel coordinates into straight lines in parameter space. The algorithm accumulates votes for each possible line in the parameter space, and the parameter combination with the highest votes corresponds to the actual lane line. Line fitting uses the least squares method to calculate the slope and intercept parameters of the lane line. A negative slope for the left lane indicates a slope from the upper left to the lower right, while a positive slope for the right lane indicates a slope from the lower left to the upper right. The intercept parameter determines the intersection of the lane line with the image boundary. The calculated slope and intercept parameters are converted to lane line coordinate data in the actual road coordinate system through coordinate transformation, forming standardized road geometry information. Infrared thermal imagers measure temperature distribution by detecting infrared radiation emitted by target objects. The interactive operation area includes objects of different materials, such as concrete floors, metal gantry cranes, and containers, which have significant differences in thermal radiation characteristics. During the temperature distribution scan, the temperature of the concrete floor is generally close to the ambient temperature. Metal equipment has good thermal conductivity and changes temperature rapidly. Containers, as large metal structures, heat up significantly when exposed to sunlight. The region segmentation process divides the thermal image into different temperature zones based on a temperature difference threshold. The set temperature threshold is dynamically adjusted according to the season and time of day, with the summer threshold set at 35 degrees Celsius and the winter threshold set at 25 degrees Celsius. The temperature difference threshold segmentation algorithm traverses each pixel in the thermal image, marking pixels with temperatures above the threshold as equipment areas and pixels with temperatures below the threshold as ground areas, forming a binary region segmentation image. The boundary coordinates of the interactive work area are obtained using a contour extraction algorithm. The algorithm tracks continuous boundary pixels along the region boundary, recording the coordinate position of each boundary point to form a closed contour coordinate sequence.
[0031] Morphological closing filtering removes noise and corrects the shape of the extracted interactive work area boundary. The closing operation consists of two steps: dilation followed by erosion. The dilation operation expands the contour using a structuring element, a 3×3 square kernel. During the dilation process, the eight neighboring pixels surrounding each contour pixel are marked as contour points, thus connecting previously broken contour segments. The erosion operation shrinks the expanded contour using the same structuring element. During the erosion process, a pixel retains the contour label only if its eight neighboring pixels belong to the contour; all other pixels have their contour labels removed. Noise interference is eliminated through connected component analysis. The algorithm calculates the area of each connected region. Connected regions with an area less than a set threshold are considered noise points and removed from the contour. Complete area contour extraction ensures that the boundary of the interactive work area is continuous and closed. The algorithm checks the connectivity of the contour and connects any broken contours using linear interpolation to generate location information that precisely describes the location and shape of the interactive work area.
[0032] Coordinate system unification resolves coordinate inconsistencies between sensor data. Lane line coordinate data is derived from the image coordinate system of the CCD camera, while interactive work area marker location information is derived from the temperature image coordinate system of the infrared thermal imager. These two coordinate systems differ in their origin locations, coordinate axis orientations, and scale ratios. This unification establishes a sensor calibration relationship. A calibration plate is placed within the terminal environment to generate a transformation matrix between the two coordinate systems. This transformation matrix includes rotation, translation, and scaling parameters. The coordinate transformation calculation converts pixel coordinates in the image coordinate system into physical coordinates in the actual road coordinate system, taking into account parameters such as the camera's installation height, pitch angle, and azimuth. A global coordinate reference system is established with the terminal gate as its origin, with the horizontal axis pointing toward the terminal's interior and the vertical axis pointing transversely. All sensor data is uniformly converted to this global coordinate system. Path environment data in a unified coordinate format includes lane line start and end coordinates, lane width information, interactive work area vertex coordinates, and area information. The data is stored in a standard JSON format, facilitating data retrieval and processing by subsequent modules.
[0033] An example illustrates the complete operation of the acquisition module: After a container truck enters the terminal, a CCD camera begins scanning the road ahead, capturing an image of both lanes. The left lane line consists of a series of pixel coordinates, including point A (100 pixels x 200 pixels y), point B (150 pixels x 300 pixels y), and so on, totaling 200 boundary points. The edge detection algorithm calculates the gradient of each pixel and finds that the gradient of the lane line area is significantly higher than that of the road surface area, successfully extracting the pixel coordinate sequence that defines the boundary between the left and right lanes. The Hough transform algorithm counts the votes for lines in parameter space and finds that the line with a slope of 0.8 and an intercept of 50 receives the highest number of votes, corresponding to the actual left lane line. The right lane line has a slope of 0.7 and an intercept of 400. Simultaneously, an infrared thermal imager scans the interactive operation area, detecting a gantry crane temperature of 40°C and a ground temperature of 30°C. Using a threshold of 35°C, region segmentation is performed, identifying the gantry crane boundary contour as consisting of eight key vertices. Morphological closing operations revealed two breakpoints in the original contour. Dilation connected the breakpoints, and erosion removed any excess protrusions, resulting in a complete rectangular contour. Coordinate normalization converted the image coordinates to road coordinates. The road coordinates for the left lane's starting point are 2 meters east by 0 meters south, and the end point is 50 meters east by 8 meters south. The road coordinates for the center of the interactive work area are 30 meters east by 15 meters south, covering an area of 400 square meters.
[0034] In a specific embodiment, the extraction module 102 includes:
[0035] a processing unit, configured to input the lane line coordinate data into an input layer of a ResNet-50 network for data preprocessing, normalize the coordinate data, and convert it into a standard input format of 224×224 pixels;
[0036] an operation unit, configured to perform a 7×7 convolution operation on the data in the standard input format through a first convolution block, set a step size of 2, and perform batch normalization processing to extract a primary feature map;
[0037] A generation unit is used to sequentially input the primary feature map into four residual blocks for deep feature extraction processing, where each residual block contains three convolutional layers and a skip connection structure to generate a multi-scale deep feature map;
[0038] A dimensionality reduction unit is used to perform global average pooling dimensionality reduction processing on the multi-scale depth feature map, compress the feature map into a 2048-dimensional feature vector and generate a guide line feature vector;
[0039] The transformation unit is used to perform linear transformation processing on the guide mark feature vector through a fully connected layer, calculate the probability distribution of each pixel point belonging to the path reference point through a softmax activation function, identify the highest probability position and output the coordinates of the path reference point.
[0040] Specifically, the processing unit performs data preprocessing after receiving lane line coordinate data transmitted by the acquisition module. The original format of the lane line coordinate data consists of an irregular sequence of pixel coordinate points, each of which records its horizontal and vertical position in the image. Normalization converts the original coordinate values from the pixel scale to a standardized numerical range. Specifically, each coordinate value is subtracted from the minimum value of that dimension and then divided by the numerical range of that dimension, ensuring that all coordinate values are between 0 and 1. Data format conversion rearranges the one-dimensional coordinate sequence into a two-dimensional matrix structure with both rows and columns set to 224, forming a standard square input format. When the number of original coordinate points is insufficient, vacant positions are filled with zeros. When the number of coordinate points exceeds the limit, key points are selected through evenly spaced sampling, generating 224×224 pixel standard input format data. In this standard input format, each matrix element represents the feature strength at the corresponding location. Element values are 1 at locations where lane lines pass, and 0 at background areas, forming a binary feature representation.
[0041] The computational unit performs preliminary feature extraction on the standard input data through the first convolution block. The 7×7 convolution kernel contains 49 weight parameters. As the convolution kernel slides over the input data, it calculates the sum of the products of all elements in the coverage area and the corresponding weights. A stride of 2 means that the convolution kernel moves 2 pixels at a time, resulting in an output size of 112×112 for a 224×224 input. During the convolution operation, the value at each output position is equal to the dot product of all input pixels in the kernel coverage area with the kernel weights, plus a bias term to obtain the eigenvalue. Batch normalization calculates the mean and variance of each feature channel for all samples in the batch. The mean is then subtracted from each eigenvalue, divided by the standard deviation, multiplied by a learnable scaling parameter, and added to a learnable offset parameter. This normalization operation addresses the problem of internal covariate shift during deep network training, making network training more stable and efficient. The primary feature map contains 64 feature channels, each of which captures different feature patterns of the input data, such as basic geometric features such as horizontal edges, vertical edges, and diagonal edges.
[0042] The generation unit sequentially feeds the primary feature map into four residual blocks for deep feature learning. Each residual block employs a skip connection structure to address the vanishing gradient problem in deep networks. The first residual block contains three convolutional layers. The first layer uses a 1×1 convolution kernel for channel dimension compression, the second layer uses a 3×3 convolution kernel for spatial feature extraction, and the third layer uses a 1×1 convolution kernel for channel dimension restoration. Skip connections directly add the input of the residual block to the output of the third convolution layer. This allows the network to learn a residual mapping rather than a direct mapping, simplifying learning. The four residual blocks process feature information at different scales. The first residual block focuses on local details, while subsequent residual blocks gradually expand the receptive field to incorporate broader contextual information. During deep feature extraction, the spatial resolution of the feature map decreases while the number of channels increases, resulting in a multi-scale deep feature map that combines low-level edge and texture features with high-level semantic features. This multi-scale feature map simultaneously captures both the local geometry and global spatial layout of lane lines, laying the foundation for subsequent path reference point localization.
[0043] The dimensionality reduction unit performs global average pooling on the multi-scale deep feature maps, compressing the spatial feature information into a fixed-length feature vector. Global average pooling calculates the average value of each feature channel across the entire spatial range. Assuming the spatial dimensions of a feature channel are height H and width W, the pooling result for that channel is equal to the arithmetic mean of all H×W elements in that channel. Dimensionality reduction converts the original spatially structured feature map into a one-dimensional feature vector with a fixed length of 2048 dimensions, where each dimension represents the intensity value of a high-level semantic feature. The guide marking feature vector contains all the key feature information of the input lane data. These features have been transformed from raw pixel-level information into a high-level semantic representation through layers of abstraction in the deep network. Elements of different dimensions in the feature vector correspond to different visual patterns, such as straight line features, curve features, and intersection features. These features, combined, fully describe the geometric and semantic properties of the lane lines.
[0044] The transformation unit performs a linear transformation on the guide marking feature vector using a fully connected layer. This layer contains 2048×224×224 weight parameters, and the value of each output position is equal to the inner product of the input feature vector and the corresponding weight vector. The linear transformation maps the high-dimensional feature vector back to the spatial coordinate domain, with the output size being 224×224, consistent with the spatial resolution of the original input. The softmax activation function probabilistically normalizes the output of the linear transformation, calculating the probability of each spatial position belonging to a path reference point, with the sum of all probabilities equal to 1. During the probability distribution calculation, the softmax function calculates the exponential function of the output value at each position and then divides the exponential value at each position by the sum of all exponential values to obtain the probability value for each position. Path reference points are identified by finding the maximum value in the probability distribution. The location with the highest probability corresponds to a key feature point of the lane line, typically located at the center or a turning point of the lane line. The coordinates of the location with the highest probability are converted from normalized coordinates to actual image coordinates through an inverse coordinate transformation. The output path reference point coordinates contain the precise horizontal and vertical coordinates of the point in the image.
[0045] In a specific embodiment, the computing unit is configured to:
[0046] Perform a 7×7 convolution kernel sliding window scan on the standard input format data, moving 2 pixel positions each time and calculating the convolution value to generate a 112×112 pixel convolution output matrix;
[0047] Input the convolution output matrix into the batch normalization layer for data normalization, calculate the batch mean and variance, perform normalization transformation, and output a standardized feature matrix;
[0048] Performing a nonlinear transformation of the ReLU activation function on the standardized feature matrix, setting negative values to zero and keeping positive values unchanged, to generate an activated feature map;
[0049] Downsample the activated feature map through a 3×3 max pooling layer, set the step size to 2 and select the maximum value in the pooling window, and output a pooled feature map of 56×56 pixels;
[0050] The pooled feature map is reorganized in channel dimension, the depth dimension of the feature map is adjusted to 64 channels while maintaining the spatial resolution, and a primary feature map is generated.
[0051] Specifically, the computation unit applies a 7×7 convolution kernel sliding window to standard input data. The kernel scans row by row, starting from the upper left corner of the input data. The convolution value at each position is calculated by summing the element-by-element product of the kernel weight and the corresponding input pixel value. During the sliding window scan, the kernel shifts rightward by two pixels at a time. When it reaches the end of a row, it shifts down by two pixels and returns to the beginning of the row to continue scanning. This shifting strategy, with a stride of two, reduces the output size of the 224×224 input data to 112×112 after scanning. The convolution value calculation involves multiplying 49 weight parameters with the corresponding 49 input pixel values. All these products are accumulated and then added with a bias parameter to produce the final convolution output value at that position. During the scanning process, the value at each output position reflects the characteristic strength of the input data in that area. The convolution value is higher at the lane edge and lower at flat road surfaces. The resulting 112×112 pixel convolution output matrix captures the local characteristic information of the input data. The batch normalization layer receives the convolutional output matrix and calculates statistics for each feature channel for all samples in the current batch. Data normalization requires calculating the batch mean and batch variance separately. The batch mean is calculated by summing the values at the corresponding position of all samples in the current batch and dividing it by the number of samples. The batch variance is calculated by summing the squared differences between the values at the corresponding position and the batch mean, and then dividing it by the number of samples. Normalization transforms each eigenvalue by subtracting the batch mean and dividing it by the square root of the batch variance. This transform adjusts the eigenvalue distribution to a standard normal distribution with a mean close to zero and a variance close to one. The normalized values are then multiplied by a learnable scaling parameter gamma and added to a learnable offset parameter beta. These two parameters allow the network to adaptively adjust the scale and position of the feature distribution during training. The output normalized feature matrix maintains the same spatial size of 112×112 as the input, but the value distribution is more stable, which facilitates the training and convergence of subsequent network layers.
[0052] The ReLU activation function performs a nonlinear transformation on the standardized feature matrix. The calculation rule of this function is to keep the original value unchanged when the input value is greater than zero, and set it to zero when the input value is less than or equal to zero. The nonlinear transformation process traverses each element in the standardized feature matrix. The negative value zeroing operation eliminates the negative feature response, and only retains the positive feature response for subsequent processing. The property of positive values remaining unchanged ensures that the effective feature information is fully preserved. The nonlinear characteristics of the ReLU function introduce nonlinear modeling capabilities to deep networks, enabling the network to learn complex feature mapping relationships. After activation, the zero-value area in the feature map corresponds to the location where the feature is not obvious in the input data, and the non-zero-value area corresponds to the location where the feature is significant. This sparse feature representation reduces the computational complexity and highlights the key feature information.
[0053] The max pooling layer downsamples the activated feature map, with the 3×3 pooling window covering nine adjacent pixel positions at a time as it slides across the feature map. Maximum selection within the pooling window compares these nine pixel values and selects the maximum value as the pooled output for that position. This maximum selection retains the strongest feature response in the local area and suppresses weaker noise information. A stride of 2 causes the pooling window to move two pixels at a time. This downsampling method compresses the 112×112 input feature map to a 56×56 output size, reducing the spatial resolution by half while retaining the most important feature information. Downsampling reduces the amount of data required for subsequent calculations and expands the network's receptive field, enabling subsequent network layers to capture a wider range of contextual information. Each pixel position in the pooled feature map represents a feature summary of a larger area in the original input.
[0054] The channel dimension reorganization process adjusts the depth structure of the pooled feature map. The number of channels in the original pooled output is related to the channel configuration of the input data. The reorganization process normalizes the number of channels to 64 feature channels. The depth dimension adjustment is achieved by increasing or decreasing the number of feature channels. When the original number of channels is less than 64, the number of channels is increased by copying existing channels or zero-padding. When the original number of channels is greater than 64, the number of channels is reduced by channel selection or linear combination. Spatial resolution preservation ensures that the reorganized feature map maintains a spatial size of 56×56. Each spatial location contains a 64-dimensional feature vector describing various characteristic attributes of that location. The primary feature map contains the basic visual features extracted from the input lane line data through convolution, normalization, activation, and pooling.
[0055] In a specific embodiment, the fusion module 103 is configured to:
[0056] The coordinates of the path reference point and the current GPS position data of the external container truck are used to construct a state vector for initialization processing, and the position, speed and heading angle of the external container truck in the terminal coordinate system are set as state variables to generate the motion state vector of the external container truck;
[0057] Establishing a vehicle motion prediction model in a terminal environment based on the motion state vector of the external container truck to perform state prediction processing, taking into account the motion characteristics of the external container truck turning and going straight in the terminal, and outputting predicted position coordinates and prediction error covariance;
[0058] The guide marking feature vector is converted into observation data of the outer truck relative to the lane centerline for observation update processing, and the observation residual of the outer truck deviating from the guide path is calculated to generate a Kalman gain coefficient;
[0059] Based on the Kalman gain coefficient, the predicted position of the external container truck is adaptively weighted and fused, and the fusion weight is dynamically adjusted in combination with the GPS positioning accuracy and the visual observation confidence to output the positioning result of the external container truck in the terminal;
[0060] Deviation analysis is performed based on the positioning result of the external container truck and the center line of the terminal guidance path, and the vertical distance of the external container truck from the lane center line and the angle between the vehicle head direction and the path direction are calculated to generate the lateral deviation amount and heading angle offset degree.
[0061] Specifically, the fusion module 103 receives the path reference point coordinates output by the extraction module and the current position data of the container truck acquired by the onboard GPS receiver. The state vector construction process converts these two types of spatial position information into a unified data structure. Initialization processing converts GPS latitude and longitude coordinates into plane coordinates in the terminal's local coordinate system through a coordinate system transformation. The terminal coordinate system is based on the gate location as the origin, with east as the positive X-axis, north as the positive Y-axis, and elevation as the positive Z-axis. The state variable set includes six dimensions: the container truck's position coordinates, velocity components, and heading angle. The position coordinates record the X and Y coordinates of the container truck in the terminal coordinate system, the velocity components record the instantaneous velocity of the container truck in the X and Y directions, and the heading angle records the angle of the container truck's head relative to the X-axis of the coordinate system. The container truck's motion state vector organizes these six state variables into a column vector, with the first and second elements representing the position coordinates, the third and fourth elements representing the velocity components, the fifth element representing the heading angle, and the sixth element representing the angular velocity, forming a mathematical representation of the complete motion state of the container truck. The vehicle motion prediction model establishes mathematical equations of motion based on the state vectors of the external container truck. The state prediction process considers the motion constraints of the external container truck as a large vehicle in the terminal environment. The motion characteristics of the terminal environment include physical parameters such as turning radius constraints, maximum acceleration constraints, and angular velocity constraints. When turning, the external container truck requires a large turning radius to avoid collisions with roadside equipment. When moving straight, the external container truck maintains a constant speed or uniform acceleration. The motion prediction model uses discretized state transition equations. The position coordinate at the next moment is equal to the current position plus the product of the velocity and the time interval. The velocity at the next moment is equal to the current velocity plus the product of the acceleration and the time interval. The heading angle at the next moment is equal to the current heading angle plus the product of the angular velocity and the time interval. The predicted position coordinates are calculated using the motion equations to determine the expected position of the external container truck at the next time step. The prediction error covariance describes the uncertainty of the prediction results. The diagonal elements of the covariance matrix represent the prediction variance of each state variable, while the off-diagonal elements represent the correlation between different state variables.
[0062] The observation update process converts the guide marking feature vector output by the extraction module into observation information about the relative position of the external card. The guide marking feature vector contains the position and orientation of the lane centerline. A geometric transformation is used to calculate the relative relationship between the current position of the external card and the lane centerline. The observation data conversion process extracts key geometric parameters from the guide marking feature vector, including the lane centerline equation parameters and the position of the external card in the image. A coordinate transformation is then used to convert the image coordinates into physical distances and angles in the road coordinate system. The observation residual is calculated by comparing the predicted external card position with the observed actual position. Each element of the residual vector represents the prediction error in different dimensions. The Kalman gain coefficient is determined based on the relative magnitude of the prediction error covariance and the observation noise covariance. When the prediction error is large, the Kalman gain increases, relying more on the observation data for state updates. When the observation noise is large, the Kalman gain decreases, relying more on the prediction result.
[0063] Adaptive weighted fusion processing dynamically adjusts fusion weights based on GPS positioning accuracy and visual observation confidence. GPS positioning accuracy is assessed by satellite signal strength and geometric distribution factors. High signal strength and uniform satellite distribution indicate high GPS accuracy, while low accuracy indicates low accuracy. Visual observation confidence is assessed by image quality and feature matching. Confidence is high when the image is clear and lane markings are distinct, while confidence is low when the image is blurry or severely obscured. A weight adjustment mechanism dynamically assigns fusion weights based on the real-time performance of the two sensors. When GPS accuracy exceeds visual observation, the weight of GPS data is increased, while when visual observation quality is superior to GPS, the weight of visual data is increased. The fusion calculation combines the predicted and observed states through a weighted average. The final state estimate is equal to the sum of the predicted state multiplied by the Kalman gain and the observation residual. The output represents the precise positioning of the outer card within the terminal, containing the optimal estimates of position coordinates, velocity, and heading angle.
[0064] Deviation analysis geometrically compares the fused precise positioning results of the external truck with the predefined terminal guidance path. The centerline of the guidance path is defined by a series of control points and connecting curves, describing the ideal trajectory the truck should follow. Vertical distance calculation is performed using a point-to-line distance formula or a point-to-curve distance algorithm. The algorithm finds the point on the guidance path closest to the external truck's current position and then calculates the vertical distance from that point as the lateral deviation. The angle between the truck's heading and the path direction is calculated using a vector angle formula. The heading angle of the external truck represents the heading direction, and the tangent direction of the guidance path at the nearest point represents the desired travel direction. The angle between these two direction vectors is the heading angle deviation. A positive lateral deviation indicates that the external truck is deviating to the right of the path, while a negative value indicates that it is deviating to the left. A positive heading angle deviation indicates that the truck is heading to the right, while a negative value indicates that it is heading to the left. These two deviation parameters provide critical feedback for the subsequent path replanning algorithm.
[0065] In one embodiment, the prediction module 104 is configured to:
[0066] The historical traffic records of external container trucks are divided into functional areas according to the gate area, buffer area and interactive operation area, and the residence time and transfer number of external container trucks in each area are counted to generate the historical data of external container truck area transfers;
[0067] Based on the historical data of the external container truck area transfer, a three-state Markov chain transfer matrix is constructed to perform probability calculation processing, calculate the state transition probability of the external container truck from the gate area to the buffer area and from the buffer area to the interactive operation area, and output the inter-area transfer probability matrix;
[0068] The interactive operation area identification location information is used as a constraint condition to perform Poisson distribution modeling on the arrival time interval of external container trucks, and the arrival rate parameters are calculated according to the historical arrival frequency of different time periods to generate the arrival time distribution model of external container trucks;
[0069] Based on the inter-regional transfer probability matrix and the external container truck arrival time distribution model, the future traffic state prediction calculation process is performed, and the change in the number of external container trucks in each region within 15 minutes is predicted through the Markov chain state equation, and the transfer probability of each functional area is output;
[0070] The queue length is predicted based on the transfer probability of each functional area and the number of container trucks outside the current area. The cumulative number and average waiting time of vehicles in the buffer area and the interactive operation area are calculated to generate the future queue length.
[0071] Specifically, after receiving historical traffic data for container trucks recorded by the terminal monitoring system, prediction module 104 divides the functional areas according to the terminal's spatial layout. The gate area is the first area for container trucks entering the terminal and includes identity verification and vehicle information registration. The buffer zone, located between the gate area and the interactive operation area, serves as a temporary parking area for container trucks awaiting job scheduling and route allocation. The interactive operation area is the core area for container trucks and automated equipment to conduct container loading and unloading operations. Functional area division is achieved by analyzing the GPS trajectory data of container trucks and assigning each trajectory point to a corresponding functional area based on its spatial coordinate range. The coordinate range of the gate area is the rectangular area near the terminal entrance, the coordinate range of the buffer zone is the strip area in the middle of the terminal, and the coordinate range of the interactive operation area is the rectangular area within the gantry crane's operating range. Dwell time statistics are calculated by calculating the time span between consecutive trajectory points of a container truck within each area. A zone transfer event is recorded when a container truck moves from one area to another. Transfer count statistics are calculated by accumulating the total number of zone transfer events for all container trucks within a specified time period. The regional transfer history data of external container trucks includes the entry time, departure time, stay time of each external container truck in each area, as well as the transfer path and transfer time between areas. After cleaning and formatting, these data form a structured historical data set.
[0072] The three-state Markov chain transition matrix is constructed based on historical data on the area transfers of external container trucks. This Markov chain abstracts the movement of external container trucks within the terminal as a random transition process between three discrete states. State transition probabilities are calculated by counting the frequency of external container trucks transitioning from one state to another and dividing it by the total number of transitions from that state. The transition probability from the gate area to the buffer area is equal to the number of times all external container trucks enter the buffer area from the gate area divided by the total number of times they depart from the gate area. The transition probability from the buffer area to the interactive operation area is equal to the number of times all external container trucks enter the interactive operation area from the buffer area divided by the total number of times they depart from the buffer area. The transition matrix is a square matrix with three rows and three columns. The matrix elements represent the probability of transitioning from the corresponding row state to the corresponding column state. The sum of the elements in each row equals one, indicating the completeness of the probability distribution. The inter-area transition probability matrix includes not only the transition probabilities between directly adjacent areas but also the probability of inter-area transitions, such as the probability of an external container truck moving directly from the gate area to the interactive operation area or from the interactive operation area back to the gate area. A zero element in the matrix indicates the absence of a corresponding transition path.
[0073] Poisson distribution modeling uses the location information of interactive operation zone markers as a constraint to statistically model the arrival time intervals of external container trucks. The Poisson distribution describes the probability distribution of the number of random events per unit time. The location information of the interactive operation zone markers defines the spatial range of external container truck arrival events. Only external container trucks entering the designated interactive operation zone are counted as arrival events. Arrival events for different interactive operation zones are modeled separately to reflect the usage characteristics of each zone. Historical arrival frequency is calculated by counting the average number of external container trucks arriving at the interactive operation zone during different time periods. Time periods include morning peak, regular daytime, evening peak, and nighttime. The arrival rate parameter for each time period is calculated by dividing the total number of arrivals during that period by the duration of the period. The arrival rate parameter reflects the density of external container truck arrivals. A higher arrival rate during the morning peak indicates more frequent arrivals, while a lower arrival rate during the night indicates less frequent arrivals. The external container truck arrival time distribution model includes Poisson distribution parameters for different time periods and corresponding time window definitions. The model can predict the probability distribution of external container truck arrivals at any point in time.
[0074] The future traffic state prediction calculation process is based on a comprehensive analysis of the inter-regional transition probability matrix and the arrival time distribution model of external container trucks. The Markov chain state equation describes the time evolution of the system state. The basic form of the state equation is that the state probability distribution at the next moment is equal to the state probability distribution at the current moment multiplied by the transition probability matrix. The state distribution at each future moment is obtained through iterative calculation. The prediction calculation divides the 15-minute prediction window into multiple time steps, each lasting 1 minute. Through 15 iterations, the trend of the number of external container trucks in each region over that 15-minute period is determined. The change in the number of external container trucks in each region is the result of the combined effects of two factors: the inter-regional transfer flow of existing external container trucks and the incremental contribution of newly arrived external container trucks to the population in each region. The transition probability output for each functional region contains the real-time transition probabilities between regions at each time step. These probabilities are dynamically adjusted based on the current system state and historical transition patterns.
[0075] Queue length prediction is based on the transition probability of each functional area and the current number of trucks in each area. Queueing theory models the waiting process of trucks in each area as a queuing system. The cumulative number of vehicles is calculated by subtracting the number of trucks leaving the area from the number of trucks entering the area during each time step. The cumulative number in the buffer zone reflects the queue of trucks awaiting work, while the cumulative number in the interactive work area reflects the number of trucks currently loading or unloading. Average waiting time is calculated based on Little's law in queuing theory: the average waiting time is equal to the average queue length in the system divided by the truck arrival rate. As the buffer queue length increases, the average waiting time increases accordingly, and as the service efficiency of the interactive work area decreases, the waiting time increases further. Future queue lengths are calculated by predicting the cumulative number of trucks in each area within a future time window. The prediction results include minute-by-minute queue lengths for both the buffer and interactive work areas over the next 15 minutes.
[0076] In one embodiment, the planning module 105 is configured to:
[0077] The transition probability is used as a road weight parameter to perform graph theory modeling on the terminal road network, and dynamic weight values are assigned to road segments according to the congestion probability of each functional area to generate a weighted terminal road network topology map;
[0078] Initializing the starting parameters of the ant colony algorithm based on the lateral deviation to perform pheromone distribution setting processing, reducing the pheromone concentration of the path segment with large lateral deviation by 20%, and outputting an initial pheromone concentration matrix;
[0079] Executing an ant colony search algorithm to perform path optimization processing based on the initial pheromone concentration matrix, setting 50 virtual ants to explore paths on the dock road network, calculating the total cost function of each path, and generating multiple candidate path solutions;
[0080] Performing path evaluation processing on the multiple candidate path solutions in combination with the heading angle deviation, adding a penalty factor to the path segments with heading angle deviation exceeding 5 degrees, screening out the optimal congestion avoidance path, and outputting a congestion avoidance detour node sequence;
[0081] Steering control calculation processing is performed based on the congestion avoidance detour node sequence and the current position of the external container truck, and the steering angle and turning radius between adjacent nodes are calculated by geometric methods to generate a steering angle correction instruction.
[0082] Specifically, after receiving the transition probability data output by the prediction module, the planning module 105 performs graph modeling to abstract the terminal road network into a mathematical graph structure. Nodes in the graph represent key locations within the terminal, including intersections, turning points, and parking spaces, and edges represent road segments connecting the nodes. The road weight parameter assignment process converts the transition probabilities for each functional area calculated by the prediction module into travel costs for the corresponding road segments. Areas with high transition probabilities indicate high traffic volume from external container trucks and are prone to congestion, and corresponding road segments are assigned higher weights to reflect the difficulty of travel. Congestion probability calculation is based on historical traffic data and current prediction results. When the number of external container trucks in a functional area exceeds the designed capacity, the congestion probability of that area increases, and the weight of the road segments connecting that area is correspondingly increased. Dynamic weight assignment is adjusted based on real-time traffic conditions. During the morning rush hour, the weight of road segments leading to the interactive operation area is increased, while during the nighttime, the weight of all road segments is decreased. Weight values range from 1 to 10, with higher values indicating higher travel costs. The weighted terminal road network topology graph contains the spatial coordinates of all nodes, the connection relationships between nodes, and the dynamic weight value of each edge. The graph structure provides the calculation basis for the subsequent path search of the ant colony algorithm.
[0083] The pheromone distribution setting process initializes the parameters of the ant colony algorithm based on the lateral deviation output by the fusion module. The pheromone concentration reflects the historical usage frequency and preference of a path segment. Initial parameter initialization sets the pheromone concentration of all road segments to the same baseline value, which is then adjusted based on the current lateral deviation of the external collection card. Path segments with large lateral deviations indicate that the external collection card is likely to deviate from the correct trajectory on that route. These path segments have poor navigation quality and should be prioritized in route selection. Pheromone concentration reduction is achieved by multiplying the pheromone concentration of path segments with lateral deviations exceeding a threshold by a 20% reduction factor to ensure a significant decrease in the attractiveness of deviating path segments. The pheromone concentration matrix is constructed based on the road network topology. The rows and columns of the matrix correspond to nodes in the road network, and the matrix elements represent the pheromone concentration of the path segments between the corresponding nodes. After the initial pheromone concentration matrix is output, the ant colony algorithm uses this matrix for path search. Path segments with high pheromone concentrations are more likely to be selected by virtual ants, while those with low pheromone concentrations are less likely to be selected.
[0084] The ant colony search algorithm performs path optimization based on an initial pheromone concentration matrix. Virtual ants are computational units that simulate the foraging behavior of biological ants. Each ant independently searches for a path and records its results. At the start of path exploration, 50 virtual ants are randomly distributed at the current location of the external truck. Each ant selects its next node to move to based on pheromone concentration and heuristic information. Node selection probability is calculated by combining pheromone concentration and distance. Nodes with high pheromone concentration and short distances are more likely to be selected, while nodes with low pheromone concentration or long distances are less likely to be selected. During the path optimization process, each ant maintains a taboo table to record visited nodes to avoid meaningless loops. A complete path search is completed when the ant reaches the target location. The total cost function calculates the comprehensive cost of each path, which includes three components: total path length, expected travel time, and congestion penalty. The total path length is calculated by summing the physical distances of each path segment. The expected travel time is calculated based on the path segment weights and vehicle speed. The congestion penalty is determined by the number of high-congestion areas traversed by the path. The process of generating multiple candidate path solutions collects all valid paths found by ants, removes duplicate paths, and sorts them according to the total cost function value. The path with the lowest cost is ranked first as the preferred solution.
[0085] The path evaluation process combines candidate routes with the heading deviation output by the fusion module. The heading deviation reflects the degree to which the truck's current heading differs from the desired direction. The heading angle of each path segment is calculated by analyzing the path geometry. The heading angle remains constant for straight segments, varies continuously for curved segments, and varies significantly for sharp turns. Path segments with heading deviations exceeding 5 degrees are considered to require significant steering adjustments. These segments are subject to an additional penalty factor to reflect the complexity and risk of the steering operation. The penalty factor is calculated based on the specific heading deviation. Greater deviations increase the penalty factor. Path segments with deviations exceeding 10 degrees have the penalty factor doubled, while those with deviations exceeding 15 degrees are excluded. The overall path score is calculated by adding the total path cost to the heading penalty factor. The path with the lowest score is selected as the optimal congestion-avoiding path, which minimizes both cost and steering complexity while still meeting the target destination. The node sequence output for traffic congestion avoidance and detour includes the coordinate information of all key nodes on the optimal path and the connection relationship between the nodes. The node sequence is arranged in the order of driving, and the distance and direction information between adjacent nodes are clearly recorded.
[0086] The steering control calculation process performs a geometric analysis based on the sequence of detour nodes and the current position of the truck. Geometric methods include vector calculations and trigonometric operations to determine steering parameters. The steering angle between adjacent nodes is calculated by analyzing the angular relationship formed by three consecutive points. The truck's current position, the next node, and the next-next node form a triangle, and the steering angle is equal to the angle between the truck's current direction of travel and the target direction. Turning radius calculation is based on the truck's kinematic constraints and road geometry. The minimum turning radius of large trucks is limited by their wheelbase length and steering angle. Turning on narrow roads requires a larger turning radius to avoid collisions with roadside facilities. The geometric method uses a vector cross product to determine the steering direction. A positive cross product indicates a left turn, while a negative cross product indicates a right turn. The absolute value of the cross product reflects the magnitude of the steering angle. The steering angle correction instruction generation process converts the calculated steering angle and turning radius into a control instruction that can be executed by the external container truck driver or the automatic driving system. The instruction contains the specific value of the steering angle, a clear indication of the steering direction and the recommended driving speed. The steering angle is in degrees and is accurate to one decimal place. The steering direction is clearly indicated as left or right turn. The recommended speed is determined according to the turning radius and road conditions.
[0087] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A full-path guidance system for container trucks outside a container terminal, characterized in that the system include: The acquisition module is used to acquire and process real-time images of the driving path of the external container terminal through the vehicle-mounted multimodal image acquisition unit to generate lane line coordinate data and interactive operation area identification position information; An extraction module is used to input the lane line coordinate data into a ResNet-50 convolutional neural network for feature extraction processing, output a guide line feature vector, and identify the coordinates of the path reference point; The fusion module is used to perform position fusion calculation processing of the external container truck through adaptive Kalman filtering according to the coordinates of the path reference point and the characteristic vector of the guide mark, and calculate the lateral deviation and the heading angle deviation degree, including: constructing a state vector with the coordinates of the path reference point and the current GPS position data of the external container truck for initialization processing, setting the position, speed and heading angle of the external container truck in the terminal coordinate system as state variables, and generating the motion state vector of the external container truck; establishing a vehicle motion prediction model in the terminal environment according to the motion state vector of the external container truck for state prediction processing, taking into account the motion characteristics of the external container truck when turning and going straight in the terminal, and outputting the predicted position coordinates and prediction error. The guide line feature vector is converted into observation data of the external container truck relative to the lane centerline for observation update processing, and the observation residual of the external container truck deviating from the guide path is calculated to generate the Kalman gain coefficient; based on the Kalman gain coefficient, the predicted position of the external container truck is adaptively weighted and fused, and the fusion weight is dynamically adjusted in combination with the GPS positioning accuracy and the visual observation confidence, and the positioning result of the external container truck within the terminal is output; the positioning result of the external container truck is analyzed and processed according to the deviation from the centerline of the terminal guide path, and the vertical distance of the external container truck from the lane centerline and the angle between the vehicle head pointing and the path direction are calculated to generate the lateral deviation and the heading angle offset degree; A prediction module is used to input the historical passage records of external container trucks into a Markov chain model in combination with the interactive operation area identification location information to perform arrival time prediction processing, predict the transfer probability of each functional area and calculate the future queue length, including: dividing the historical passage records of external container trucks into functional areas according to the gate area, buffer area and interactive operation area, counting the residence time and transfer number of external container trucks in each area, and generating historical data of external container truck area transfers; constructing a three-state Markov chain transfer matrix based on the historical data of external container truck area transfers for probability calculation processing, calculating the state transition probability of the external container truck from the gate area to the buffer area and from the buffer area to the interactive operation area, and outputting an inter-area transition probability matrix; The interactive operation area identification location information is used as a constraint to perform Poisson distribution modeling on the arrival time interval of external container trucks. The arrival rate parameters are calculated based on the historical arrival frequency of different time periods to generate an external container truck arrival time distribution model. Future traffic state prediction and calculation are performed based on the inter-regional transition probability matrix and the external container truck arrival time distribution model. The change in the number of external container trucks in each area within 15 minutes is predicted using the Markov chain state equation, and the transition probability of each functional area is output. The queue length is predicted based on the transition probability of each functional area and the current number of external container trucks in each area. The cumulative number of vehicles and the average waiting time in the buffer zone and the interactive operation area are calculated to generate the future queue length. A planning module is used to perform path replanning processing through an improved ant colony algorithm based on the transfer probability and the lateral deviation, generate a detour node sequence based on the heading angle deviation degree and output a steering angle correction instruction, including: using the transfer probability as a road weight parameter to perform graph theory modeling processing on the terminal road network, assigning dynamic weight values to road segments according to the congestion probability of each functional area, and generating a weighted terminal road network topology map; initializing the starting parameters of the ant colony algorithm based on the lateral deviation to perform pheromone distribution setting processing, reducing the pheromone concentration of the path segment with a large lateral deviation by 20%, and outputting an initial pheromone concentration matrix; according to the The initial pheromone concentration matrix is used to perform path optimization using an ant colony search algorithm. Fifty virtual ants are set up to explore paths on the terminal road network, and the total cost function of each path is calculated to generate multiple candidate path plans. Path evaluation is performed on the multiple candidate path plans in combination with the heading angle deviation degree. A penalty factor is added to path segments with heading angle deviations exceeding 5 degrees. The optimal congestion avoidance path is screened out, and a congestion avoidance detour node sequence is output. Steering control calculation processing is performed based on the congestion avoidance detour node sequence and the current position of the external container truck. The steering angle and turning radius between adjacent nodes are calculated using a geometric method to generate a steering angle correction instruction.
2. The full-path guidance system for container trucks outside the automated container terminal according to claim 1 is characterized in that: The acquisition module is used to: The CCD camera is used to continuously scan and photograph the dock pavement, and edge detection is performed based on pixel grayscale value analysis to extract the lane boundary pixel coordinate sequence; Inputting the lane line boundary pixel coordinate sequence into the Hough transform algorithm for straight line fitting processing, calculating the lane line slope and intercept parameters and generating lane line coordinate data; Use an infrared thermal imager to scan the temperature distribution of the interactive operation area, perform area segmentation based on the temperature difference threshold, and identify the boundary contour coordinates of the interactive operation area; Performing morphological closing operation filtering on the boundary contour coordinates of the interactive operation area to eliminate noise interference and extract the complete area contour to generate interactive operation area identification position information; The lane line coordinate data and the interactive operation area identification position information are subjected to coordinate system unification processing, a terminal global coordinate reference system is established, and path environment data in a unified coordinate format is output.
3. The full-path guidance system for container trucks outside the automated container terminal according to claim 1 is characterized in that: The extraction module comprises: a processing unit, configured to input the lane line coordinate data into an input layer of a ResNet-50 network for data preprocessing, normalize the coordinate data, and convert it into a standard input format of 224×224 pixels; an operation unit, configured to perform a 7×7 convolution operation on the data in the standard input format through a first convolution block, set a step size of 2, and perform batch normalization processing to extract a primary feature map; A generation unit is used to sequentially input the primary feature map into four residual blocks for deep feature extraction processing, where each residual block contains three convolutional layers and a skip connection structure to generate a multi-scale deep feature map; A dimensionality reduction unit is used to perform global average pooling dimensionality reduction processing on the multi-scale depth feature map, compress the feature map into a 2048-dimensional feature vector and generate a guide line feature vector; The transformation unit is used to perform linear transformation processing on the guide mark feature vector through a fully connected layer, calculate the probability distribution of each pixel point belonging to the path reference point through a softmax activation function, identify the highest probability position and output the coordinates of the path reference point.
4. The full-path guidance system for container trucks outside the automated container terminal according to claim 3 is characterized in that: The computing unit is used for: Perform a 7×7 convolution kernel sliding window scan on the standard input format data, moving 2 pixel positions each time and calculating the convolution value to generate a 112×112 pixel convolution output matrix; Input the convolution output matrix into the batch normalization layer for data normalization, calculate the batch mean and variance, perform normalization transformation, and output a standardized feature matrix; Performing a nonlinear transformation of the ReLU activation function on the standardized feature matrix, setting negative values to zero and keeping positive values unchanged, to generate an activated feature map; Downsample the activated feature map through a 3×3 max pooling layer, set the step size to 2 and select the maximum value in the pooling window, and output a pooled feature map of 56×56 pixels; The pooled feature map is reorganized in channel dimension, the depth dimension of the feature map is adjusted to 64 channels while maintaining the spatial resolution, and a primary feature map is generated.
Citation Information
Patent Citations
Automatic driving vehicle obstacle avoidance path planning method, vehicle and readable storage medium
CN113985896A
Port logistics dynamic path planning method and system based on multi-source data fusion
CN120160637A