A method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking
Through dynamic reinforcement learning and Kalman particle tracking methods, real-time and accurate monitoring of the lifting state of the locking card is achieved, solving the problem of frequent locking card accidents and improving the safety and efficiency of container transportation.
Patent Information
- Application Number
- CN202410951263.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-07-16
AI Technical Summary
It is difficult for the prior art to realize real-time accurate monitoring of the lifting status of the trucks during container transportation, resulting in frequent accidents of trucks and lifting, which poses safety hazards.
Using a method based on dynamic reinforcement learning and Kalman particle tracking, the container and truck bracket positions are located through the YOLOv8 segmentation model, combined with the adaptive Kalman particle filtering tracking algorithm for real-time target tracking, and through dynamic reinforcement learning, intelligent decision-making of lifting motion trends is realized, and alarm signals are output.
The monitoring accuracy and alarm accuracy of the lifting status of the truck are improved, ensuring that the alarm signal is issued in time when the truck is lifted, and secondary damage is prevented.
Smart Images

Figure CN118823066B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of port operations, and in particular to a container truck anti-lifting method based on dynamic reinforcement learning and Kalman particle tracking. Background Art
[0002] In the container shipping industry, truck lifting accidents are a common safety hazard that cannot be ignored. These accidents often occur when a gantry crane (such as an RTG / RMG) is lifting a container to the yard. Because the truck's locking pins and the container's locking holes fail to fully release, the gantry crane accidentally lifts the truck along with it, posing a serious threat to the safety of personnel and vehicles.
[0003] Current technologies rely heavily on manual inspection and laser detection. These methods are not only inefficient but can also lead to false alarms and missed detections in certain situations, making accurate real-time monitoring difficult. Therefore, to address this challenge, the present invention proposes an innovative solution for intelligently determining and monitoring the lifting status of container trucks in real time. This approach can promptly detect and prevent potential container truck lifting accidents, significantly improving the safety and efficiency of container transportation and meeting the high safety standards of modern ports. Summary of the Invention
[0004] In order to overcome the above-mentioned shortcomings of the prior art, the present invention proposes a method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking. First, the positions of the container and the container truck bracket are accurately located from the video stream images captured by the camera through the YOLOv8 segmentation model. Then, the Kalman particle filter tracking algorithm optimized by reinforcement learning is used to realize real-time motion tracking of both the container and the container truck bracket. Dynamic reinforcement learning is introduced to make decisions on the lifting motion trend of the container truck. When the lifting strategy is met, the controller outputs an alarm signal to the external system for a protective response to prevent secondary injuries to personnel and the container truck.
[0005] The technical solution adopted by the present invention to solve the technical problem is: a method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking, comprising the following steps:
[0006] Step 1: Decode image data using the real-time collected truck video stream data;
[0007] Step 2: Use the YOLOv8 segmentation model to locate the position of the container and the truck bracket;
[0008] Step 3: Use the adaptive Kalman particle filter tracking algorithm to track the container and the truck bracket in real time;
[0009] Step 4: Implement intelligent decision-making on the lifting movement trend of container trucks through dynamic reinforcement learning.
[0010] Compared with the prior art, the present invention has the following positive effects:
[0011] The present invention proposes an adaptive dynamic policy optimization ADPO (Adaptive Dynamic Policy Optimization) reinforcement learning method, which has two core ideas: First, the Kalman particle filter algorithm combines the linear prediction ability of the Kalman filter and the nonlinear processing ability of the particle filter, and optimizes the tracking accuracy and robustness by adaptively and dynamically adjusting the particle distribution, further improving the speed and accuracy of target tracking; second, dynamic policy adjustment. Unlike traditional reinforcement learning algorithms, the ADPO adaptive dynamic policy reinforcement learning algorithm can adjust tracking parameters according to real-time tracking results to cope with different environments and target states, thereby improving the speed and accuracy of tracking two targets: containers and container truck brackets. Secondly, the present invention proposes to realize lifting decisions through a dynamic reinforcement learning algorithm, thereby improving the accuracy of the container truck lifting alarm, thereby ensuring that when the container truck is lifted, the system can promptly send an alarm signal to the external system for protective response. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The present invention will now be described by way of example with reference to the accompanying drawings, in which:
[0013] Figure 1 It is a schematic flow diagram of the present invention. DETAILED DESCRIPTION
[0014] The container truck anti-lifting method based on dynamic reinforcement learning and Kalman particle tracking of the present invention comprises the following steps:
[0015] Step S1: Two cameras installed on the side of the gantry crane operation lane collect real-time video stream data of the truck and transmit it to the controller through the Ethernet switch. FFmpeg is used to decode image frames from the video stream.
[0016] In step S2, the image is input into the YOLOv8 segmentation model to perform instance segmentation of the container and the truck bracket. The segmentation results of multiple objects in the image are post-processed to form the minimum enclosing rectangle of the segmentation.
[0017] Step S3, using the Kalman particle filter tracking algorithm optimized by the adaptive dynamic strategy reinforcement learning algorithm to perform real-time target tracking on the container and the truck bracket in the image;
[0018] Step S4 uses a reinforcement learning-based proximal policy optimization (PPO) algorithm, combining a multilayer perceptron and a Gaussian distribution as the policy network. The input is the location information of the top edge center of the truck bracket segmentation target and the bottom edge center of the container segmentation target. The network outputs the mean and standard deviation of the action, and thus the probability distribution of the alarm action, implementing an intelligent truck anti-lift alarm.
[0019] Specifically, the container truck anti-lifting system used in the present invention includes a controller, an Ethernet switch and two cameras, wherein:
[0020] (1) The switch connects the camera and the controller for data transmission.
[0021] (2) The camera is a front-end video data acquisition device used to collect video stream data of the working lane. The present invention installs two cameras on the working side. Each camera performs detection and logical judgment independently. The system takes the union of the judgment results of the camera image segmentation and tracking to perform lifting protection. The specific implementation is described with two cameras, as follows:
[0022] Two cameras are installed on the working side. They are mounted horizontally and symmetrically, facing the front and rear of the truck, respectively. Adjustment is possible based on site conditions. The front camera focuses on the rear of the truck's carriage, while the rear camera focuses on the front. The imaging ranges of the two cameras intersect in the middle of the truck, ensuring that both cameras cover the entire truck's carriage. Within the imaging range of a single camera, the upper edge of the truck's carriage is parallel to the bottom line of the image and falls within a vertical range of 3 / 7 to 4 / 7 of the image.
[0023] (3) The controller serves as the main data processing unit in the present invention and is responsible for real-time processing of the video data collected by the camera connected to the controller, detecting the operating status of the gantry crane, and comprehensively judging whether a lifting accident has occurred based on the video stream data of the operating side camera. When a lifting accident is detected, the controller immediately outputs an alarm signal to the external system for a protective response to prevent the staff and the container truck from suffering secondary injuries.
[0024] The system, which incorporates a controller, a switch, and two cameras, collaborates with external systems to implement alarm and automatic protection functions. Once the system is operational, the controller captures the camera's video stream in real time and performs video decoding, image segmentation, and target tracking. It then analyzes the image segmentation results and the output of a reinforcement learning algorithm to determine whether the gantry crane is lifting the container truck. While the gantry crane is operating, the controller analyzes the movement trends of the container truck and container based on the image segmentation and target tracking results of consecutive frames to determine the lifting logic. If the judgment is that the container truck is lifting, the controller immediately outputs an alarm signal to the external system for a protective response.
[0025] The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking is as follows: Figure 1 As shown, the following steps are included:
[0026] Step S1: After the controller is powered on, the system automatically runs, loads the camera communication parameters in the configuration file, uses the corresponding communication protocol to connect the camera, log in and verify, decode and obtain the stream, and other initialization tasks. After the initialization is completed, the system collects the video stream data captured by the camera on the operating side in real time, and uses ffmpeg to decode the image data from the collected video stream data;
[0027] Step S2: Use the trained YOLOv8 image segmentation model to segment the container and the truck bracket from the image, obtain a segmentation map of the container and the truck bracket, and perform post-processing on the segmentation map to form a minimum enclosing rectangle for the segmentation. The details are as follows:
[0028] In step S201, an advanced YOLOv8 image segmentation model based on deep learning is selected, and the model is trained using a dataset of labeled containers and truck carriages to ensure that the model can accurately segment the two objects in the image at the pixel level.
[0029] Step S202: Before image segmentation, the input image is subjected to necessary preprocessing such as scaling and normalization to meet the input requirements of the model.
[0030] Step S203: Image segmentation: Input the preprocessed image into the trained deep learning YOLOv8 image segmentation model to obtain the category label of each pixel. Based on the category label, the image is segmented into different regions, each corresponding to an object or background.
[0031] Step S204, extract the target area, and identify and calculate the target area to be tracked in the segmentation result. The segmentation result obtained from the deep learning YOLOv8 image segmentation model is a probability map (or score map), in which each pixel has a corresponding category probability. Set the threshold to 0.5 to convert the probability map into a binary map (only the target and the background). Use Gaussian filtering to remove noise in the image segmentation results, retain more edge details, and make the segmentation results more refined. Since threshold processing may produce some small, discontinuous areas, or there are holes inside the target area, the expansion in the morphological operation is used to fill these holes and connect small areas. After the expansion operation, the connectivity analysis algorithm is used to perform a depth-first search from the binary map to find the connected areas of the pixels with a value of 1 in the binary map. The largest connected area is selected as the target area, and the bounding box of the target area is calculated.
[0032] Step S3 uses the Kalman particle filter tracking algorithm optimized by the adaptive dynamic strategy reinforcement learning algorithm to perform real-time target tracking on the two targets in the image: the container and the truck bracket. The adaptive Kalman particle filter tracking algorithm combines the linear prediction capability of the Kalman filter with the nonlinear processing capability of the particle filter. By adaptively and dynamically adjusting the particle distribution, it improves the speed, accuracy, and robustness of target tracking. Using the adaptive dynamic strategy reinforcement learning algorithm, the tracking parameters are adjusted according to real-time data to cope with different environments and target states, further improving the tracking accuracy of the Kalman particle filter. The specific steps are as follows:
[0033] Step S301, system initialization:
[0034] Set the initial parameters and determine the initial positions of the truck bracket and container. Initialize the particle filter and set the number of particles, initial positions, and weights. Use the target box positions of the truck bracket and container predicted by the YOLOv8 segmentation model to initialize the particle filter particle set. The particle set represents the possible distribution of the target state. Let the particle set be Where i=1,2,...,N, N is the number of particles. Initially, each particle From is centered and has a covariance of is sampled from a normal distribution.
[0035] Setting up the reinforcement learning model:
[0036] Define the state space: Define the Euclidean distance between the predicted position of the truck carriage and container and the actual position of the target as the state.
[0037] Define the action space: An action is an adjustment to the tracking-related parameters.
[0038] Design reward function: The reward function is based on the accuracy and real-time performance of tracking. This invention designs two reward functions: tracking accuracy reward function and stability reward.
[0039] The tracking error e is the Euclidean distance between the predicted position and the true position, and a reward function R based on the tracking error is defined accuray (e), the function is defined as follows:
[0040]
[0041] Here, α is a positive hyperparameter that controls the magnitude of the reward. When the error is small, the reward is close to α, and as the error increases, the reward becomes negative.
[0042] In order to encourage tracking stability, this paper defines a stability reward function R based on the rate of change of tracking error between consecutive frames. stability(s). Assume that s is the rate of change (absolute value) of the tracking error between adjacent frames, a positive reward is given for a smaller rate of change, and a negative reward is given for a larger rate of change. The specific reward function is as follows:
[0043]
[0044] Where γ is a positive hyperparameter and σ is a threshold for the rate of change. This reward function encourages the policy network to maintain tracking stability when objects such as containers and truck carriages move or when the environment changes.
[0045] Step S302: Kalman particle filter tracking
[0046] (1) Kalman filter prediction:
[0047] Based on the target state at the previous moment, use Kalman filtering to predict the target state at the current moment. Let the target state be The prediction equation of Kalman filter is:
[0048]
[0049] Where A is the state transfer matrix, B is the input matrix, is the estimated value of the state at the previous moment, u k-1 The input at the previous moment.
[0050] (2) Particle set filter update: the state estimation of the Kalman filter is used as the input of the particle filter, and the particle position and weight of the particle filter are updated according to the real-time observation data. Get the observation data z at the current moment k , for each particle Calculate its observation likelihood Calculate the new weight of each particle based on the observation likelihood and particle weight The weight update formula is:
[0051]
[0052] Use recursive filters for prediction and correction to estimate the internal state of a dynamic system.
[0053] (3) Adaptive particle resampling and weight adjustment. Resampling is performed based on the particle weights. Different from the traditional resampling method, this invention designs an adaptive strategy. First, the effective sample size N of the particle weight is calculated. eff , N eff is the number of effective particles in the particle set, which reflects the efficiency of the particle set in representing the posterior probability distribution. The formula is as follows:
[0054]
[0055] When N eff When it is less than the threshold value N / 2, it means that the particle weight distribution is too concentrated and resampling is required. However, in order to avoid particle depletion (that is, all particles become very similar), an adaptive strategy is used to dynamically adjust the frequency and method of resampling. eff When the weight is lower than the threshold N / 2 but higher than another lower threshold N / 4, only the particles with lower weights are partially resampled to balance the diversity of particles and tracking accuracy. In this strategy, particles with high weights are retained, while particles with low weights are resampled according to certain rules. eff If it is less than N / 4, a full resampling is performed. The following is a detailed description of how to implement this step:
[0056] (1) Calculate particle weight:
[0057] First, according to the current observation data z k , calculate each particle The observation likelihood of And use this likelihood to update the particle weight
[0058] (2) Calculate the effective sample size:
[0059] Calculate the effective sample size N using the updated weights eff .
[0060] (3) Check whether partial resampling or full resampling is required:
[0061] If N eff Below a preset threshold N / 2 but above another lower threshold N / 4, a partial resampling is performed. eff If it is less than N / 4, it means that the particle weight distribution is very concentrated, and a full resampling is performed to avoid particle depletion.
[0062] (4) Determine the number of resampling particles during partial resampling:
[0063] Set N resample The number is 0.7*N eff , ensuring that the number of retained particles is sufficient to reflect the main features of the posterior probability distribution.
[0064] (5) Resample low-weight particles:
[0065] Sort the particles by weight from low to high. Select the NN with the lowest weight resampleParticles are resampled. For each particle that needs to be resampled, a random selection is made based on its weight. The lower the weight, the lower the probability of being selected, but there is still a possibility of being selected. For each selected low-weight particle, it is first copied, that is, it is replaced with a new particle, and the state of the new particle is the same as the original particle. In order to increase the diversity of the particle set, the present invention adds a small random perturbation to the state of the copied particle. This perturbation should be small enough to avoid introducing excessive deviations, but large enough to increase the diversity of the particles. A perturbation factor δ is set to 0.01, which determines the amplitude of the perturbation. In each iteration, the perturbation factor is dynamically adjusted according to the tracking results and the distribution of particle weights. Specifically, the perturbation factor is adjusted based on the variance of the particle weights. When the variance of the particle weights is large, it means that the diversity of the particle set is good. At this time, the perturbation factor can be appropriately reduced to improve the accuracy. When the variance of the particle weights is small, it means that the diversity of the particle set is poor. At this time, the perturbation factor can be appropriately increased to increase the diversity. The specific adjustment formula can be expressed as:
[0066]
[0067] Among them, δ base is the basic disturbance factor, δ is the adjusted disturbance factor, δ w is the standard deviation of the current particle weight, is the maximum value of the standard deviation of particle weights.
[0068] To ensure that it does not cause the particle to deviate too far from the true state, while effectively increasing diversity. For each replicated particle, add a random perturbation to each dimension of its state vector x. This perturbation can be a random perturbation with a mean of 0 and a variance of δ. 2 Sampling from a normal distribution yields:
[0069] x new =x original +N(0,δ 2 )
[0070] x original is the original state vector, N(0,δ 2 ) represents a value with a mean of 0 and a variance of δ 2 Normal distribution, x new is the state vector after adding the perturbation. This ensures that each particle can be properly perturbed in each dimension, thereby increasing the diversity of the entire particle set.
[0071] The weight of the new particle is re-evaluated and normalized according to the current observation data. The specific formula is:
[0072]
[0073] Among them, p(z k |x i ) is particle x i In the observed data z k The likelihood under w i is the normalized weight of the new particle.
[0074] Through the above method, the present invention can gradually increase the diversity of particles while maintaining the approximate accuracy of the particle set to the posterior distribution, thereby improving the performance of the Kalman particle filter tracking algorithm.
[0075] (6) Retain high-weight particles:
[0076] Except for the resampled particles, the remaining high-weight particles remain unchanged.
[0077] (7) Iterative execution algorithm:
[0078] The updated particle set (including resampled particles and retained high-weight particles) is used to continue the algorithm's subsequent observation update, state estimation, and other steps. During tracking, the position and range of the particles are dynamically adjusted based on the target's motion in the image and the output of the YOLOv8 image segmentation model. When the target moves, the particle position is updated based on the target position information provided by the YOLOv8 image segmentation model; when the target shape or size changes, the particle range is updated based on the target area information provided by the YOLOv8 image segmentation model.
[0079] This partial resampling strategy can update low-weight particles while maintaining a portion of high-weight particles, thereby balancing particle diversity and tracking accuracy to a certain extent. This strategy is more flexible than full resampling, especially when the particle weight distribution is not particularly extreme.
[0080] Step S303: Adaptive dynamic strategy reinforcement learning:
[0081] (1) State perception:
[0082] Sense the Euclidean distance between the predicted position of each adaptive Kalman particle filter tracking algorithm and the actual position of the target.
[0083] (2) Select an action:
[0084] Based on the state-aware distance between the predicted box and the actual box, an adaptive dynamic ε-greedy strategy is used to adaptively select the parameters that need to be adjusted for tracking.
[0085] (3) Execute action:
[0086] Based on the adaptively selected parameter type, the reinforcement learning model is used to adjust the parameter in real time to obtain the closeness between the new position of the target and the actual position of the target.
[0087] (4) Observation feedback:
[0088] Observe the target position information after the action is performed and the tracking error between the new position and the target's true position.
[0089] (5) Update rewards:
[0090] The reward value is updated based on the observation results. The update policy uses the adaptive policy adjustment type algorithm Q-learning to update the reward value of the reinforcement model to adapt to the new environment and target state.
[0091] Step S304: Real-time target tracking:
[0092] The target state updated by the Kalman particle filter algorithm optimized by reinforcement learning is obtained, and the above tracking steps are repeated to continuously update the estimated value of the target state to achieve real-time tracking of both the container and the truck bracket.
[0093] Step S4, determine whether the container truck is lifted:
[0094] Based on the reinforcement learning proximal policy optimization (PPO) algorithm, a multilayer perceptron and Gaussian distribution are combined as the policy network. The input is the position information of the center of the upper edge of the truck bracket segmentation target and the center of the bottom edge of the container segmentation target. The network can output the mean and standard deviation of the action, and thus output the probability distribution of the action (in this case, the action is either alarm or no alarm). The details are as follows:
[0095] Step S401: Data preprocessing. Define the positional state information of the center point of the upper edge of the truck carriage segmentation target and the center point of the container segmentation target as input. Converting this positional information into a numerical form that can be processed by the deep learning neural network and normalizing it to the range of [0, 1] helps the neural network better process the data and accelerates the training process.
[0096] In step S402, the data is input into a multilayer perceptron (MLP). Using the trained MLP, the mean and standard deviation of the alarm activity are output. The input layer contains only the container's location information and has two neurons, each containing the x and y coordinates of the target's location. The hidden layer consists of one layer of neurons, each followed by a ReLU activation function. The output layer has two neurons: one for outputting the mean of the action and the other for outputting the standard deviation of the action.
[0097] Step S403, reinforcement learning proximal policy optimization PPO algorithm is applied. The multi-layer perceptron of step S402 is used as the policy network, and the mean and standard deviation of the alarm action are output according to the position information of the center point of the upper edge of the truck bracket segmentation target and the center point of the bottom edge of the container segmentation target. Interact with the environment, collect data (state, action, reward and next state), and use this data to update the policy network. A policy network was trained using the reinforcement learning proximal policy optimization PPO algorithm. The goal of the network is to learn a policy that can select the optimal action (alarm or not alarm) based on the current state to maximize the long-term cumulative reward. By constantly interacting with the environment and collecting new data, the system can continuously update the policy network and iteratively optimize it so that it can gradually learn a better anti-lifting strategy. The detailed description is as follows:
[0098] 1. The status is the position information of the center point of the upper edge of the truck bracket segmentation target and the center point of the lower edge of the container segmentation target.
[0099] 2. The action is to alarm when the container truck is lifted or not.
[0100] 3. The description of the rewards is as follows:
[0101] (1) Successful detection of a truck lifting: When the system successfully detects that a truck is in a dangerous state of lifting or about to be lifted, and takes timely preventive measures such as issuing an alarm, the system will give a positive reward. If the system issues an early warning and successfully prevents the truck from being lifted, the reward value will be relatively high; if the truck is already partially lifted but the system still responds in time, the reward value will be slightly lower.
[0102] (2) No lifting event occurs: If the system does not detect any lifting event within a period of time and the trucking operation proceeds normally, the system will give a small positive reward. This reward encourages the system to maintain a stable operating state and continue to monitor potential risks.
[0103] (3) Lifting event occurs: If the system fails to detect the lifting of the container truck in time or fails to effectively prevent the occurrence of the lifting event, the system will give a large negative reward to reflect the negative impact of the lifting event on system performance and safety.
[0104] The present invention designs a shape excitation function and defines key variables:
[0105] d(t) is the real-time distance between the center point of the upper edge of the truck bracket segmentation target and the center point of the bottom edge of the container segmentation target, and (t) is the time step.
[0106] d safeThis is a safety distance threshold, with a reference value of 0.4 meters, which can be adjusted based on the specific conditions on site. When the distance between the truck bracket and the container is greater than this threshold, the system considers the truck to be in a safe state without being lifted.
[0107] d danger This is a dangerous distance threshold, with a reference value of 0.2 meters, which can be adjusted according to the specific conditions on site. When the distance between the truck bracket and the container is less than this threshold, the system considers the truck to be in a dangerous state of being lifted.
[0108] The following is the specific design of the shape activation function:
[0109]
[0110] Among them, R safe is a positive value, with a reference value set to 10, and rewards are given when the distance between the truck bracket and the container is greater than the safety threshold. β is a coefficient set to 1.2, which is used to adjust the rate at which rewards decrease when transitioning from a safe state to a dangerous state. danger It is a negative value, set to -10, which gives a penalty when the distance between the truck bracket and the container is less than the danger threshold. crash A large negative value, set to -20, gives a severe penalty when a hoisting event actually occurs.
[0111] This shape reward function takes into account the distance between the center point of the upper edge of the truck carriage segmentation target and the center point of the bottom edge of the container segmentation target, and gives different rewards or penalties depending on the range of distance. When the distance is above the safe threshold, a positive reward is given; when the distance is below the dangerous threshold, a negative reward is given; when the distance is between the two, the reward value decreases linearly with the distance.
[0112] 4. In the next state, after executing the action, the system needs to collect the new position information of the center point of the upper edge of the truck bracket segmentation target and the center point of the bottom edge of the container segmentation target.
[0113] Through the above steps, the container truck anti-lifting system can use reinforcement learning methods to interact with the environment and optimize by continuously updating the policy network, so that it can quickly and accurately detect being lifted and output alarm signals to the external system for protective response, preventing workers and container trucks from suffering secondary injuries.
[0114] Through the above steps, we can combine the adaptive dynamic strategy reinforcement learning algorithm with the Kalman particle filter to build an efficient and accurate hybrid tracking system, achieve real-time target tracking of the container truck bracket and container, and accurately determine whether the container truck is lifted through the reinforcement learning algorithm.
Claims
1. A method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking, characterized by: The steps include: Step 1: Decode image data using the real-time collected truck video stream data; Step 2: Use the YOLOv8 segmentation model to locate the position of the container and the truck bracket; Step 3: Use the adaptive Kalman particle filter tracking algorithm to track the container and the truck bracket in real time: Step S301, system initialization: set initial parameters and determine the initial positions of the truck bracket and container; initialize the particle filter, set the number of particles, initial positions and weights, and use the target frame positions of the truck bracket and container predicted by the YOLOv8 segmentation model to initialize the particle filter particle set; Step S302, Kalman particle filter tracking: (1) Kalman filter prediction: Based on the target state at the previous moment, the Kalman filter is used to predict the target state at the current moment; (2) Particle set filter update: The state estimate of the Kalman filter is used as the input of the particle filter, and the particle positions and weights of the particle filter are updated according to the real-time observation data; (3) Use the updated particle set to continue performing subsequent observation updates and state estimation; Step S303: Adaptive dynamic strategy reinforcement learning: (1) State perception: Calculate the Euclidean distance between the predicted position predicted by the Kalman filter and the actual position of the target; (2) Execute the action: Adaptively select the parameters that need to be adjusted for tracking based on the Euclidean distance between the predicted position and the target's true position, and then use the reinforcement learning model to adjust the parameters in real time to obtain the proximity between the target's new position and the target's true position; (3) Observation feedback: Calculate the tracking error between the new target position after the action is performed and the target's actual position; (4) Update reward: Update the reward value based on the tracking error; Step S304: Real-time target tracking: Obtain the target state updated by the Kalman particle filter algorithm after reinforcement learning optimization, repeat the above tracking steps, and continuously update the estimated value of the target state; Step 4: Implement intelligent decision-making on the lifting movement trend of container trucks through dynamic reinforcement learning.
2. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 1 is characterized in that: During the tracking process, the position and range of the particles are dynamically adjusted according to the movement of the target in the image and the output of the YOLOv8 image segmentation model: when the target moves, the position of the particles is updated according to the target position information provided by the YOLOv8 image segmentation model; when the shape or size of the target changes, the range of the particles is updated according to the target area information provided by the YOLOv8 image segmentation model.
3. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 1 is characterized in that: The particle set filter update method includes the following steps: (1) Update the particle weight according to the following formula Where z k is the current observation data, For each particle The observation likelihood of , N is the number of particles; (2) Calculate the effective sample size N according to the following formula eff : (3) Determine whether partial resampling or full resampling is required; (4) Determine the number of resampling particles N during partial resampling resample 0.7*N eff ; (5) Resample low-weight particles.
4. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 3 is characterized in that: The method to determine whether partial resampling or full resampling is needed is: When N eff When N is lower than N / 2 but higher than N / 4, partial resampling is performed; when N eff If N is less than N / 4, a full resampling is performed.
5. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 3 is characterized in that: The method of resampling low-weight particles is to sort the particles from low to high according to their weights and select the NN with the lowest weight. resample Particles are resampled. For each particle that needs to be resampled, it is randomly selected according to its weight. For each selected low-weight particle, it is first copied. Then, for each copied particle, a random perturbation is added to each dimension of its state vector x by scanning the following formula: x new =x original +N(0,δ 2 ) x original is the original state vector, N(0,δ 2 ) represents a value with a mean of 0 and a variance of δ 2 Normal distribution, x new is the state vector after adding the disturbance; where: δ is the adjusted disturbance factor, which is calculated as follows: Among them, δ base is the basic perturbation factor, δ w is the standard deviation of the current particle weight, is the maximum value of the standard deviation of particle weights.
6. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 1, characterized in that: The reward functions used by reinforcement learning models include: (1) Reward function R based on tracking error accuray (e) Among them, α is a positive hyperparameter, e is the tracking error, and ∈ is the threshold of the tracking error; (2) Stability reward function R based on the rate of change of tracking error between consecutive frames stability (s): Where s is the rate of change of tracking error between adjacent frames, γ is a positive hyperparameter, and σ is the threshold of the rate of change.
7. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 1 is characterized in that: The method described in step 4 for realizing intelligent decision-making on the lifting movement trend of the container truck through dynamic reinforcement learning is as follows: using a combination of a multi-layer perceptron and a Gaussian distribution as a policy network, taking the position information state of the center point of the upper edge of the container truck bracket segmentation target and the center point of the bottom edge of the container segmentation target as input, using the trained multi-layer perceptron to output the mean and standard deviation of the alarm action, thereby outputting the action of the container truck lifting alarm or no alarm, and then performing stimulation; interacting with the environment, and using the collected state, action, and stimulation data to update the policy network.
8. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 7, characterized in that: The following shape excitation function is used to excite the action of the container truck lifting alarm or no alarm: Among them, R safe is a positive value, β is a coefficient, R danger is a negative value, R crash is a large negative value, d(t) is the real-time distance between the center point of the upper edge of the truck bracket segmentation target and the center point of the bottom edge of the container segmentation target, (t) is the time step, d safe is a safety distance threshold, d danger is a danger distance threshold.
9. The method for preventing container trucks from lifting based on dynamic reinforcement learning and Kalman particle tracking according to claim 1, characterized in that: The method for locating the position of the container and the truck bracket using the YOLOv8 segmentation model in step 2 is as follows: Step S201, training a YOLOv8 image segmentation model using a dataset of labeled containers and truck carriages; Step S202: Before image segmentation, the input image is scaled and normalized to meet the input requirements of the model; Step S203: Image segmentation: The preprocessed image is input into the trained YOLOv8 image segmentation model to obtain the category label of each pixel. The image is segmented into different regions according to the category label, and each region corresponds to an object or background. Step S204: extract the target area: Use Gaussian filtering to remove noise from the image segmentation results; use dilation to fill holes generated after threshold processing and connect small areas; after the dilation, use the connectivity analysis algorithm to search the binary image for connected areas of pixels with a value of 1 in the binary image, then select the largest connected area as the target area, and calculate the bounding box of the target area.
Citation Information
Patent Citations
Re-sampling particle filtering algorithm based on Gaussian disturbance
CN106296727A
Container truck anti-hoisting method and system based on machine vision and deep learning
CN113177431A