Omnidirectional mobile AGV deviation correction control method and system based on deep learning

Through the combination of deep learning YOLOv8-Pose algorithm and PID controller, the high-precision and stability correction of AGV trolleys are achieved, solving the problem that AGV trolleys cannot accurately reach the pickup point in the existing technology, and reducing the risk of goods falling.

CN120276429APending Publication Date: 2025-07-08FUJIAN ZKLJAN INTELLIGENT EQUIP ANDTECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510200582.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the deviation correction control, existing AGV cars have problems such as torque mismatch in the drive system, control signal interference and improper controller parameter settings, which leads to the inability to accurately reach the pickup point, and the existing deviation correction methods are insufficient in terms of accuracy and stability.

Method used

The YOLOv8-Pose key point detection algorithm based on deep learning is used to combine with the PID controller to identify the QR code on the cargo pallet through the camera, calculate the position deviation information, and use the McNum wheel kinematic equation to adjust the motion parameters of the AGV trolley to achieve accurate deviation correction.

Benefits of technology

The position and posture correction accuracy of the AGV trolley is improved, the complexity is reduced, the stability and smoothness of correction is enhanced, the risk of goods falling, and the difficulty of control is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276429A_ABST
    Figure CN120276429A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of combination of image processing and AGV control, and provides an omnidirectional mobile AGV deviation correction control method and system based on deep learning, and the method comprises the following steps: S10, enabling an AGV to reach a pickup area; s20, the camera shoots the cargo tray to obtain a detection image containing the QR code; s30, performing position and key point identification on the QR code of the detection image; s40, calculating pose deviation information between the QR code and the standard QR code identification frame; s50, the PID controller compares the pose deviation information with target expectation information; s60, deviation correction is carried out on the AGV trolley; and S70, judging that the QR code is aligned with the standard QR code identification frame, and finishing deviation correction. According to the method, the position and attitude of the AGV are corrected in an image processing mode, meanwhile, the position and attitude calculation capability of the AGV is enhanced, the precision is kept, the complexity is reduced, and the stability and smoothness of correction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing combined with AGV trolley control technology, and in particular to a deep learning-based omnidirectional mobile AGV deviation correction control method and system. Background Art

[0002] With the development and progress of science and technology and the major changes in shopping methods, logistics sorting equipment plays an increasingly important role in the fields of production, sorting, transportation and resource allocation. According to the latest development report on logistics and warehousing equipment, the coverage rate of automated equipment in most sorting warehouses in my country is still very low, and the sorting work is mainly completed by manpower, with labor costs accounting for more than 30% of the total cost. In order to reduce labor costs and improve the space utilization of sorting warehouses, intelligent unmanned warehousing has gradually become the mainstream research direction in the field of logistics. With the deepening of research, Automated Guided Vehicle (AGV) as a new type of automated logistics equipment that can reduce handling time, management and labor costs, has begun to enter the industry's field of vision.

[0003] In recent years, with the continuous advancement of science and technology, AGV guidance methods have also continued to develop. There are many commonly used AGV guidance methods, such as electromagnetic guidance, magnetic tape guidance, QR code guidance, optical guidance and other fixed path guidance, as well as inertial guidance, SLAM guidance, GPS guidance and other free path guidance; different guidance methods can be selected according to the type, structure and application scenario of the AGV itself.

[0004] The QR code guidance is composed of components such as a code reader and a central controller. The central controller processes the code reader data to obtain the AGV position and posture information. This method is one of the common AGV guidance methods on the market and is relatively simple to implement. Although the development of AGV based on QR code navigation has matured, there is still room for improvement in the correction method and the overall system composition. For example, in the actual operation of the AGV, factors such as drive system torque mismatch, control signal interference, and improper controller parameter settings will prevent the AGV itself from arriving at the pickup point in the correct posture. In addition, during the transportation of cargo pallets, if the track brakes, the cargo pallets may also not arrive at the unloading point in the correct posture.

[0005] At present, the main correction method for AGV carts is to use the central camera of the AGV car body to identify the QR code in the middle of the cargo pallet, and assist in adjusting the position and posture of the AGV to ensure that the AGV car can accurately pick up the goods and avoid the risk of goods falling during the pickup and delivery process due to its own factors and improper position of the cargo pallet.

[0006] In the prior art, in deviation correction control, there is a type that only uses a controller for deviation correction, and its deviation correction methods include methods such as PID control, fuzzy control, sliding mode control, optimal control, and genetic algorithm. PID control is easy to implement but may not be sufficient to handle highly complex or nonlinear systems; genetic algorithm and optimal control are suitable for complex optimization problems, but the computational cost may be high; fuzzy control has strong adaptability but may lack precision; sliding mode control has strong robustness but may produce jitter. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide an omnidirectional mobile AGV deviation correction control method and system based on deep learning, which uses image processing to correct the position and attitude of the AGV trolley, and at the same time enhances the calculation ability of the position and attitude of the AGV trolley, maintaining both accuracy and reducing complexity, and improving the stability and smoothness of deviation correction.

[0008] The present invention is implemented by the following technical solutions:

[0009] An omnidirectional mobile AGV deviation correction control method based on deep learning, comprising the following steps:

[0010] S10. The AGV trolley arrives at the picking area. The body of the AGV trolley has a camera, and the goods tray located in the picking area has a QR code.

[0011] S20. The camera captures the goods tray to obtain a detection image containing the QR code, and then sends the detection image to the upper computer.

[0012] S30. The upper computer uses the YOLOv8-Pose key point detection algorithm based on deep learning to identify the position and key points of the QR code in the detection image, obtaining the QR code position and key point information. The upper computer also sets a standard QR code recognition frame at the center of the detection image.

[0013] S40. Using the QR code position and key point information, calculate the pose deviation information between the QR code and the standard QR code recognition frame. The pose deviation information includes a position deviation angle, a position deviation amount, and an attitude deviation angle, and then transmit the pose deviation information to the PID controller.

[0014] S50. The PID controller stores preset target expectation information, and the target expectation information includes that the error distance between the QR code and the standard QR code recognition frame is less than the threshold pixel, and the error angle between the QR code and the standard QR code recognition frame is less than the threshold angle.

[0015] The PID controller compares the pose deviation information with the target expected information. When the pose deviation information does not conform to the target expected information, it goes to S60; when the pose deviation information conforms to the target expected information, it goes to S70;

[0016] S60. The PID controller corrects the deviation of the AGV cart. According to the preset deviation correction control amount and combining with the kinematic equation of the Mecanum wheel, it adjusts the motion parameters of the AGV cart. The motion parameters include the moving direction, moving speed, rotation direction, and rotation speed, so that the AGV cart carrying the camera moves towards the center direction of the QR code, and then goes to S20;

[0017] S70. Determine that the QR code is aligned with the standard QR code recognition frame, and end the deviation correction.

[0018] An omnidirectional mobile AGV control system based on deep learning includes an AGV cart, and the AGV cart executes the above method;

[0019] The AGV cart includes a vehicle body, a camera, an upper board, a main board, a PID controller, a slave board, an MCU, a motor drive module, motors, and Mecanum wheels. The camera is arranged at the center of the top surface of the vehicle body. The main board, PID controller, slave board, MCU, motor drive module, and motors are all arranged inside the vehicle body. The Mecanum wheels are arranged on the chassis of the vehicle body;

[0020] The camera is used to photograph the cargo tray to obtain a detection image containing a QR code;

[0021] The main board is used to receive the detection image and then forward it to the upper computer;

[0022] The upper computer is used to obtain the QR code position key point information from the detection image, set a standard QR code recognition frame at the center of the detection image, and calculate the pose deviation information between the QR code and the standard QR code recognition frame by using the QR code position key point information;

[0023] The PID controller is used to receive the pose deviation information, and then compare the pose deviation information with the preset target expected information to obtain deviation correction control information by computer;

[0024] The slave board is used to receive the deviation correction control information and then forward it to the MCU;

[0025] The MCU is used to generate a PWM signal according to the deviation correction control information;

[0026] The motor drive module is used to adjust the rotational speed of the motor according to the duty cycle of the PWM signal and control the steering of the motor according to the phase of the PWM signal;

[0027] The motor is used to drive the Mecanum wheel, so that the AGV moves straight along the specified path to the target position or performs an action of turning to the specified angle.

[0028] Compared with the background technology, the beneficial technical effects or advantages of the present invention:

[0029] 1. Deploy the YOLOv8-Pose network model in an omnidirectional mobile AGV with Mecanum wheels, use the QR code that identifies the center of the goods tray to position the AGV car, the YOLOv8-Pose network model detects the position and key points of the QR code in the detection image, and gradually adjusts the relative position between the QR code and the standard QR code recognition frame at the center of the detection image. The PID controller drives the Mecanum wheels of the AGV car for position correction, reduces costs and achieves high-precision access to goods, prevents accidents of goods falling during transportation, has broad industrial application prospects, can also reduce the control difficulty, and improves the safety and stability of the logistics system; The present invention corrects the position and attitude of the AGV car by means of image processing, and at the same time enhances the calculation ability of the position and attitude of the AGV car, which not only maintains the accuracy but also reduces the complexity, and improves the stability and smoothness of the correction.

[0030] 2. Introduce the ECA attention mechanism and the Slim-neck architecture into the YOLOv8-Pose network model at the same time to obtain an efficient QR code corner detection algorithm. The ECA attention mechanism can improve the attention to the QR code corners, enhance the feature extraction ability and detection accuracy; The Slim-neck architecture improves the detection ability for QR codes of different scales. Brief Description of the Drawings

[0031] The following further describes the present invention with reference to the accompanying drawings in conjunction with the embodiments.

[0032] Figure 1 It is the flowchart of the AGV car deviation correction in the embodiment of the present invention.

[0033] Figure 2 It is the comparison diagram of the pose deviation of the AGV car in the embodiment of the present invention.

[0034] Figure 3 It is the comparison diagram of the pose deviation between the QR code and the standard QR code recognition frame in the embodiment of the present invention.

[0035] Figure 4 It is the schematic diagram of using Labelme image annotation in the embodiment of the present invention.

[0036] Figure 5 This is the flowchart of model training in the embodiment of the present invention.

[0037] Figure 6 This is the structural diagram of the YOLOv8n-Pose network model in the embodiment of the present invention.

[0038] Figure 7 This is the structural diagram of the ES-YOLO network model in the embodiment of the present invention

[0039] Figure 8 This is the schematic diagram of the ECA attention mechanism module in the embodiment of the present invention.

[0040] Figure 9 This is the schematic diagram of the GSConv module in the embodiment of the present invention

[0041] Figure 10 This is the schematic diagram of the VoV-GSCSP module in the embodiment of the present invention.

[0042] Figure 11 This is the schematic diagram of Mosaic data augmentation in the embodiment of the present invention

[0043] Figure 12 This is the comparison diagram of the QR code detection effect of YOLOv8n-Pose in the embodiment of the present invention.

[0044] Figure 13 This is the definition diagram of the pose deviation parameters of the AGV trolley based on the QR code in the embodiment of the present invention.

[0045] Figure 14 This is the schematic diagram of the deviation situation of the AGV trolley in the embodiment of the present invention.

[0046] Figure 15 This is the schematic diagram of the motion analysis of the Mecanum wheel in the embodiment of the present invention.

[0047] Figure 16 This is the schematic diagram of the hardware framework of the AGV trolley in the embodiment of the present invention.

[0048] Figure 17 This is the structural diagram of the AGV trolley in the embodiment of the present invention. Detailed implementation manners

[0049] Refer to Figures 1 to 17 , the preferred embodiment of the present invention.

[0050] An omnidirectional mobile AGV deviation correction control method based on deep learning includes the following steps:

[0051] S10. The AGV trolley arrives at the picking area. The body of the AGV trolley is equipped with a camera, and the goods tray located in the picking area has a QR code;

[0052] S20. The camera captures the goods tray to obtain a detection image containing a QR code, and then sends the detection image to the host computer;

[0053] S30. The host computer uses the YOLOv8-Pose key point detection algorithm based on deep learning to identify the position and key points of the QR code in the detection image, obtaining the QR code position and key point information. The host computer also sets a standard QR code recognition frame at the center of the detection image;

[0054] S40. Using the QR code position and key point information, calculate the pose deviation information between the QR code and the standard QR code recognition frame. The pose deviation information includes a position deviation angle, a position deviation amount, and an attitude deviation angle, and then transmit the pose deviation information to the PID controller;

[0055] S50. The PID controller stores preset target expectation information, where the target expectation information includes that the error distance between the QR code and the standard QR code recognition frame is less than a threshold pixel, and the error angle between the QR code and the standard QR code recognition frame is less than a threshold angle;

[0056] Among them, the threshold pixel is 25 pixels, and the threshold angle is 2°.

[0057] The PID controller compares the pose deviation information with the target expectation information. When the pose deviation information does not meet the target expectation information, go to S60; when the pose deviation information meets the target expectation information, go to S70;

[0058] S60. The PID controller corrects the deviation of the AGV cart. According to the preset deviation correction control amount, combined with the kinematic equation of the Mecanum wheel, adjust the motion parameters of the AGV cart. The motion parameters include the moving direction, moving speed, rotation direction, and rotation speed, so that the AGV cart carrying the camera moves in the direction of the center of the QR code, and go to S20;

[0059] S70. Determine that the QR code is aligned with the standard QR code recognition frame, and end the deviation correction.

[0060] After ending the deviation correction, the QR code at the center of the goods tray is accurately located at the central position of the current image captured by the camera, so that the AGV cart is accurately positioned directly below the center of the goods tray. The AGV cart picks up the goods tray to prevent accidents of goods falling during the transportation of the AGV cart.

[0061] Detailed description of the method of the present invention:

[0062] (1) The AGV cart arrives at the goods picking area, takes a picture of the bottom surface of the goods tray through a camera to obtain a detection image containing a QR code, and uses the YOLOv8n-Pose key point detection algorithm based on deep learning to identify the position and key points of the QR code located at the center of the bottom surface of the goods tray in the detection image.

[0063] (2) Using the position information and key point information of the QR code in the detection image obtained in (1), applying mathematical formulas and combining with the actual length size of the QR code in pixels, calculate the position deviation angle, position deviation amount, and attitude deviation angle of the QR code at the center of the goods tray deviating from the standard QR code recognition frame located at the center of the detection image, and transmit the calculated position deviation angle, position deviation amount, and attitude deviation angle as pose deviation information to the PID controller.

[0064] (3) After receiving the pose deviation information of the AGV cart, the PID controller compares this pose deviation information with the preset target expectation information to determine the deviation between the current position and attitude and the target expectation position and attitude. Based on this deviation, the PID controller calculates an appropriate control amount to adjust the moving direction, moving speed, rotating direction, and rotating speed of the AGV cart.

[0065] (4) Using the PID control amount obtained in (3), decompose the moving direction and moving speed of the AGV cart into the independent rotational speeds of the four wheels according to the Mecanum wheel kinematic equation, so that the AGV cart corrects the deviation in displacement. When the error distance between the QR code at the center of the goods tray and the standard QR code recognition frame at the center of the detection image is less than 25 pixels, the correction of the AGV cart's displacement ends. Then, decompose the rotating direction and rotating speed of the AGV cart into the independent rotational speeds of the four wheels according to the Mecanum wheel kinematic equation, so that the AGV cart corrects the deviation in attitude. When the error angle between the QR code at the center of the goods tray and the standard QR code recognition frame at the center of the detection image is less than 2°, the correction of the AGV cart's attitude ends.

[0066] (5) During the correction process of the AGV cart in (1) to (4), through the PC host computer, a detection image containing the QR code at the center of the goods tray can be received in real time. By correcting each consecutive frame captured by the camera, the AGV cart can continuously center the QR code at the center of the goods tray, reduce the deviation between the QR code at the center of the goods tray and the standard QR code recognition frame at the center of the detection image until the QR code at the center of the goods tray is aligned with the standard QR code recognition code at the center of the detection image, and ensure that the QR code at the center of the goods tray is accurately located at the central position of the detection image in the current frame captured by the camera.

[0067] As Figure 2As shown in the figure, it is a comparison diagram of the pose deviation of the AGV trolley. The comparison diagram of the pose deviation of the AGV trolley randomly places the AGV trolley at any position under the cargo tray, and a total of 50 groups of displacement deviations and angle deviations of the AGV trolley before and after deviation correction are measured and recorded. Figure 2 On the left side is the comparison diagram of displacement deviation, and on the right side is the comparison diagram of angle deviation. In the comparison diagram of displacement deviation, values greater than zero indicate that the AGV body (camera) is in front of the QR code, values less than zero indicate that it is behind the QR code, and values equal to zero indicate that the AGV is directly below the QR code. In the comparison diagram of angle deviation, values greater than zero indicate that the front of the AGV trolley is deviated to the left, values less than zero indicate that the front of the AGV trolley is deviated to the right, and values equal to zero indicate that the front is facing straight. It can be seen from the comparison diagrams of displacement deviation and angle deviation that after the deviation correction control of the AGV trolley, the running accuracy of the AGV trolley is improved, and both the displacement deviation and the angle deviation show good convergence. The displacement deviation between the AGV trolley and the QR code can be stably maintained within 1.5 cm, and the angle deviation between the AGV trolley and the QR code can be stably maintained within 2°.

[0068] As Figure 3 shown, it is a comparison diagram of the pose deviation between the QR code and the standard QR code recognition frame obtained by photographing the cargo tray with the 26th group of AGV trolleys selected from 50 groups of measurement data. On the left side of the figure is the relative position between the standard QR code recognition frame of the pre-deviation correction detection image and the cargo tray QR code, and on the right side is the relative position between the standard QR code recognition frame of the post-deviation correction detection image and the cargo tray QR code. It can be seen from the figure that after the deviation correction control, the pose of the AGV is effectively corrected, which proves the feasibility and practicability of the deviation correction method.

[0069] The present invention only needs to paste a QR code with cargo information at the center of the cargo tray. When the AGV trolley reaches the designated picking area in the warehouse, it identifies the QR code at the center of the cargo tray through the YOLOv8n-Pose network model to determine the pose information of the AGV trolley, and uses the obtained pose information for auxiliary deviation correction. The present invention enables the AGV trolley to be accurately positioned directly below the center of the cargo tray, greatly reducing the risk of the cargo falling during its transportation of the cargo, reducing the labor cost and operation cost, and improving the management efficiency.

[0070] Furthermore, in the step S30, the YOLOv8n-Pose key point detection algorithm based on deep learning specifically includes the following steps:

[0071] S31. Obtain the QR code data set and data augmentation: Select the publicly available QR code data set provided by the Baidu PaddlePaddle AI Studio platform, and at the same time preprocess the images in the data set, including rotation, cropping, adding noise, adjusting brightness and contrast, to increase the quantity and complexity of the data set pictures;

[0072] S32, image annotation and data set division: LabelMe software is used to annotate the data set after data enhancement, marking the positions of the QR code bounding boxes and key points, and dividing it into training set, validation set and test set in a ratio of 8:1:1;

[0073] S33. Build the YOLOv8-Pose-ES network model: Based on the YOLOv8-Pose network model, add the ECA attention mechanism module to the backbone network, and use the VoV-GSCSP module and GSConv module in the neck network to replace the C2f module and Conv module respectively, forming a Slim-neck architecture.

[0074] S34, training model and parameter adjustment optimization model: first use the Mosaic data enhancement method on the data set of S32, then select four images in the training set for splicing, and then use the data set that has been enhanced by the Mosaic data enhancement method to train the YOLOv8-Pose-ES network model. The loss function is calculated and the model parameters are updated in each iteration to minimize the value of the loss function until the model converges, and finally a PT format model is obtained;

[0075] S35, model reasoning acceleration: convert the PT format model into an ONNX format model, use the ONNX format model to detect the images in the test set, and verify the performance of the ONNX format model.

[0076] Beneficial effects of this technical solution: Based on the YOLOv8-Pose network, the ECA attention mechanism module is added to the backbone network to enhance the network's feature extraction capability and detection accuracy. The VoV-GSCSP module and the GSConv module are used in the neck network to replace the C2f module and the Conv module, respectively, thereby forming a Slim-neck lightweight architecture to reduce the complexity of the model. Finally, an improved YOLOv8 network model is obtained.

[0077] Four images in the training set are selected for stitching to increase the number of detection targets and enrich the background information of the detection objects.

[0078] Convert the PT format model generated by training to an ONNX format model to accelerate reasoning.

[0079] like Figure 6 As shown in the figure, the YOLOv8n-Pose network model structure diagram. The YOLOv8n-Pose network model structure mainly includes four parts: input part, backbone network, neck network and output part.

[0080] The input part is responsible for scaling the input image to the size required for training.

[0081] The backbone network adopts the CSPNet structure, which divides the gradient flow and makes it propagate along different network paths, effectively solving the problem of gradient disappearance and reducing the computational cost; the C2f module is used to learn features, obtaining richer gradient flow information while ensuring a lightweight effect; the SPFF structure is adopted to improve the model inference speed. The backbone network effectively extracts features and generates feature maps of three different scales.

[0082] The neck network adopts a structure that combines the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN), effectively integrating the top-down and bottom-up information flows in the network, providing more semantic information and location information, and enhancing the detection performance of objects of different sizes. The neck network aggregates different feature parameters extracted by the backbone network to enhance the hierarchical structure of the feature maps, thereby improving the feature fusion ability of the network.

[0083] The output part uses feature maps of different scales to identify the category, location, and key point information of objects of different sizes, replaces the coupled head with a decoupled head, and directly predicts the center of the object using Anchor-free when outputting the predicted object location, reducing the complexity and dependence on the size and shape of the predefined anchor boxes, and at the same time accelerating the model convergence speed and inference speed.

[0084] The converted ONNX format model is used to detect the images in the test set to verify its model performance.

[0085] IOU is used to calculate the AP for keypoint detection, and OKS is used to calculate the AP for keypoint detection. When IOU or OKS exceeds the threshold T, the detection result is considered correct; when IOU or OKS is less than the threshold T, the detection result is considered incorrect.

[0086] Precision reflects the ability of the model to identify targets, that is, among the targets correctly identified by the model as a certain category, the proportion of the targets that truly belong to that category.

[0087] Recall measures the comprehensiveness of the model in finding targets, that is, the proportion of all true targets found by the model among all actual existing targets.

[0088]

[0089] Among them, TP represents positive samples and is correctly identified; FP represents negative samples that are incorrectly identified as positive samples; FN represents positive samples that are incorrectly identified as negative samples.

[0090] Calculating the area under the P-R curve yields AP. Among them, the bounding box mAP50 is the average value of AP measured for all classes when IOU = 0.50, and the bounding box mAP50-95 is the average value of the AP means for all classes when OKS = 0.50, 0.55, …, 0.95 (step size 0.05).

[0091] Calculating the area under the P-R curve yields AP. Among them, the keypoint mAP50 is the average value of AP measured for all classes when OKS = 0.50, and the keypoint mAP50-95 is the average value of the AP means for all classes when OKS = 0.50, 0.55, …, 0.95 (step size 0.05).

[0092]

[0093] The model's computational cost, that is, the number of floating-point operations (FLOPs), represents the total number of floating-point operations required for the model to run once. FLOPs is a key metric for measuring the computational complexity and inference efficiency of the model.

[0094] The number of model parameters (Parameters) refers to the total number of parameters that the model needs to learn and adjust, and is an important metric for evaluating the complexity and capacity of the model.

[0095] The number of frames per second (FPS) processed is an important metric for measuring the detection speed of the model in a video stream or image sequence. This metric is based on the computational results of a single-sample batch on the test set in the CPU and GPU environments.

[0096]

[0097] Among them, T preprocess is the image preprocessing time, T inference is the time during model inference, and T postprocess is the image postprocessing time.

[0098] A higher mAP value represents a better detection effect. The experimental environment of this invention is 64-bit Windows 10, using the Pytorch 1.12.0 deep learning framework and CUDA 11.6, and the programming language is Python 3.8.18. The hardware configuration includes an AMD Ryzen 7 5800H CPU (main frequency 3.20GHz) and an NVIDIA GeForce RTX 3060 GPU (video memory 12GB). The same hyperparameters are used for the algorithms before and after improvement, and the input image size is 640×640. When optimizing the neural network, the Stochastic Gradient Descent (SGD) optimizer is used, and the learning rate is adjusted through the cosine annealing strategy. The initial learning rate is 0.01, the weight decay rate is 0.0005, the momentum parameter is 0.937, the number of training epochs is 300, and the batch size is 8.

[0099] To better verify the improvement of the ECA attention mechanism and the lightweight Slim-neck architecture on the performance of the improved YOLOv8n-Pose network and evaluate the effectiveness of each component, ablation experiments are conducted for comparison, and the experimental results are shown in Table 1-2.

[0100] Table 1 Influence of different improvement strategies on the detection accuracy of the model

[0101]

[0102] Table 2 Influence of different improvement strategies on the detection performance of the model

[0103]

[0104] The experimental results show that after introducing different modules into the original network, various indicators of the model have been improved. Especially when the ECA attention mechanism and the Slim-neck architecture are introduced simultaneously, the improved network has increased by 0.3% and 1.5% respectively on mAP50 and mAP50-95 for boxes, reaching 99.5% and 88.4%; on mAP50 and mAP50-95 for poses, it has increased by 1.6% and 1.1% respectively, reaching 97.4% and 93.3%. In addition, the detection speeds on the CPU and GPU have increased by 0.3 f / s and 0.7 f / s respectively, reaching 14.2 f / s and 59.6 f / s. It proves that the Slim-neck architecture significantly enhances the feature extraction ability of the ECA attention mechanism, making the improved network not only meet the lightweight requirements but also have advantages in detection accuracy and speed, and can effectively meet the needs of detection tasks.

[0105] To further compare the improvement effect of this paper on the network, the algorithms before and after improvement are compared with the current mainstream key point detection algorithms, and the experimental results are shown in Table 3-4.

[0106] Table 3 Comparative Experiments of Different Mainstream Models (Model Performance)

[0107]

[0108] Table 4 Comparative Experiments of Different Mainstream Models (Model Performance)

[0109]

[0110] The experimental results show that the improved algorithm is superior to other algorithms in the comprehensive evaluation of detection accuracy and speed, especially in terms of the number of model parameters, computational complexity, and detection speed. Although its detection accuracy is slightly lower than that of the YOLOv7-w6-Pose algorithm, it has significant advantages in lightweight and detection speed. In summary, under the conditions of extremely few parameters and fast detection speed, the improved network still ensures higher detection accuracy, showing that it is more suitable for the QR code detection task compared to other algorithms.

[0111] To demonstrate the performance of the improved model in detecting QR codes and their corner points, three QR code images are selected to test YOLOv8n-Pose and ES-YOLO. In Figure 12 , on the left is the effect of YOLOv8n-Pose detecting QR codes, and on the right is the effect of the ES-YOLO of the present invention detecting QR codes.

[0112] As Figure 12 shown, the comparative diagram of the detection effect of YOLOv8n-Pose on QR codes is elaborated. It can be seen from the figure that the improvement effect of ES-YOLO is significantly better than that of the YOLOv8-Pos model. In the QR code detection scenario, due to the deformation of the QR code, the accuracy of corner point positioning is affected, and there are also certain errors in the border positioning. The improved model is closer to the real position in QR code corner point positioning, and the detected border is also more in line with the real QR code, which can include the QR code content to the greatest extent, and the edge deviation is relatively reduced. These detection results further verify the applicability of the improved algorithm in the QR code detection task.

[0113] Furthermore, in S33, the backbone grid uses the CSP-Darknet53 structure, and the ECA attention mechanism module is embedded in the layer before the SPPF module of the backbone network;

[0114] The steps to obtain the ECA attention mechanism module are as follows:

[0115] S33-1. Through global average pooling (GAP), the two-dimensional feature map with input size H×W in each channel is converted into a scalar, simplifying the spatial information and retaining the global features. The specific calculation formula is as follows:

[0116]

[0117] Among them, y is the global feature vector, and x i is the i-th feature map with an input size of H×W;

[0118] S33-2. The one-dimensional convolutional layer Conv1d performs a linear transformation on the global feature vector y output by the global average pooling GAP to learn the interaction weights between different channels and capture the dependencies between channels. The specific calculation formula is as follows:

[0119] w = σ(C1D k (y)) Formula 2;

[0120] Among them, σ is the Sigmoid activation function, w is the channel weight, and y is the global feature vector generated by GAP;

[0121] S33-3. The value of k is adaptively determined by the mapping of the channel dimension C to adjust the interaction range between channels and optimize the feature extraction process. The specific calculation formula is as follows:

[0122]

[0123] Among them, γ and b are adjustment parameters, with values of 2 and 1, k is the convolutional kernel size, and |t|odd represents the odd number closest to t.

[0124] The beneficial effects of this technical solution: Through the global average pooling GAP, the global feature information of each channel can be effectively obtained to capture the dependencies between channels in the next step.

[0125] Through the local receptive field of the one-dimensional convolution, the mutual relationship between channels can be captured more efficiently, and then these weights are mapped to the range of 0 to 1 through the Sigmoid activation function to generate channel attention weights, and finally multiplied by the original feature layer to achieve the extraction of important features.

[0126] To improve the local interaction efficiency, the value of k is adaptively determined by the mapping of the channel dimension C, thereby adjusting the interaction range between channels and optimizing the feature extraction process.

[0127] Through the mapping relationship ψ, high-dimensional channels have a larger interaction range, while low-dimensional channels have a smaller interaction range through the use of non-linear mapping; therefore, it can adaptively enable different-sized channel feature maps to obtain cross-channel information interactions suitable for the channel size of the feature map.

[0128] Furthermore, in the S33, the neck network is a structure that combines the feature pyramid network FPN with the path aggregation network PAN. The GSBottleneck module is designed based on the GSConv module and the Conv module. The VoV-GSCSP module is designed based on the GSBottleneck module using the one-shot aggregation method. VoV-GSCSP and GSConv are used to replace the C2f module and Conv module in the neck network respectively to form a Slim-neck architecture.

[0129] Beneficial effects of this technical solution: The GSConv module is obtained by specifically including the following steps: first, cross-channel information interaction is achieved through standard convolution; then, group convolution is used to improve the computational efficiency of the structure; finally, the feature maps generated by the two convolution operations are cascaded, and the feature channels are rearranged through channel shuffling, thereby improving the information flow between features and outputting the enhanced feature map. This method reduces redundant information while achieving lightweight network and effectively improving detection accuracy.

[0130] In order to speed up the prediction calculation speed, the images in the CNN backbone network usually need to be gradually converted from spatial information to channel information. Each spatial compression and channel expansion may cause some semantic information to be lost. Although dense convolution maximizes the retention of hidden connections between channels, sparse convolution cuts off these connections. GSConv reduces the adverse effects of DSC's defects on the model by retaining channel connections as much as possible, while making full use of DSC's advantages. Specifically, GSConv maintains the transmission of semantic information during spatial compression and channel expansion, and effectively avoids the problem of decreased feature extraction efficiency that may be caused by dense convolution by maintaining the connection between channels.

[0131] like Figure 10 As shown in the figure, the schematic diagram of the VoV-GSCSP module. The design of the VoV-GSCSP module is as follows: The GSBottleneck module reduces the complexity and parameter amount of the module through GSConv, and enhances the feature transfer through residual connection and feature fusion, optimizes the gradient flow, and avoids the loss of information during training. On this basis, a module called VoV-GSCSP is designed by combining the GSBottleneck module with the OSA method. First, the input feature map is divided into two parts, one part is subjected to feature extraction through Conv, and the other part is subjected to more refined feature extraction through the GSBottleneck module. Finally, the two parts of the feature map are cascaded through the Conv operation to realize cross-channel information interaction, thereby enhancing the feature extraction capability of the network.

[0132] Reconstruct the neck network of YOLOv8n-Pose using lightweight convolutional GSConv and VoV-GSCSP modules to form an architecture called Slim-neck. First, fuse feature maps of different scales through upsampling and concatenation operations; then, use GSConv to achieve lightweight and efficient feature extraction; subsequently, further perform feature extraction and fusion through the VoV-GSCSP module. Finally, use the Pose branch for object detection. Specifically, the Slim-neck architecture can extract more refined features, effectively prevent network degradation, and improve the detection ability of the model.

[0133] Furthermore, after the neck network extracts three feature maps, three detection modules of different scales are respectively connected to the three feature maps. Each detection module serves as a sliding window detector for QR code classification, bounding box, and key point regression on each layer of the feature pyramid structure.

[0134] Furthermore, in S34, the loss function includes a class loss function, a bounding box loss function, and a key point loss function;

[0135] The class loss function uses the BCE loss function. The class loss function is used to determine whether the predicted target box encloses the correct class. The calculation formula is as follows:

[0136]

[0137] where N is the total number of samples, p i is the prediction probability of the model for the i-th sample, and the value range is 0 - 1; y i is the true label of the i-th sample, with the positive class taking the value 1 and the negative class taking the value 0;

[0138] The bounding box loss function uses the CIOU loss function and the DFL loss function. The bounding box loss function is used to determine whether the predicted target box can tightly enclose the target object. The calculation formula is as follows:

[0139]

[0140] where a is:

[0141]

[0142] where v is

[0143]

[0144] where, w and h are the width and height of the true box, w gt and h gtis the width and height of the prediction box, c represents the diagonal length of the minimum bounding rectangle of the candidate box and the ground truth box, a is the weight function, v is used to measure the consistency of the aspect ratio, and ρ(b, b gt ) is the square of the Euclidean distance between the center point b of the prediction box and the center point b gt of the ground truth box;

[0145] DFL(S i , S i+1 ) = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 )) Formula VIII;

[0146] where S is the probability output value; it can be seen from Formula VIII that when y and y i+1 are very close and the S probability output is large, DFL is small, making the distribution approach the annotation center;

[0147] The key point loss function adopts the OKS loss function and the BCE loss function. The key point loss function is used to judge whether the predicted key points exist and can closely approach the target key points. The calculation formula is as follows:

[0148]

[0149] where p i represents the i-th key point of the p-th target, d pi represents the Euler distance between the key point prediction value and the true value, s p represents the scale factor of the target, v pi represents the visibility of the key point, σ i represents the standard deviation between the i-th key point annotation value and the true value, δ is the visible point calculation function, and OKS is used to measure the similarity between the predicted key points and the true key points, with a value range of 0 - 1.

[0150] The beneficial effect of this technical solution: The class loss adopts the BCE loss function. The class loss function is used to judge whether the predicted target box encloses the correct class. q is the predicted target score. If it is the true class, then q is the loU of the prediction and the true value; if it is other classes, q is 0.

[0151] The DFL loss function can make the network focus on the values near the target y faster, increase their probabilities, and accelerate convergence.

[0152] OKS is used to measure the similarity between the predicted key points and the true key points, with a value range of 0 - 1. The closer the value is to 1, the more similar the model's predicted key points are to the true annotation.

[0153] Further, in S40, the pose deviation information between the QR code and the standard QR code recognition frame is calculated, specifically as follows:

[0154] S41. Use a wide-angle camera to collect a detection image of the bottom surface of the goods pallet. The size of the detection image is set to W×H. Taking the center coordinates (W / 2, H / 2) of the detection image as the reference point, determine the center point (x img , y img ) of the standard QR code recognition frame at the center of the detection image, and determine the position of the standard QR code recognition frame at the center of the detection image according to the representation of the actual length of the QR code on the goods pallet in pixels;

[0155] S42. Use the YOLOv8-Pose key point detection algorithm to identify the detection image, locate the QR code bounding box and four key points; starting from the lower left corner of the QR code at the center of the standard goods pallet in the clockwise direction, mark the corner points respectively. The lower left corner is the first corner point, the upper left corner is the second corner point, the upper right corner is the third corner point, and the lower right corner is the fourth corner point; calculate the center point (x qr , y qr ) of the QR code at the center of the goods pallet according to the average value of the coordinates of the four corner points of the QR code at the center of the goods pallet. The specific calculation formula is as follows:

[0156]

[0157] where (x qr , y qr ) is the center point coordinate of the QR code at the center of the goods pallet, and (x1, y1), (x2, y2), (x3, y3), and (x4, y4) are the coordinates of the lower left corner, upper left corner, upper right corner, and lower right corner of the QR code at the center of the goods pallet;

[0158] S43. Construct a Cartesian coordinate system with the center point of the standard QR code recognition frame as the origin, define the positive direction of the X-axis as 0°, and the increasing direction of the angle as clockwise; use the center point (x img , y img ) of the standard QR code recognition frame at the center of the detection image and the center point (x qr , y qr ) of the QR code at the center of the goods pallet to calculate the angle between the line connecting these two points and the horizontal rightward line. This angle is the position deviation angle, and the position deviation angle represents the orientation of the center of the goods pallet relative to the AGV cart, where the front corresponds to 90°, the rear corresponds to 270°, the right corresponds to 0°, and the left corresponds to 180°; at the same time, calculate the center point (x0, y0) of the standard QR code recognition frame at the center of the detection image and the center point (x qr , y qr) The straight-line distance between them, which is the position deviation amount. The positive or negative of the position deviation amount is determined by the value range of the position deviation angle. When the position deviation angle is between 0° and 180°, it means that the center of the goods tray is in front of the AGV vehicle, and at this time the position deviation amount is negative; when the position deviation angle is between 180° and 360°, it means that the center of the goods tray is behind the AGV vehicle, and at this time the position deviation amount is positive. The specific calculation formulas for the position deviation angle θ and the position deviation amount d are as follows:

[0159] θ rad = arctan2(y qr -y img , x qr -x img ) Formula Eleven;

[0160]

[0161] Where, (x qr , y qr ) is the center point coordinates of the QR code of the goods tray center, (x img , y img ) is the center point coordinates of the standard QR code recognition frame of the detection image center, θ rad is the radian between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (-π, π), and θ deg is the angle between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (0°, 360°);

[0162] S44: Construct a Cartesian coordinate system with the center point (x qr , y qr ) of the QR code of the goods tray center as the origin. The forward direction of the AGV vehicle is 90 degrees. The angle value taken from the positive Y-axis direction and rotated clockwise to the negative Y-axis direction is positive, while the angle value taken from the positive Y-axis direction and rotated counterclockwise to the negative Y-axis direction is negative; use the upper left corner coordinates (x2, y2) and the upper right corner coordinates (x3, y3) of the QR code of the goods tray center to calculate the midpoint (x m , y m ) between the two points, and calculate the included angle between the line connecting the two points (x qr , y qr ) and (x m , y m ) and the vertically upward straight line. This included angle is the attitude deviation angle; the attitude deviation angle The specific calculation formula is as follows:

[0163]

[0164] Among them, (x qr , y qr ) is the center point coordinate of the QR code at the center of the goods pallet, (x m , y m ) is the midpoint coordinate of the left and right corner points of the QR code at the center of the goods pallet, (x2, y2) and (x3, y3) are the upper left and upper right coordinates of the QR code at the center of the goods pallet respectively. is the radian between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (-π, π). is the angle between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (-π, π).

[0165] As Figure 13 shown, the AGV car pose deviation parameter definition diagram based on the QR code. The AGV pose deviation parameter definition diagram is the definition diagram of the AGV car position deviation amount, position deviation angle, and attitude deflection angle. As Figure 14 shown, the AGV car offset situation diagram. The AGV car offset situation diagram is the position and attitude of the AGV relative to the QR code at the center of the goods pallet during the deviation correction process.

[0166] Furthermore, in S50, the PID controller compares the pose deviation information with the preset target expectation information, specifically:

[0167] The pose deviation information is input into the PID controller. Through the adjustment of the proportional term P, integral term I, and derivative term D of the PID controller, a control output signal is finally generated.

[0168] The calculation formula for the output of the PID controller is as follows:

[0169]

[0170] Among them, K p , K i , K d are the gain coefficients of proportional, integral, and derivative respectively, T i is the integral time constant, T d is the derivative time constant, e(t) is the deviation signal, and u(t) is the output of the controller.

[0171] Beneficial effects of this technical solution: By obtaining the real-time status signal of the target system, comparing it with the preset expected value, calculating the deviation value between the two, inputting the deviation signal into the PID controller, and adjusting through the proportional term P, integral term I, and derivative term D of the PID controller, finally generating a control output signal. Among them, the proportional control part generates a control signal proportional to the error according to the magnitude of the error; the integral control part adjusts by accumulating historical errors to reduce the steady-state error of the system; the derivative control part corrects by predicting the trend of error change, thereby improving the response speed and stability of the system.

[0172] Further, in S60, the kinematic equation of the Mecanum wheel is specifically as follows:

[0173] S61. The center of mass of the body of the AGV and the four Mecanum wheels is regarded as a rigid body, and there is sufficient friction between the rollers of the Mecanum wheels and the ground. Analyze the right front wheel in the top view of the AGV, name it Mecanum wheel B, establish a robot coordinate system with the center position of the AGV as point O, stipulate that the forward direction is the positive direction of the X-axis, that is, 90° in the Cartesian coordinate system, the leftward direction is the positive direction of the Y-axis, that is, 180° in the Cartesian coordinate system, the upward direction is the positive direction of the robot Z-axis, and the counterclockwise direction around the Z-axis is the positive rotation direction; decompose the overall speed of the vehicle into the speeds of the center of mass of Mecanum wheel B moving along the X and Y directions. The center-of-mass speed of Mecanum wheel B is V B_X and V B_y with the following calculation formulas:

[0174]

[0175] Among them, V B_X is the left-right moving speed of the center of mass of Mecanum wheel B, and the rightward movement is positive; V B_y is the front-back moving speed of the center of mass of Mecanum wheel B, and the forward movement is positive; V x is the front-back moving speed of the AGV, and the forward movement is positive; V y is the left-right moving speed of the AGV, and the forward movement is positive; ω z is the rotational speed of the AGV around point O, and the counterclockwise direction is positive; r x is half of the wheelbase between the two side wheels, and r y is half of the wheelbase between the front and rear wheels;

[0176] S62. The V B_X and V B_y of the center-of-mass speed of the Mecanum wheel B are also generated by the combination of V B轮 and V B辊 with the following calculation formulas:

[0177] V B-x = VB轮 +V B辊 ×sinβ Formula 21;

[0178] V B-y =-V B辊 ×cosβ Formula 22;

[0179] where V B轮 is the linear velocity of Mecanum wheel B, with forward being positive; V B辊 is the linear velocity of the roller in contact with the ground of Mecanum wheel B that generates relative sliding;

[0180] S63. By combining Formulas 20 to 23 of S61 and S62, we can obtain:

[0181]

[0182] where β is the angle between the wheel axis and the roller axis;

[0183] S64. The axial direction of each roller of the Mecanum wheel forms a 45° angle with the wheel axis,

[0184] Substituting β = 45° and into Formulas 24 and 25, we can obtain:

[0185] V B轮 =V x +V y +ω z ×L×(sin(α B )+cos(α B )) Formula 25;

[0186] S65. Substituting L×sin(α B ) = r y and L×cos(α B ) = r x into Formula 26, we can obtain:

[0187] V B轮 =V x +V y +ω z ×(r x +r y ) Formula 26;

[0188] S66: Repeat S61 to S65 to derive V A轮 、V C轮 、V D轮 wheels and V x 、V y 、ω zThe linear velocity direction of the rollers of McReel A, McReel C, and McReel D in contact with the ground is related to the left-right moving speed of the center of mass of McReel B, and V A-y =V A辊 ×cosβ、V C-y =V C辊 ×cosβ、V D-y =-V D辊 ×cosβ, according to the inverse solution formula of Mecanum wheel kinematics, the target linear velocity or target angular velocity of the four wheels can be calculated from the moving speed and rotation speed of the omnidirectional mobile AGV, as follows:

[0189]

[0190] V A轮 =R×ω A 、V B轮 =R×ω B 、V C轮 =R×ω C 、V D轮 =R×ω D Substituting the above formula into the inverse kinematics equation of the Mecanum wheel, we can get:

[0191]

[0192] Where R is the radius of the wheel; ω B is the angular velocity of the hub of McDonnell Douglas B, with counterclockwise rotation being positive;

[0193] S67. The direct solution formula of Mecanum wheel kinematics is obtained by combining the target linear velocity or target angular velocity of the four wheels to calculate the moving speed and rotation speed of the omnidirectional mobile AGV, as follows:

[0194] V A轮 =R×ω A 、V B轮 =R×ω B 、V C轮 =R×ω C 、V D轮 =R×ω D Substituting the above formula into the equation, we can get the forward kinematic equation of the Mecanum wheel:

[0195]

[0196] like Figure 15 As shown, the Mecanum wheel motion analysis diagram is a plan view obtained by looking down at the AGV vehicle and performing the Mecanum wheel motion analysis of steps S61-S67 on it.

[0197] An omnidirectional mobile AGV control system based on deep learning, including an AGV cart, and the AGV cart executes the above method;

[0198] The AGV cart includes a vehicle body, a camera, an upper board, a main board, a PID controller, a slave board, an MCU, a motor drive module, motors, and Mecanum wheels. The camera is arranged at the center of the top surface of the vehicle body. The main board, the PID controller, the slave board, the MCU, the motor drive module, and the motors are all arranged inside the vehicle body. The Mecanum wheels are arranged on the chassis of the vehicle body;

[0199] The camera is used to photograph the goods tray to obtain a detection image containing a QR code;

[0200] The main board is used to receive the detection image and then forward it to the host computer;

[0201] The host computer is used to obtain the QR code position key point information from the detection image, set a standard QR code recognition frame at the center of the detection image, and use the QR code position key point information to calculate the pose deviation information between the QR code and the standard QR code recognition frame;

[0202] The PID controller is used to receive the pose deviation information, then compare the pose deviation information with the preset target expectation information, and the computer obtains the deviation correction control information;

[0203] The slave board is used to receive the deviation correction control information and then forward it to the MCU;

[0204] The MCU is used to generate a PWM signal according to the deviation correction control information;

[0205] The motor drive module is used to adjust the speed of the motor according to the duty cycle of the PWM signal and control the steering of the motor according to the phase of the PWM signal;

[0206] The motors are used to drive the Mecanum wheels to make the AG cart go straight along the specified path to the target position or perform an action of turning to the specified angle.

[0207] Detailed description of the system of the present invention:

[0208] The AGV vehicle has image acquisition and image processing functions. The main board collects video frames through an externally connected wide-angle camera, uses the YOLOv8-Pose key point detection algorithm of deep learning to calculate and infer the pose deviation information of the AGV vehicle in real time, and uses a PID controller to calculate and output motion parameters such as the moving direction, moving speed, rotation direction, and rotation speed of the AGV vehicle. The main board packages and sends the control instruction data to the slave board through the I2C communication method. The main board packages and sends this feedback information to the host computer through wireless communication for visual display of the AGV operation status.

[0209] After receiving the control data packet, the MCU of the slave board is responsible for generating the corresponding PWM signal according to the received data packet. The motor drive module changes the duty cycle of the PWM signal to adjust the motor speed and changes the phase of the PWM signal to control the motor steering. The AGV vehicle chassis is controlled by the motor, so that the AGV vehicle goes straight along the specified path to the target position or performs a turning action to a specified angle, thereby making real-time corrections to the lateral deviation, longitudinal deviation, and angle deviation that occur during operation.

[0210] As Figures 16 to 17 shown, the hardware framework diagram of the AGV vehicle and the schematic diagram of the AGV vehicle. In the hardware framework diagram of the AGV vehicle and the schematic diagram of the AGV vehicle, the reference numerals shown in the figure are wide-angle camera 1, main board 2, slave board 3, and Mecanum wheel 4.

[0211] The parameters of the wide-angle camera 1 are as follows: the resolution is 1280×720, the focal length is 3.51mm, it has a wide-angle field of view of 120 degrees, the horizontal viewing angle is 85°, the vertical viewing angle is 69°, and there is no distortion. This camera uses a CMOS sensor, the interface is USB, the operating temperature range is -10°C to 60°C, and the operating humidity range is 15% to 85%.

[0212] Furthermore, the four Mecanum wheels are arranged on the chassis of the vehicle body in an O-type installation method. The Mecanum wheel includes a hub and a plurality of rollers fixedly arranged on the outer periphery of the hub. The angle between the axis of the hub and the axis of the rollers is 45°. The Mecanum wheel has three degrees of freedom, namely rotating around the axis of the hub, moving in a direction perpendicular to the axis of the roller in contact with the ground, and rotating around the contact point between the hub and the ground.

[0213] According to the shape image of the contact part of the rollers of the four wheels of the Mecanum wheel chassis projected onto the ground, it is divided into an O-type installation method and an X-type installation method. The omnidirectional AGV vehicle commonly uses the O-type installation method.

Claims

1. An omnidirectional mobile AGV deviation correction control method based on deep learning, characterized in that, It includes the following steps: S10. The AGV cart arrives at the picking area. The body of the AGV cart is equipped with a camera, and the goods tray located in the picking area has a QR code; S20. The camera captures the goods tray to obtain a detection image containing the QR code, and then sends the detection image to the host computer; S30. The host computer uses the YOLOv8-Pose key point detection algorithm based on deep learning to identify the position and key points of the QR code in the detection image, obtaining the QR code position and key point information. The host computer also sets a standard QR code recognition frame at the center of the detection image; S40. Using the QR code position and key point information, calculate the pose deviation information between the QR code and the standard QR code recognition frame. The pose deviation information includes the position deviation angle, the position deviation amount, and the attitude deviation angle, and then transmit the pose deviation information to the PID controller; S50. The PID controller stores the preset target expectation information. The target expectation information includes that the error distance between the QR code and the standard QR code recognition frame is less than the threshold pixel, and the error angle between the QR code and the standard QR code recognition frame is less than the threshold angle; The PID controller compares the pose deviation information with the target expectation information. When the pose deviation information does not meet the target expectation information, go to S60; when the pose deviation information meets the target expectation information, go to S70; S60. The PID controller corrects the deviation of the AGV cart. According to the preset deviation correction control amount, combined with the kinematic equation of the Mecanum wheel, adjust the motion parameters of the AGV cart. The motion parameters include the moving direction, the moving speed, the rotating direction, and the rotating speed, so that the AGV cart carrying the camera moves in the direction of the center of the QR code, and go to S20; S70. Determine that the QR code is aligned with the standard QR code recognition frame, and end the deviation correction.

2. The omnidirectional mobile AGV deviation correction control method based on deep learning according to claim 1, wherein, In S30, the YOLOv8n-Pose key point detection algorithm based on deep learning specifically includes the following steps: S31. Obtain the QR code data set and data augmentation: Select the publicly available QR code data set provided by the Baidu PaddlePaddle AI Studio platform, and at the same time preprocess the images in the data set, including rotation, cropping, adding noise, adjusting brightness and contrast, to increase the number and complexity of the data set images; S32. Image annotation and data set division: Use the LabelMe software to annotate the data set after data augmentation, annotate the positions of the QR code bounding box and key points, and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1; S33. Build the YOLOv8-Pose-ES network model: On the basis of the YOLOv8-Pose network model, add an ECA attention mechanism module to the backbone network, and use the VoV-GSCSP module and the GSConv module to replace the C2f module and the Conv module respectively in the neck network to form a Slim-neck architecture; S34. Training the model and tuning the parameters to optimize the model: First, use the Mosaic data augmentation method on the dataset in S32, then select four images from the training set for splicing, and then use the dataset processed by the Mosaic data augmentation method to train the YOLOv8-Pose-ES network model. Calculate the loss function and update the model parameters in each iteration to minimize the value of the loss function until the model converges, and finally obtain the PT format model; S35. Model inference acceleration: Convert the PT format model to the ONNX format model, and use the ONNX format model to detect the images in the test set to verify the performance of the ONNX format model.

3. A method for correcting deviation control of an omnidirectional mobile AGV based on deep learning according to claim 2, characterized in that In S33, the backbone grid uses the CSP-Darknet53 structure, and the ECA attention mechanism module is embedded in the layer before the SPPF module of the backbone network; The steps to obtain the ECA attention mechanism module are as follows: S33-1. Convert the two-dimensional feature map with input size H×W in each channel into a scalar through global average pooling (GAP) to simplify the spatial information and retain the global features. The specific calculation formula is as follows: Among them, y is the global feature vector, and x i is the i-th feature map with an input size of H×W; S33-2. The one-dimensional convolutional layer Conv1d performs a linear transformation on the global feature vector y output by the global average pooling (GAP) to learn the interaction weights between different channels and capture the dependencies between channels. The specific calculation formula is as follows: w = σ(C1D k (y)) Formula 2; where σ is the Sigmoid activation function, w is the channel weight, and y is the global feature vector generated by GAP; S33-3. Adaptively determine the value of k through the mapping of the channel dimension C to adjust the interaction range between channels and optimize the feature extraction process. The specific calculation formula is as follows: where γ and b are adjustment parameters with values of 2 and 1, k is the convolutional kernel size, and |t|odd represents the odd number closest to t.

4. The omnidirectional mobile AGV deviation correction control method based on deep learning according to claim 3, wherein, In S33, the neck network is a structure that combines the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN). The GSbottleneck module is designed based on the GSConv module and the Conv module. The VoV-GSCSP module is designed using the one-shot aggregation method based on the GSBottleneck module. The VoV-GSCSP and GSConv are used to replace the C2f module and the Conv module in the neck network respectively to form the Slim-neck architecture.

5. The omnidirectional mobile AGV deviation correction control method based on deep learning according to claim 4, characterized in that After the neck network extracts three feature maps, three detection modules with different scales are connected to the three feature maps respectively. Each detection module serves as a sliding window detector for QR code classification, bounding box regression, and keypoint regression on each layer of the feature pyramid structure.

6. A method for correcting deviation control of an omnidirectional mobile AGV based on deep learning according to claim 2, characterized in that, In S34, the loss function includes the class loss function, the bounding box loss function, and the keypoint loss function; The class loss function uses the BCE loss function. The class loss function is used to determine whether the predicted target box encloses the correct class. The calculation formula is as follows: where N is the total number of samples, p i is the predicted probability of the model for the i-th sample, with a value range of 0 - 1; y i is the true label of the i-th sample, taking the value 1 for the positive class and 0 for the negative class; The border loss function adopts the CIOU loss function and the DFL loss function. The border loss function is used to determine whether the predicted target box can tightly frame the target object. The calculation formula is as follows: Where a is: Where v is where w and h are the width and height of the ground truth box, w gt and h gt are the width and height of the predicted box, c represents the diagonal length of the minimum bounding rectangle of the candidate box and the ground truth box, a is the weight function, v is used to measure the consistency of the aspect ratio, and ρ(b, b gt ) is the square of the Euclidean distance between the center point b of the predicted box and the center point b gt of the ground truth box; DFL(S i ,S i+1 ) = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 )) Equation (8); where S is the probability output value; it can be seen from Equation VIII that when y is very close to y i+1 and the probability output S is large, DFL is small, causing the distribution to approach the annotation center; The key point loss function uses the OKS loss function and the BCE loss function. The key point loss function is used to determine whether the predicted key point exists and whether it is close enough to the target key point. The calculation formula is as follows: where p i represents the i-th key point of the p-th target, d pi represents the Euler distance between the key point prediction value and the true value, s p represents the scale factor of the target, v pi represents the visibility of the key point, σ i represents the standard deviation between the labeled value and the true value of the i-th key point, δ is the visible point calculation function, and OKS is used to measure the similarity between the predicted key point and the true key point, with a value range of 0-1.

7. A full-directional mobile AGV deviation correction control method based on deep learning according to claim 1, characterized in that In the step S40, the posture deviation information between the QR code and the standard QR code recognition frame is calculated, specifically: S41. Use a wide-angle camera to collect a detection image of the bottom surface of the goods tray. The size of the detection image is set to W×H. Taking the center coordinates (W / 2, H / 2) of the detection image as a reference point, determine the center point (x img , y img ) of the standard QR code recognition frame at the center of the detection image. Determine the position of the standard QR code recognition frame at the center of the detection image according to the representation of the actual length of the QR code on the goods tray in pixels; S42. Use the YOLOv8-Pose key point detection algorithm to identify the detected image, and locate the QR code bounding box and four key points. Starting from the lower left corner of the QR code at the center of the standard goods pallet in a clockwise direction, mark the corner points respectively. The lower left corner is the first corner point, the upper left corner is the second corner point, the upper right corner is the third corner point, and the lower right corner is the fourth corner point. Calculate the center point (x qr , y qr ) of the QR code at the center of the goods pallet according to the average value of the coordinates of the four corner points of the QR code at the center of the goods pallet. The specific calculation formula is as follows: Among them, (x qr , y qr ) are the central point coordinates of the QR code at the center of the goods pallet, and (x1, y1), (x2, y2), (x3, y3) and (x4, y4) are the coordinates of the lower left corner, upper left corner, upper right corner and lower right corner of the QR code at the center of the goods pallet; S43. A Cartesian coordinate system is constructed with the center point of the standard QR code recognition frame as the origin, the positive direction of the X-axis is defined as 0°, and the increasing direction of the angle is clockwise; using the center point (x img , y img ) of the standard QR code recognition frame at the center of the detection image and the center point (x qr , y qr ) of the QR code at the center of the goods tray, calculate the angle between the line connecting these two points and the straight line horizontally to the right. This angle is the position deviation angle, and the position deviation angle represents the orientation of the center of the goods tray relative to the AGV cart. Among them, the front corresponds to 90°, the back corresponds to 270°, the right corresponds to 0°, and the left corresponds to 180°; at the same time, calculate the straight-line distance between the center point (x0, y0) of the standard QR code recognition frame at the center of the detection image and the center point (x qr , y qr ) of the QR code at the center of the goods tray. This distance is the position deviation amount, and the positive and negative of the position deviation amount are determined by the value range of the position deviation angle. When the position deviation angle is between 0° and 180°, it means that the center of the goods tray is in front of the AGV cart, and at this time the position deviation amount is negative; when the position deviation angle is between 180° and 360°, it means that the center of the goods tray is behind the AGV cart, and at this time the position deviation amount is positive; the specific calculation formulas for the position deviation angle θ and the position deviation amount d are as follows: θ rad = arctan2(y qr - y img , x qr - x img ) Formula XI; Among them, (x qr , y qr ) is the center point coordinate of the QR code on the center of the goods pallet, (x img , y img ) is the center point coordinate of the standard QR code recognition frame at the center of the detected image, θ rad is the radian between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (-π, π), and θ deg is the angle between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (0°, 360°); S44: Take the center point of the QR code at the center of the cargo pallet (x qr ,y qr ) as the origin to construct a Cartesian coordinate system. The AGV trolley moves in a 90-degree forward direction. The angle taken from the positive direction of the Y axis in a clockwise direction to the negative direction of the Y axis is a positive number, while the angle taken from the positive direction of the Y axis in a counterclockwise direction to the negative direction of the Y axis is a negative number. Use the upper left corner coordinates (x2, y2) and the upper right corner coordinates (x3, y3) of the center QR code of the cargo pallet to calculate the midpoint (x m ,y m ), and calculate (x qr ,y qr ) and (x m ,y m ) The angle between the line connecting the two points and the vertical straight line is the attitude deviation angle; attitude deviation angle The specific calculation formula is as follows: Among them, (x qr , y qr ) is the center point coordinate of the QR code at the center of the goods pallet, (x m , y m ) is the midpoint coordinate of the left and right corner points of the QR code at the center of the goods pallet, (x2, y2) and (x3, y3) are the upper left and upper right coordinates of the QR code at the center of the goods pallet respectively, is the radian between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (-π, π), is the angle between the positive X-axis direction of the straight line between (x qr , y qr ) and (x qr , y qr ), with a value range of (-π, π).

8. A method for correcting deviation control of an omnidirectional mobile AGV based on deep learning according to claim 1, characterized in that, In S50, the PID controller compares the posture deviation information with the preset target expected information, specifically: The posture deviation information is input into the PID controller, and the control output signal is finally generated through the adjustment of the proportional term P, the integral term I and the differential term D of the PID controller; The PID controller output is calculated as follows: Among them, K p , K i , K d are the gain coefficients of proportional, integral and differential respectively, T i is the integral time constant, T d is the differential time constant, e(t) is the deviation signal, and u(t) is the output of the controller.

9. A full-directional mobile AGV deviation correction control method based on deep learning according to claim 1, characterized in that In the S60, the kinematic equation of the Mecanum wheel is specifically: Regarding the body of the AGV vehicle and the center of mass of the four Mecanum wheels as a rigid body, and assuming that there is sufficient friction between the rollers of the Mecanum wheels and the ground, analyze the right front wheel in the top view of the AGV vehicle, named Mecanum wheel B. Establish a robot coordinate system with the center position of the AGV vehicle as point O. It is stipulated that the forward direction is the positive direction of the X-axis, that is, 90° in the Cartesian coordinate system, the leftward direction is the positive direction of the Y-axis, that is, 180° in the Cartesian coordinate system, the upward direction is the positive direction of the robot Z-axis, and the counterclockwise direction around the Z-axis is the positive rotation direction; decompose the overall speed of the vehicle into the speeds of the center of mass of Mecanum wheel B moving along the X and Y directions. The V B_X and V B_y of the center of mass speed of Mecanum wheel B have the following calculation formulas: Among them, V B_X is the left - right moving speed of the centroid of Mecanum wheel B, with right - moving being positive; V B_y is the front - back moving speed of the centroid of Mecanum wheel B, with forward - moving being positive; V x is the front - back moving speed of the AGV cart, with forward - moving being positive; V y is the left - right moving speed of the AGV cart, with forward - moving being positive; ω z is the rotational speed of the AGV cart around point O, with counter - clockwise being positive; r x is half of the wheelbase between the two side wheels, r y is half of the wheelbase between the front and rear wheels; S62, the V of the centroid velocity of the Mecanum wheel B B_X and V B_y and again from V B轮 and V B辊 merged and generated, and its calculation formula is as follows: V B-x = V B轮 + V B辊 × sin β Formula XXI; V B-y = -V B辊 × cosβ Formula XXII; Among them, V B轮 is the linear velocity of Mecanum wheel B, with forward being positive; V B辊 is the linear velocity at which the roller in contact with the ground of Mecanum wheel B generates relative sliding; S63, combining formula 20 to formula 23 of S61 and S62, we can get: Among them, β is the angle between the wheel axis and the roller axis; S64, the axial direction of each roller of the Mecanum wheel is 45 degrees to the wheel axis, Substituting β = 45° and into Formula 24 and Formula 25, we can obtain: V B轮 = V x + V y + ω z × L × (sin(α B ) + cos(α B )) Formula 25; S65. Substitute L×sin(α B ) = r y and L×cos(α B ) = r x into Equation (26), we get: V B轮 = V x + V y + ω z × (r x + r y ) Equation 26; S66: Repeat S61 to S65 to derive V A轮 , V C轮 , V D轮 wheel and V x , V y , ω z The relationship between; The linear velocity direction of the rollers in contact with the ground of Mecanum wheels A, C, and D is related to the left - right movement speed of the centroid of Mecanum wheel B, and V A-y = V A辊 × cosβ, V C-y = V C辊 × cosβ, V D-y = - V D辊 × cosβ. According to the inverse kinematic formula of Mecanum wheels, the target linear velocity or target angular velocity of the four wheels can be obtained from the moving speed and rotational speed of the omnidirectional mobile AGV vehicle as follows: Let V A轮 = R×ω A 、V B轮 = R×ω B 、V C轮 = R×ω C 、V D轮 = R×ω D Substituting the above into the formula, the inverse kinematic equation of the Mecanum wheel can be obtained: Among them, R is the radius of the Mecanum wheel; ω B is the angular velocity when the hub of Mecanum wheel B rotates, and counterclockwise rotation is positive; S47. The direct solution formula of Mecanum wheel kinematics is obtained by combining the target linear velocity or target angular velocity of the four wheels to calculate the moving speed and rotation speed of the omnidirectional mobile AGV, as follows: Let V A轮 = R×ω A 、V B轮 = R×ω B 、V C轮 = R×ω C 、V D轮 = R×ω D Substituting the above into the equation, the forward kinematic equation of the Mecanum wheel can be obtained:

10. An omnidirectional mobile AGV control system based on deep learning, characterized in that, Comprising an AGV vehicle, the AGV vehicle performs the method according to any one of claims 1 to 10; The AGV trolley includes a body, a camera, an upper board, a main board, a PID controller, a slave board, an MCU, a motor drive module, a motor and a Mecanum wheel. The camera is arranged at the center of the top surface of the body, the main board, the PID controller, the slave board, the MCU, the motor drive module and the motor are all arranged inside the body, and the Mecanum wheel is arranged on the chassis of the body; The camera is used to photograph the cargo pallet to obtain a detection image containing the QR code; The mainboard is used to receive the detection image and then forward it to the host computer; The host computer is used to obtain the key point information of the QR code position from the detection image, set a standard QR code recognition frame at the center of the detection image, and use the key point information of the QR code position to calculate the posture deviation information between the QR code and the standard QR code recognition frame; The PID controller is used to receive the posture deviation information, and then compare the posture deviation information with the preset target expected information, so that the computer obtains the correction control information; The slave board is used to receive the correction control information and then forward it to the MCU; The MCU is used to generate a PWM signal according to the deviation correction control information; The motor driving module is used to adjust the speed of the motor according to the duty cycle of the PWM signal, and to control the direction of the motor according to the phase of the PWM signal; The motor is used to drive the Mecanum wheel to make the AG car go straight along a specified path to a target position or turn to a specified angle.

11. A omnidirectional mobile AGV control system based on deep learning according to claim 10, characterized in that The four Mecanum wheels are arranged on the chassis of the vehicle body in an O-shaped installation manner. The Mecanum wheel includes a hub and a plurality of rollers fixedly arranged on the outer periphery of the hub. The included angle between the axis of the hub and the axis of the rollers is 45°. The Mecanum wheel has three degrees of freedom, namely, rotating around the axis of the hub, moving in a direction perpendicular to the axis of the rollers in contact with the ground, and rotating around the contact point between the hub and the ground.

Citation Information

Cited By

  • Visual inspection and positioning system of cable hoist loading and unloading vehicle

    CN121074347A