A lightweight automatic unloading method and system for a throwing tool based on computer vision
By applying a lightweight material throwing automatic unloading method based on computer vision in agricultural machinery, using deep learning models to analyze video data and adjust the position of material throwing tools, the problems of high cost and low real-time performance in the existing technology are solved, and efficient and automated agricultural machinery unloading operations are achieved.
Patent Information
- Application Number
- CN202510273333.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing automatic unloading technology is difficult to effectively apply in the low-cost computing environment of small and medium-sized enterprises due to high cost, low real-time and high dependence on labor.
The lightweight material throwing tool automatic unloading method based on computer vision is adopted to collect video data through imaging equipment, analyze video data using pre-trained deep learning models, calculate the adjustment pulse information of the material throwing tool, and adjust the position of the material throwing tool in real time through the unloading controller to realize automatic unloading.
It significantly reduces the system's calculation cost, improves feature extraction capabilities and adaptability, realizes automatic unloading that operates efficiently on low-cost computing equipment, and improves the automation and accuracy of the operation process.
Smart Images

Figure CN119769284B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural machinery automation, and particularly to a lightweight thrower automatic unloading method and system based on computer vision. Background Art
[0002] In recent years, significant progress has been made in the unmanned technology of agricultural equipment, and the automatic unloading of harvesters is also developing towards the unmanned direction. At present, the harvesting and unloading link mainly relies on the traditional manual control unloading mode, but there has also been research starting to explore automatic unloading technology. Currently, the research on automatic unloading mainly focuses on the fixed-point automatic unloading mode and the side-harvesting and side-automatic unloading mode.
[0003] The fixed-point unloading mode mainly relies on the harvester to set a fixed unloading point in advance. After reaching the given unloading point, the harvester automatically unloads, and then the harvester returns to harvest. Different unloading points need to be set for different working locations, which limits the operation flexibility and affects the overall efficiency.
[0004] The side-harvesting and side-automatic unloading mode is that the transport vehicle always maintains a relative position with the operating harvester. After the transport vehicle's grain bin is full, it returns to the storage point, and a new transport vehicle continues to maintain the relative position with the harvester and continues the side-harvesting and side-unloading state; the advantage of this mode is that the harvester can keep harvesting without stopping, improving the operation efficiency. However, most of the existing research relies on unmanned systems and various detection devices and cannot be widely promoted.
[0005] Combining the traditional manual control unloading mode and the side-harvesting and side-automatic unloading mode may be a new solution. With the development of deep learning and the help of image recognition technology, fixing the image acquisition device on the thrower of the harvester to perform real-time recognition on the transport vehicle and adjusting the pose of the thrower according to the recognition result has become a feasible solution.
[0006] However, the existing image recognition methods based on deep learning construct complex models to ensure the algorithm accuracy, and there are problems such as high training costs and high model computing power requirements. Small and medium-sized enterprises usually lack high-end GPUs or dedicated AI acceleration hardware and mostly use embedded devices for inference training. Under the limited cost and computing power, the existing methods are difficult to be effectively applied. Therefore, designing an image recognition technology that can still operate efficiently in an environment with limited cost and computing power, without using an unmanned system to control the harvester and the transport vehicle and being able to ensure side-harvesting and side-automatic unloading, has important practical significance and application value. Summary of the Invention
[0007] In view of the above deficiencies in the prior art, the present invention provides a lightweight thrower automatic unloading method and system based on computer vision, which solves the problems of high cost, low real-time performance, and high dependence on manual labor in the prior art.
[0008] To achieve the above-mentioned invention purpose, the technical solution adopted by the present invention is as follows: A lightweight automatic unloading method for a throwing tool based on computer vision, comprising the following steps:
[0009] S1. Collect the video of the agricultural harvesting site through an imaging device to obtain video data;
[0010] S2. Receive the video data through an image processing module and analyze the video data using a pre-trained deep learning model to obtain the position data of the transport vehicle; calculate the adjustment pulse information of the throwing tool according to the position data of the transport vehicle;
[0011] S3. According to the adjustment pulse information of the throwing tool, the attitude of the throwing tool is adjusted in real time through a unloading controller to achieve automatic unloading.
[0012] The present invention also provides a system for implementing the above-mentioned lightweight automatic unloading method for a throwing tool based on computer vision, comprising:
[0013] An imaging device for collecting the video of the agricultural harvesting site to obtain video data;
[0014] An image processing module for receiving the video data and analyzing the video data using a pre-trained deep learning model to obtain the position data of the transport vehicle; calculating the adjustment pulse information of the throwing tool according to the position data of the transport vehicle;
[0015] A unloading controller for adjusting the attitude of the throwing tool in real time according to the adjustment pulse information of the throwing tool to achieve automatic unloading;
[0016] A throwing tool for unloading.
[0017] The beneficial effects of the present invention are as follows:
[0018] 1. By reconstructing the YOLOV8 model, the present invention proposes a deep learning model, which not only significantly enhances the feature extraction ability and adaptability of the system, but also reduces the computational cost of the system. The alternately stacked ShuffleBlock and AugShuffleBlock structures in the backbone part make full use of the information in the middle layer of the feature map, avoid information waste, and promote cross-layer information mixing to improve the model performance. At the same time, the length of the optimization path in the module part is reduced, which can fully alleviate the risk of gradient disappearance and explosion inside the module, and can identify the transport vehicle in the agricultural harvesting scene faster and more accurately.
[0019] 2. All AugShuffleBlocks in the backbone part of the deep learning model proposed in the present invention adopt the technology of XBN (Batch Normalization Free) to replace the traditional Batch Normalization (BN), effectively preventing the accumulation of estimation drift in the network and reducing the performance degradation caused by distribution drift during the testing process. This enables the model to maintain stable performance in different application scenarios, especially important in a highly variable environment such as the agricultural harvesting scenario.
[0020] 3. The neck of the deep learning model proposed in the present invention adopts an Omni-Dimensional Dynamic Convolution layer (ODConv). By assigning different attention values to the position of the convolution kernel, input channels, output channels, and the overall convolution kernel, it realizes multi-dimensional differential convolution operations, thus greatly enhancing the feature capture ability of convolution, not only enriching the context information but also improving the model's processing ability for complex backgrounds and target diversity.
[0021] 4. According to the video data provided by the imaging device, the present invention identifies the position of the transport vehicle by the image processing module and calculates the target position of the required throwing tool in real time. The unloading controller precisely adjusts the pose of the throwing tool according to the adjustment pulse, realizing accurate and real-time throwing, and improving the automation and accuracy of the entire operation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flowchart of the method proposed in the present invention;
[0023] Figure 2 is a schematic structural diagram of the deep learning model obtained after reconstructing YOLOV8;
[0024] Figure 3 is a schematic structural diagram of the ShuffleBlock (SFB);
[0025] Figure 4 is a schematic structural diagram of the AugShuffleBlock (AugSFB);
[0026] Figure 5 is a flowchart of the operation of the system proposed in the present invention;
[0027] Figure 6 is a structural diagram of the system proposed in the embodiment.
[0028] Among them, 1. Transport vehicle; 2. Imaging device; 3. Throwing tool; 4. Image processing module; 5. Display; 6. Unloading controller; 7. y Axis control motor; 8. x Axis control motor. DETAILED DESCRIPTION OF THE INVENTION
[0029] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0030] In one embodiment of the present invention, as Figure 1 shown, a lightweight automatic unloading method for a throwing tool based on computer vision provided by the present invention includes the following steps:
[0031] S1. Collect the video of the agricultural harvesting site through the imaging device 2 to obtain video data;
[0032] S2. Receive the video data through the image processing module 4 and analyze the video data using a pre-trained deep learning model to obtain the position data of the transport vehicle 1; calculate the adjustment pulse information of the throwing tool 3 according to the position data of the transport vehicle 1;
[0033] S3. According to the adjustment pulse information of the throwing tool 3, the attitude of the throwing tool 3 is adjusted in real time through the unloading controller 6 to achieve automatic unloading.
[0034] Among them, the deep learning model is a model obtained by reconstructing YOLOV8; in order to reduce the parameters of the model itself and ensure a certain feature extraction ability, Darknet53 in the backbone part of the YOLOV8 model is replaced by ShuffleBlock (shuffle block) and AugShuffleBlock (enhanced shuffle block) stacked alternately three times; and BN in AugShuffleBlock is replaced by XBN (batch-free normalization) to avoid the accuracy loss caused by the cumulative error during the training of BN; the convolutional layer before the neck feature fusion of the YOLOV8 model is replaced by a 3×3 ODConv (omnidimensional dynamic convolutional layer) to enhance the detail extraction ability before feature fusion and capture finer spatial features; at the same time, C2f in the neck of the YOLOV8 model is replaced by DWConv (depthwise separable convolutional layer), effectively reducing the amount of calculation and the number of parameters, while retaining the feature extraction ability; the last two Concat in the neck of the YOLOV8 model are replaced by ADD (feature-by-feature enhancement layer), which can increase the information volume of each dimension without increasing the dimension of the image.
[0035] As Figure 2As shown in the figure, the backbone of the deep learning model obtained by reconstructing YOLOV8 includes: a convolutional layer, a first shuffle block, a first enhanced shuffle block, a second shuffle block, a second enhanced shuffle block, a third shuffle block, and a third enhanced shuffle block connected in sequence; the input of the convolutional layer is the input of the entire deep learning model.
[0036] The neck of the deep learning model includes: a first full-dimensional dynamic convolutional layer, a first upsampling layer, a first splicing layer, a first depthwise separable convolutional layer, a second full-dimensional dynamic convolutional layer, a second upsampling layer, a second splicing layer, a second depthwise separable convolutional layer, a third depthwise separable convolutional layer, a first feature addition enhancement layer, a fourth depthwise separable convolutional layer, a fifth depthwise separable convolutional layer, a second feature addition enhancement layer, and a sixth depthwise separable convolutional layer connected in sequence;
[0037] Among them, the input end of the first full-dimensional dynamic convolutional layer is connected to the output end of the third enhanced shuffle block; the other input end of the first splicing layer is connected to the output end of the second enhanced shuffle block; the other input end of the second splicing layer is connected to the output end of the first enhanced shuffle block; the other input end of the first feature addition enhancement layer is connected to the output end of the second full-dimensional dynamic convolutional layer; the other input end of the second feature addition enhancement layer is connected to the output end of the first full-dimensional dynamic convolutional layer; the outputs of the second depthwise separable convolutional layer, the fourth depthwise separable convolutional layer, and the sixth depthwise separable convolutional layer are respectively used as the inputs of the head of the deep learning model.
[0038] Through the above reconstruction of the YOLOV8 model, the model can be made more lightweight while ensuring a certain accuracy, suitable for deployment on devices with low-cost computing power, and reducing the cost pressure on enterprises.
[0039] As Figure 3 , Figure 4 shown, the specific steps for extracting image features from the data to be analyzed using the backbone of the deep learning model are as follows:
[0040] B11. Through the Shuffleblock (shuffle block), the input feature map , is divided into two parts along the channel dimension, and the number of channels of each part is ;
[0041] B12. Perform 1×1 pointwise convolution on to obtain , and its expression is ; perform 3×3 depthwise separable convolution on with a stride of 2 to obtain , and its expression is ; perform Use 1×1 pointwise convolution to adjust the number of channels and obtain , and its expression is ;
[0042] B13. Perform 3×3 depthwise separable convolution on to obtain , and its expression is ; Perform 1×1 pointwise convolution on to adjust the number of channels and obtain , and its expression is ;
[0043] B14. Concatenate and together to obtain , and its expression is ; Perform channel shuffle on to interleave the channels from different parts and obtain a new feature map , and its expression is ;
[0044] B21. Perform further feature extraction on , through the AugShuffleBlock (enhanced shuffle block); Divide the feature map into two parts according to the set splitting ratio r (less than 0.5); where r is an adjustable parameter used to control the computational cost and redundancy of the module;
[0045] B22. Perform 1×1 pointwise convolution on to obtain , and its expression is , The number of channels of remains ; Perform 3×3 depth convolution on with a stride of 2 to obtain ;
[0046] B23. Perform channel cross on to obtain , and its expression is ;
[0047] B24. Perform channel cross on to obtain , and its expression is ;
[0048] B25. Concatenate with to obtain , and its expression is ;
[0049] B26. Perform a 1×1 point-by-point convolution on to adjust the number of channels and obtain , whose expression is ;
[0050] B27. Concatenate with to obtain ; its expression is ;
[0051] B28. Concatenate with to obtain , whose expression is ; Perform channel shuffle on to obtain , whose expression is .
[0052] In the neck structure of the deep learning model, every time the feature map passes through ODConv, the following operations are performed on the input feature map:
[0053]
[0054] Among them, represents the input of ODConv, represents the output of ODConv, is the convolution kernel 's attention scalar, is the newly introduced attention mechanism, which is calculated along the spatial dimension and input channel dimension of the convolution kernel respectively; represents the attention mechanism along the output channel dimension. The deep learning model enhances the representation ability of the backbone network through the feature fusion method that combines the bottom-up and top-down paths, avoiding the rigid matching of the target size and network depth.
[0055] In this embodiment, the specific steps for training the deep learning model include:
[0056] A1. Use an imaging device to collect agricultural harvesting site videos, and determine whether the data collected by the imaging device meets the requirements (whether an object is photographed, whether the positive and negative samples are balanced, whether there is serious blurring or occlusion, etc.). Use the data that meets the requirements as the original data;
[0057] A2. Set the size of each frame of image in the original data to 320×320 to obtain data in a unified format;
[0058] A3. After manually annotating the data obtained in the unified format, use data augmentation techniques such as random rotation, saturation change, brightness change, horizontal flipping, Gaussian noise addition, and salt-and-pepper noise addition to increase the diversity of the images and expand the number of image samples to obtain a training set;
[0059] A4. Input the training set into the deep learning model to obtain the position data of the target, that is, the position data of the transport vehicle 1;
[0060] A5. Calculate the final detection loss; its calculation expression is:
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068] where is the final detection loss, is the CIoU loss of the bounding box, is the cross-entropy loss of the object category, is the binary cross-entropy loss of the object score; 、 、 are the weight factors assigned to each loss term; the center coordinates and size of the ground truth box are , represents the center abscissa of the ground truth box, represents the center ordinate of the ground truth box, represents the width of the ground truth box, represents the height of the ground truth box; the center coordinates and size of the predicted box are , represents the center abscissa of the predicted box, represents the center ordinate of the predicted box, represents the width of the predicted box, represents the height of the predicted box; represents the sum of the widths of the ground truth boxes; represents the sum of the heights of the predicted boxes; is the balance parameter; is the parameter used to measure the consistency of the aspect ratio;c is the number of categories, is the indicator variable for category c in the true label, is the predicted probability for category c ; is the true label, is the predicted score; is used to measure the overlap degree between the predicted bounding box and the true bounding box, is the bounding box regression loss function; is the area of the overlapping part between the true bounding box and the predicted bounding box; is pi; tan is the tangent trigonometric function;
[0069] A6. Perform backpropagation using the calculated final detection loss, and update the parameters of the deep learning model using the Adam optimizer to reduce the value of the loss function;
[0070] A7. Determine whether the loss function no longer decreases or the accuracy of the deep learning model does not improve significantly. If so, end the training; otherwise, return to step A4.
[0071] To further verify the advantages of the deep learning model described in the present invention compared with other existing mainstream models, a comparative experiment was conducted, including YOLOv3-tiny, YOLOv5-n, YOLOv6-n, YOLOv7-tiny, YOLOv8-n, and SSD300. The experimental results are shown in Table 1:
[0072] Table 1
[0073]
[0074] The experimental results show that the number of parameters and the amount of computation required by the deep learning model proposed in the present invention are less than those of other mainstream models, greatly reducing the demand for computing resources and enabling the solution to proceed smoothly under the conditions of low computing power and low cost.
[0075] The present invention also provides a system for implementing the above-mentioned lightweight automatic unloading method of the throwing tool 3 based on computer vision, including an imaging device 2, which is mostly a high-definition camera and is used to collect video data of the agricultural harvesting site; an image processing module 4, which is used to receive the video data and analyze the video data using a pre-trained deep learning model to obtain the position data of the transport vehicle 1; calculate the adjustment pulse information of the throwing tool 3 according to the position data of the transport vehicle 1; a unloading controller 6, which is used to adjust the pose of the throwing tool 3 in real time according to the adjustment pulse information of the throwing tool 3 to achieve automatic unloading; and a throwing tool 3, which is used for unloading.
[0076] In this embodiment, as Figure 6As shown, the imaging device 2 is installed 1 m away from the discharge port of the throwing tool 3, ensuring that its field of view can cover all areas that the transport vehicle 1 may reach. The included angle between the imaging device 2 and the throwing tool 3 is set between 30° and 45°, and can be adjusted according to the specifications of different throwing tools 3 in other embodiments. The imaging device 2 is calibrated before installation to ensure the consistency and accuracy of the image quality. In this embodiment, a display 5 is also included for displaying real-time videos and warning messages. After the imaging device 2 collects the agricultural field operation video, the video is transmitted to the image processing module 4. The image processing module 4 mainly consists of devices such as an MPU chip, a control board, a high-speed storage unit, and a RAM, and is mainly responsible for identifying the position of the transport vehicle 1 in the video data, inferring the displacement of the throwing tool 3, and transmitting the processed video to the display 5. After the lightweight deep learning model in the image processing module 4 identifies the position of the transport vehicle 1, it immediately calculates the target position of the throwing tool 3 and then uses the PID algorithm to calculate x the adjustment pulse information of the Z-axis motor 8 and y the adjustment pulse information of the Y-axis motor 7. When the pulse information meets the mechanical requirements, a discharge signal is sent to the discharge controller 6; after receiving the discharge request from the image processing module 4, the discharge controller 6 adjusts the pose of the throwing tool 3 according to x the adjustment information of the Z-axis motor 8 and y the adjustment information of the Y-axis motor 7, and performs the discharge operation after the adjustment is completed. If the transport vehicle 1 is not recognized for a period of time, a stop discharge signal is sent to the discharge controller 6, and a warning message is displayed on the display 5, and the system stops working. When the system is in the running state, it continuously receives new image inputs from the imaging device 2 and responds quickly.
[0077] The specific working process of the system proposed in the present invention is as Figure 5 shown, including:
[0078] C1. Start and initialize the system settings;
[0079] C2. After the initialization is completed, check whether the imaging device 2 is available. If it is available, use the imaging device 2 to collect the original data and enter step C4; otherwise, enter step C3;
[0080] C3. Output the reason for being unavailable and shut down the system;
[0081] C4. Load the trained deep learning model, check whether the deep learning model is loaded successfully. If it is successful, enter step C5; otherwise, enter step C3;
[0082] C5. Check whether the original data collected by the imaging device is available. If it is available, enter step C6; otherwise, return to step C1;
[0083] C6. Take the raw data collected by the imaging device as the input of the image processing module 4, and use the deep learning model to perform target recognition on the raw data to obtain the recognition result;
[0084] C7. Check the recognition result: If the transport vehicle 1 is not recognized, go to step C8; If the transport vehicle 1 is recognized, identify the center point position of the transport vehicle 1; To facilitate the uniform unloading of the throwing tool 3 onto the transport vehicle 1, divide the transport vehicle 1 into M areas, and set the time limit for each area to be , record the current time and enter step C9;
[0085] C8. Determine whether the transport vehicle 1 has not been recognized after n times of recognition. If so, send a stop unloading request, output a warning message and shut down the system; Otherwise, return to step C6;
[0086] C9. Calculate the target position of the throwing tool 3 according to the current position of the transport vehicle 1 , and determine whether the area where the target center point of the throwing tool 3 is located meets the time limit. If so, enter step C10; Otherwise, update the unloading area of the transport vehicle, and calculate the target position of the throwing tool according to the current position of the transport vehicle;
[0087] The time limit is:
[0088] The constraint condition is:
[0089]
[0090] ,
[0091] Among them, N means dividing the total time T of the entire unloading process evenly into N time periods, and the time length of each time period is ; M represents the number of divided areas of the transport vehicle, ; represents in the i th time period, the actual residence time of the transport vehicle 1 in the j th area; represents the expected residence time of the transport vehicle 1 in the j th area; represents minimization.
[0092] C10. Use the PID algorithm to calculate the adjustment pulses of the x axis motor and the y axis motor required for the throwing tool 3 to adjust from the current position to the target position; Its expression is:
[0093]
[0094]
[0095]
[0096] Among them, represents the current position coordinates of the material throwing tool, represents the target position coordinates of the material throwing tool; represents time, represents the x axis error between the target position and the current position at the moment, represents y the axis error between the target position and the current position at the moment; x is the adjustment information of the axis at the moment, y is the adjustment information of the , are respectively the proportional gains of the x axis and the y axis; , are respectively the integral gains of the x axis and the y axis; , are respectively the derivative gains of the x axis and the y axis; represents time, .
[0097] C11. Determine whether the calculated adjustment pulse meets the mechanical condition limits of the motor (the rotation angle of the material throwing tool 3 is different for different models of transport vehicle 1, that is, x the maximum and minimum pulses of the y axis motor and the
[0098] axis motor are different). If so, enter C12; otherwise, output a warning message and shut down the system;
[0099] In summary, the present invention has simple deployment, low computing power requirements, low learning cost and good detection effect.
Claims
1. A method for automatic unloading of lightweight throwing tools based on computer vision, characterized in that: The following steps are involved: S1. Collect agricultural harvesting scene videos through imaging equipment to obtain video data; S2. Receive video data through an image processing module and analyze the video data using a pre-trained deep learning model to obtain position data of the transport vehicle; calculate adjustment pulse information of the throwing tool according to the position data of the transport vehicle; The deep learning model is an optimized YOLOV8 model; that is, the Darknet53 in the backbone of the YOLOV8 model is replaced with alternately stacked ShuffleBlock and AugShuffleBlock; and the BN in the AugShuffleBlock is replaced with XBN; the convolution layer before the neck feature fusion of the YOLOV8 model is replaced with 3×3 ODConv, C2f is replaced with DWConv, and the last two Concats are replaced with ADD; The backbone of the deep learning model obtained after reconstructing YOLOV8 includes: a convolutional layer, a first shuffle block, a first enhanced shuffle block, a second shuffle block, a second enhanced shuffle block, a third shuffle block, and a third enhanced shuffle block connected in sequence; wherein the input of the convolutional layer is the input of the entire deep learning model; The neck of the deep learning model obtained after reconstructing YOLOV8 includes: a first full-dimensional dynamic convolution layer, a first upsampling layer, a first splicing layer, a first depth-separable convolution layer, a second full-dimensional dynamic convolution layer, a second upsampling layer, a second splicing layer, a second depth-separable convolution layer, a third depth-separable convolution layer, a first feature-added enhancement layer, a fourth depth-separable convolution layer, a fifth depth-separable convolution layer, a second feature-added enhancement layer, and a sixth depth-separable convolution layer, which are sequentially connected; Wherein, the input end of the first full-dimensional dynamic convolutional layer is connected to the output end of the third enhanced shuffle block; Another input end of the first splicing layer is connected to the output end of the second enhanced shuffle block; Another input end of the second splicing layer is connected to the output end of the first enhanced shuffle block; Another input end of the first feature gradual enhancement layer is connected to the output end of the second full-dimensional dynamic convolution layer; Another input end of the second feature gradual enhancement layer is connected to the output end of the first full-dimensional dynamic convolution layer; The outputs of the second depth-wise separable convolutional layer, the fourth depth-wise separable convolutional layer, and the sixth depth-wise separable convolutional layer are respectively used as inputs to the head of the deep learning model; S3. According to the adjustment pulse information of the throwing tool, the posture of the throwing tool is adjusted in real time through the unloading controller to realize automatic unloading.
2. According to the computer vision-based lightweight throwing tool automatic unloading method of claim 1, it is characterized in that: The specific steps for training a deep learning model include: A1. Use imaging equipment to collect agricultural harvesting site videos and obtain raw data; A2. Set the size of each frame of the original data to a×a to obtain data in a unified format; A3. After manually labeling the data in a unified format, the number of image samples is expanded through data enhancement technology to obtain a training set; A4. Input the training set into the deep learning model to obtain the target location data, i.e., the transport vehicle location data; A5. Calculate the final detection loss; its calculation expression is: in is the final detection loss, is the CIoU loss of the bounding box, is the cross entropy loss of the object category, Binary cross entropy loss for object scores; , , is the weight factor assigned to each loss term; the center coordinates and size of the real box are , represents the central horizontal coordinate of the real frame, Represents the central ordinate of the real frame, Indicates the width of the real box, Represents the height of the real box; the center coordinates and size of the predicted box are , Represents the central horizontal coordinate of the prediction box, Represents the central ordinate of the prediction box, Indicates the width of the prediction box, Indicates the height of the prediction box; Represents the sum of the widths of the real boxes; Represents the sum of the heights of the prediction boxes; is the balance parameter; is a parameter used to measure the consistency of aspect ratio; c is the number of categories, is the indicator variable of category c in the true label, is the predicted probability of category c, is the true label, is the predicted score; Used to measure the overlap between the predicted box and the true box. is the bounding box regression loss function; is the area of the overlap between the real box and the predicted box; is the ratio of pi; tan is the tangent trigonometric function; A6. Use the calculated final detection loss for back propagation and use the Adam optimizer to update the parameters of the deep learning model to reduce the value of the loss function. A7. Determine whether the loss function no longer decreases or whether the accuracy of the deep learning model has not been significantly improved. If so, end the training; otherwise, return to step A4.
3. The method for automatically unloading a lightweight throwing tool based on computer vision according to claim 1, characterized in that: The specific steps to extract image features using the backbone of the deep learning model are: B11, through the Shuffleblock feature map Divided along the channel dimension Two parts; B12, Yes Perform 1×1 point-by-point convolution to obtain , whose expression is ;right Perform a 3×3 depth-wise separable convolution with a step size of 2, and we get , whose expression is ;right Using 1×1 point-wise convolution, we get , whose expression is ; B13, Yes Perform a 3×3 depth-separable convolution to obtain , whose expression is ;right Perform 1×1 point-by-point convolution to obtain , whose expression is ; B14. and Put together and get , whose expression is ;right Shuffle the channels to get a new feature map , whose expression is ; B21, through AugShuffleBlock Further feature extraction is performed; the feature map According to the set split ratio r, Two parts, where r is an adjustable parameter; B22, yes Perform 1×1 point-by-point convolution and adjust the number of channels to obtain , whose expression is ;right Perform a 3×3 depth convolution with a step size of 2 to get , whose expression is ; B23, yes Perform channel crossover to obtain , whose expression is ; B24, yes Perform channel crossover to obtain , whose expression is ; B25, will and Splice and get , whose expression is ; B26, Yes Perform 1×1 point-by-point convolution and adjust the number of channels to obtain , whose expression is ; B27, will and Splice and get ; Its expression is ; B28, will and Splice and get , whose expression is ;right Shuffle the channels and get , whose expression is .
4. A system for automatic unloading of lightweight throwing tools based on computer vision according to any one of claims 1 to 3, characterized in that: include: Imaging equipment, used to collect agricultural harvesting site videos and obtain video data; An image processing module is used to receive video data and analyze the video data using a pre-trained deep learning model to obtain the position data of the transport vehicle; and calculate the adjustment pulse information of the throwing tool according to the position data of the transport vehicle; The unloading controller is used to adjust the posture of the throwing tool in real time according to the adjustment pulse information of the throwing tool to realize automatic unloading; Throwing tool, used for unloading.
5. The system according to claim 4, characterized in that The specific working steps of the system are: C1. Start and initialize system settings; C2. After initialization is completed, check whether the imaging device is available. If yes, use the imaging device to collect raw data and proceed to step C4. Otherwise, proceed to step C3. C3, output the reason for unavailability and shut down the system; C4, loading the trained deep learning model, checking whether the deep learning model is loaded, if yes, proceed to step C5, otherwise proceed to step C3; C5, check whether the original data collected by the imaging device is available, if yes, go to step C6, otherwise return to step C1; C6. Use the raw data collected by the imaging device as the input of the image processing module, use the deep learning model to perform target recognition on the raw data, and obtain the recognition result; C7, check the recognition result: if the transport vehicle is not recognized, proceed to step C8; if the transport vehicle is recognized, divide the transport vehicle into M areas, record the current time and proceed to step C9; C8, determine whether n identifications have been performed and no transport vehicle has been identified, if yes, send a stop unloading request, output a warning message and shut down the system; otherwise, return to step C6; C9. Calculate the target position of the throwing device according to the current position of the transport vehicle , and determine whether the area where the target center point of the throwing tool is located meets the time limit, if yes, proceed to step C10; otherwise, update the unloading area of the transport vehicle, and calculate the target position of the throwing tool according to the current position of the transport vehicle; C10, using PID algorithm to calculate the adjustment pulse required for the throwing tool to adjust from the current position to the target position; C11, determine whether the calculated adjustment pulse meets the mechanical condition limit, if yes, enter C12; otherwise, output a warning message and shut down the system; C12, output the unloading request to the unloading controller, and the unloading controller controls the throwing device to automatically adjust the posture and perform the throwing operation; repeat steps C6-C7.
6. The system according to claim 5, characterized in that The expression of the time limit described in step C9 is: The constraints are: , Where N means that the total time T of the entire unloading process is divided into N time periods, and the length of each time period is ; M represents the number of divided areas for the transport vehicle; represents the actual stay time of the transport vehicle in the jth area during the i-th period; represents the expected stay time of the transport vehicle in the jth area; Indicates minimization.
7. The system according to claim 6, characterized in that The PID algorithm is used to calculate the adjustment pulse required for the throwing tool to adjust from the current position to the target position. The expression is: in, Indicates the current position coordinates of the throwing tool. Indicates the target position coordinates of the throwing tool; Indicates time, express The x-axis error between the target position and the current position at the moment, express The y-axis error between the target position and the current position at the moment; for The adjustment information of the x-axis at the moment, for Adjustment information of the y-axis at the moment; , They are the proportional gains of the x-axis and y-axis respectively; , are the integral gains of the x-axis and y-axis respectively; , They are the differential gains of the x-axis and y-axis respectively; Indicates time, .
Citation Information
Patent Citations
Deep learning-based throwing tool automatic positioning and tracking system and method
CN117789098A