Palm Fruit Grabbing and Loading Control Method, Device, Equipment and Medium
Through the improved Yolov8 network model and the technology combining DWConv and Faster-PGLU modules, the problems of low efficiency of loading and transfer and difficulty in identification of palm fruits are solved, and automated loading and pedestrian safety guarantees are achieved.
Patent Information
- Application Number
- CN202411429809.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-10-12
AI Technical Summary
In the prior art, palm fruit loading and transporting mainly relies on labor, which is low in efficiency and high in cost, and it brings huge challenges to palm fruit recognition in complex lighting environments and dense stacking scenarios.
The improved Yolov8 network model is adopted, combined with the DWConv and Faster-PGLU modules, the RGB images and depth images in the palm fruit transport truck are targeted to determine the three-dimensional coordinates of the palm fruit, and automatic loading operations are carried out through the robotic arm to ensure the safety of pedestrians.
It realizes accurate identification and positioning of palm fruits in complex lighting and dense stacking scenarios, improves loading efficiency, reduces management costs, and ensures the safety of the robotic arm.
Smart Images

Figure CN119359995B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of agricultural production, and particularly to a method for controlling the grasping and loading of palm fruits, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art
[0002] Palm oil has become one of the most consumed vegetable oils in the world and is one of the main sources of daily and industrial oils. Indonesia and Malaysia are the largest palm oil exporting countries. Approximately 90% of palm oil is used for food purposes, while approximately 10% is used for industrial products. Palm oil is pressed from palm fruits, which can be processed into vegetable oil and made into various daily necessities such as cosmetics, soaps, detergents, textile oils, and biodiesel. Palm fruits are harvested from oil palm trees, and then the oil is extracted from these fruits. Since palm fruits have a high oil yield and can achieve the maximum economic benefits on limited land, the planting area of oil palm trees is increasing, thus increasing the demand for human resources. The demand for human resources conflicts with the high labor costs, and it is necessary to vigorously promote the mechanical intelligence in the park to solve the dilemma of labor shortage.
[0003] The loading and transfer of palm fruits is an important link, usually carried out manually. The traditional manual fruit loading operation has low efficiency and high management costs. During the palm fruit harvest season, tens of thousands of palm fruits are piled up by the roadside, waiting for the transfer vehicle to come and load them and transport them to the oil mill. The complex and changing outdoor environment during actual fruit grasping operations poses a huge challenge to palm fruit recognition. To improve efficiency, the fruit grasping operation must run for a long time, and at the same time, the in-vehicle computer of the transfer vehicle does not have such high performance.
[0004] In summary, in view of the problems in the prior art that the loading and transfer of palm fruits are usually carried out manually, the traditional manual fruit loading operation has low efficiency and high management costs, and the complex and changing outdoor environment during actual fruit grasping operations poses a huge challenge to palm fruit recognition, etc., the present application makes corresponding explorations to solve these problems. Summary of the Invention
[0005] The purpose of the present application is to solve the above problems and provide a method for controlling the grasping and loading of palm fruits, a corresponding device, an electronic device, and a computer-readable storage medium.
[0006] To meet the various purposes of the present application, the following technical solutions are adopted:
[0007] A method for controlling the grasping and loading of palm fruits proposed for one of the purposes of the present application includes:
[0008] In response to the palm fruit grasping and loading control instruction, obtain the RGB image of the target working area in the palm fruit transporter containing the palm fruits to be loaded and / or pedestrians, as well as the corresponding depth image of the target working area;
[0009] Use the target detection model trained to the convergence state to perform target detection on the RGB image of the target working area to determine the central two-dimensional pixel coordinates corresponding to the palm fruits to be loaded and / or pedestrians. Among them, the basic network architecture of the target detection model is an improved Yolov8 network model. The improved Yolov8 network model replaces all Conv Module modules in the original Yolov8 network model with DWConv modules and replaces all C2F modules in the original Yolov8 network model with Faster-PGLU modules. Among them, the Faster-PGLU module is constructed by combining convolutional gated linear units with the FasterNet structure introducing the PConv operator;
[0010] Align the RGB image of the target working area and the depth image of the target working area, and match the depth value in the depth image of the target working area according to the central two-dimensional pixel coordinates of the palm fruits in the RGB image of the target working area to determine the three-dimensional coordinates of the palm fruits to be loaded in the real space;
[0011] Determine the safe working space of the robotic arm of the palm fruit transporter, detect whether the central two-dimensional pixel coordinates corresponding to the pedestrians are within the safe working space. If the central two-dimensional pixel coordinates corresponding to the pedestrians are outside the safe working space, control the robotic arm to perform the loading operation according to the three-dimensional coordinates of the palm fruits to be loaded in the real space. Among them, the safe working space is larger than the operable space of the robotic arm;
[0012] If the central two-dimensional pixel coordinates corresponding to the pedestrians are within the safe working space, send a robotic arm operation stop instruction to control the robotic arm to pause the loading operation to complete the control of palm fruit grasping and loading.
[0013] Optionally, the depth factor of the improved Yolov8 network model is 0.165, the width factor is 0.125, and the maximum number of channels is 512.
[0014] Optionally, before the step of obtaining the RGB image of the target working area in the palm fruit transporter containing the palm fruits to be loaded, it includes:
[0015] In response to the image preprocessing instruction, perform image scaling, data augmentation, and normalization processing on the RGB image of the target working area containing the palm fruits to be loaded and / or pedestrians;
[0016] Adjust the RGB image of the target working area to a preset image size for image scaling;
[0017] Perform data augmentation on the RGB image of the target working area after image scaling by using random flipping, rotation, and brightness adjustment;
[0018] Perform normalization processing on the RGB image of the target working area after data augmentation.
[0019] Optionally, the step of performing object detection on the RGB image of the target working area by using a target detection model trained to a converged state to determine the central two-dimensional pixel coordinates corresponding to the palm fruits and / or pedestrians to be loaded includes:
[0020] Call an improved Yolov8 network model trained to a converged state, and input the RGB image of the target working area containing the palm fruits and / or pedestrians to be loaded into the improved Yolov8 network model;
[0021] Perform feature extraction through the DWConv module in the improved Yolov8 network model to obtain feature maps of different scales;
[0022] Use the Faster-PGLU module to process the features of different scales to strengthen the correlation between features, and identify the key features of the palm fruits and pedestrians to be loaded in the feature maps;
[0023] Use the detection head network to perform object localization and classification based on the key features of the palm fruits and pedestrians to be loaded to generate the bounding box coordinates of the palm fruits and pedestrians to be loaded;
[0024] Determine the central two-dimensional pixel coordinates corresponding to the palm fruits and / or pedestrians to be loaded according to the bounding box coordinates of the palm fruits and pedestrians to be loaded.
[0025] Optionally, the step of determining the central two-dimensional pixel coordinates corresponding to the palm fruits and / or pedestrians to be loaded according to the bounding box coordinates of the palm fruits and pedestrians to be loaded includes:
[0026] Determine the bounding box coordinates of the palm fruits to be loaded and the bounding box coordinates of the pedestrians, where the bounding box coordinates include the abscissa of the upper left corner of the bounding box, the ordinate of the upper left corner of the bounding box, the abscissa of the lower right corner of the bounding box, and the ordinate of the lower right corner of the bounding box;
[0027] Calculate and determine the first average value between the abscissa of the upper left corner of the bounding box of the palm fruits to be loaded and the abscissa of the lower right corner of the bounding box of the palm fruits to be loaded, and use the first average value as the abscissa of the central two-dimensional pixel coordinates of the palm fruits to be loaded;
[0028] Calculate and determine the second average value between the ordinate of the upper left corner of the bounding box of the palm fruit to be loaded and the ordinate of the lower right corner of the bounding box of the palm fruit to be loaded, and use the second average value as the ordinate of the central two-dimensional pixel coordinates of the palm fruit to be loaded;
[0029] Calculate and determine the third average value between the abscissa of the upper left corner of the bounding box of the pedestrian and the abscissa of the lower right corner of the bounding box of the pedestrian, and use the third average value as the abscissa of the central two-dimensional pixel coordinates of the pedestrian;
[0030] Calculate and determine the fourth average value between the ordinate of the upper left corner of the bounding box of the pedestrian and the ordinate of the lower right corner of the bounding box of the pedestrian, and use the fourth average value as the ordinate of the central two-dimensional pixel coordinates of the pedestrian to determine the corresponding central two-dimensional pixel coordinates of the palm fruit to be loaded and / or the pedestrian.
[0031] Optionally, the steps of aligning the RGB image of the target working area and the depth image of the target working area, and matching the depth value corresponding to the central two-dimensional pixel coordinates of the palm fruit in the RGB image of the target working area to determine the three-dimensional coordinates of the palm fruit to be loaded in the real space include:
[0032] Map the two-dimensional pixel coordinates into the three-dimensional coordinates of the palm fruit to be loaded in the real space, and perform normalization processing on the two-dimensional pixel coordinates, which includes:
[0033]
[0034] where, x p and y p are two-dimensional pixel coordinates; c p is the center of the image; f x , f y are the focal lengths of the image plane in the x and y directions respectively; x n , y n are normalized image coordinates;
[0035] Convert the normalized image coordinates and the depth value to the camera coordinates, which includes:
[0036] X c =x n *depth,
[0037] Y c =x n *depth,
[0038] Z c =depth,
[0039] where, Xc , Y c , Z c are the target three-dimensional coordinates in the camera coordinate system, and depth is the depth value obtained from the depth image.
[0040] Optionally, detecting whether the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operation space. If the central two-dimensional pixel coordinates corresponding to the pedestrian are outside the safe operation space, the steps of controlling the robotic arm to perform the loading operation according to the three-dimensional coordinates of the palm fruit to be loaded in the real space include:
[0041] Responding to the operable space monitoring instruction, judging whether the palm fruit to be loaded exceeds the operable space of the robotic arm based on the three-dimensional coordinates of the palm fruit to be loaded in the real space. If it exceeds, sending a warning instruction to the system interface of the palm fruit transporter to display that the current palm fruit to be loaded exceeds the operable space range on the system interface.
[0042] A palm fruit grasping and loading control device provided to meet another object of the present application includes:
[0043] An image acquisition module, configured to respond to the palm fruit grasping and loading control instruction, and acquire the RGB image of the target working area containing the palm fruit to be loaded and / or pedestrians in the palm fruit transporter and its corresponding depth image of the target working area;
[0044] A target detection module, configured to perform target detection on the RGB image of the target working area by using a target detection model trained to the convergence state to determine the central two-dimensional pixel coordinates corresponding to the palm fruit to be loaded and / or pedestrians. Among them, the basic network architecture of the target detection model is an improved Yolov8 network model. The improved Yolov8 network model replaces all Conv Module modules in the original Yolov8 network model with DWConv modules, and replaces all C2F modules in the original Yolov8 network model with Faster-PGLU modules. Among them, the Faster-PGLU module is constructed by combining convolutional gated linear units with the FasterNet structure introducing the PConv operator;
[0045] An image matching module, configured to align the RGB image of the target working area and the depth image of the target working area, and match the depth value corresponding to the central two-dimensional pixel coordinates of the palm fruit in the RGB image of the target working area to determine the three-dimensional coordinates of the palm fruit to be loaded in the real space;
[0046] The loading operation module is configured to determine the safe operation space of the robotic arm of the palm fruit transporter, detect whether the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operation space, and if the central two-dimensional pixel coordinates corresponding to the pedestrian are outside the safe operation space, control the robotic arm to perform the loading operation according to the three-dimensional coordinates of the palm fruit to be loaded in the real space, wherein the safe operation space is larger than the operable space of the robotic arm;
[0047] The safety control module is configured to send a robotic arm operation stop instruction to control the robotic arm to pause the loading operation if the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operation space, so as to complete the control of palm fruit grasping and loading.
[0048] An electronic device provided to meet another object of the present application includes a central processing unit and a memory, and the central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the palm fruit grasping and loading control method described in the present application.
[0049] A computer-readable storage medium provided to meet another object of the present application stores a computer program implemented according to the palm fruit grasping and loading control method in the form of computer-readable instructions, and when the computer program is called and run by a computer, it executes the steps included in the corresponding method.
[0050] Compared with the prior art, in the prior art, the loading and transportation of palm fruits are usually carried out manually. The traditional manual fruit loading operation has low efficiency and high management costs, and the complex and changing outdoor environment during the actual fruit grasping operation brings great challenges to palm fruit recognition. The present application includes, but is not limited to, the following beneficial effects:
[0051] First, accurate recognition and positioning of palm fruits are achieved in complex lighting and homogeneous dense stacking scenarios. The feedback results of inference using the validation set and test set data according to the improved Yolov8 network model show that the accuracy of the improved Yolov8 network model reaches 90.6%; according to the comparison of actual measurement data and algorithm positioning data, the positioning accuracy is less than 2 cm;
[0052] Second, a new lightweight convolution module is formed, and the structure of the detection model is designed in a lightweight manner, enabling the model to run on devices with low computing power. While maintaining the detection accuracy, the parameter quantity of our lightweight model is only 12.94% of the original model, and the calculation amount is only 22.22% of the original model;
[0053] Third, the palm orchard is divided into many small areas, each with a dedicated manager. The number of palm fruits produced in the area is directly related to the workers' wages. For this reason, we added a palm fruit counting and display function to the robot. In densely stacked fruit piles (greater than 50), the counting accuracy is ±1. Generally, the number of palm fruit piles in a palm orchard will not exceed 50, so accurate palm fruit counting can be achieved.
[0054] Fourthly, the hydraulic robot arm for grabbing fruits has a fixed working space due to its own length limitation. The hydraulic robot arm cannot grab the palm fruits outside the working space. We circle the fruit grabbing space of the robot arm in the real-time video detection screen to facilitate the adjustment of the relative position between the robot and the palm fruits. The palm fruits within the range of the line can be stably grabbed.
[0055] Fifth, to improve the safety of the machine, we added a pedestrian detection function. When a pedestrian is detected approaching the robot arm's workspace, the robot arm will stop immediately. We adjusted the camera angle to keep the camera in a suitable field of view for shooting. When a person's feet appear at the edge of the camera screen, the robot arm will stop working. This ensures that the robot arm will not accidentally injure pedestrians approaching.
[0056] Sixth, the operating speed of the robot is closely related to the degree of trust the user has in the robot and the operating efficiency of the robot. At present, the average automatic fruit grabbing speed of the robotic arm is comparable to the fruit grabbing speed of the robotic arm controlled by ordinary operators.
[0057] Seventh, by placing a camera in the middle position, the camera base can be rotated to allow the camera to obtain the three-dimensional coordinates of the palm fruits on the right and left sides, and guide the robotic arm to automatically grasp them. There is no need to place a camera on each side, thus saving costs.
[0058] Furthermore, by combining DWConv and Faster-PGLU modules, the improved YOLOv8 model has efficient feature extraction and powerful information expression capabilities when processing palm fruit and pedestrian target detection. This architecture can not only improve detection accuracy, but also optimize the use of computing resources, making it an ideal choice for automated operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0060] Figure 1 This is an exemplary architecture used by the palm fruit transport vehicle in the embodiments of the present application;
[0061] Figure 2 This is a schematic diagram of the fruit grabbing range of the robotic arm in the embodiment of the present application;
[0062] Figure 3 It is a schematic flowchart in the embodiments of the present application;
[0063] Figure 4 It is a schematic diagram of the improved Yolov8 network model in the embodiments of the present application;
[0064] Figure 5 It is a schematic structural diagram of the Faster-PGLU and CGLU modules in the embodiments of the present application;
[0065] Figure 6 It is a schematic diagram of the effect of pedestrian detection in the embodiments of the present application;
[0066] Figure 7 It is a schematic diagram of the interface of the target detection system in the embodiments of the present application;
[0067] Figure 8 It is a schematic framework diagram of the hierarchical control system of the palm fruit detection system in the embodiments of the present application;
[0068] Figure 9 It is a schematic framework diagram of the electrical structure of the palm fruit detection system in the embodiments of the present application;
[0069] Figure 10 It is a principle block diagram of the palm fruit grasping and loading control device in the embodiments of the present application;
[0070] Figure 11 It is a schematic structural diagram of the computer device in the embodiments of the present application. Detailed implementation manners
[0071] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation of the present application.
[0072] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.
[0073] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0074] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive and no ability to transmit, and devices with both receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices can include: cellular or other communication devices such as personal computers, tablets, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which can include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm-top computers or other devices, which are conventional laptop and / or palm-top computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, for example, it can be a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be a smart TV, a set-top box, etc.
[0075] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components revealed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.
[0076] It should be noted that the concept of "server" in this application can similarly be extended to the case of server clusters. According to the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be called through interfaces, or integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by it in the implementation of the network deployment method of this application.
[0077] One or several technical features of this application, unless expressly specified, can either be deployed on the server and accessed by the client remotely calling the online service interface provided by the server, or directly deployed and run on the client for access.
[0078] The neural network models cited or possibly cited in this application, unless expressly specified, can either be deployed on a remote server and remotely called by the client, or deployed on a client capable of handling the device and directly called. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid over-occupying the client's hardware operating resources.
[0079] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being called by the technical solution of this application.
[0080] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, these methods can be independently executed unless otherwise specified. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the expression is different, should be equivalently understood.
[0081] For the various embodiments to be disclosed in this application, unless expressly pointed out that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the needs in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.
[0082] Please refer to Figure 1, in the cab of the palm fruit transporter, a visual computing main control unit 1 and a display screen 2 are arranged. A camera 3 is installed behind the cab. To complete the automatic fruit-grabbing task, a fruit-grabbing robotic arm 4 and a palm fruit storage bin 5 are also installed on the transporter. Among them, the visual computing main control unit 1 and the camera 3 are connected by a USB3.0 transmission line; the visual computing main control unit 1 and the display screen 2 are connected by an HDMI to DP cable; a USB to CAN module is plugged into the visual computing main control unit 1, and the communication between the visual computing main control unit 1 and the robotic arm is connected through the USB to CAN module; the visual computing main control unit 1 can receive the RGB image and depth image captured by the camera 3 in real time and perform calculation and analysis using algorithms, and send the analysis results to the robotic arm control end through the USB to CAN module to guide the operation of the robotic arm. At the same time, it can rotate through the base of the camera to enable the camera to obtain the three-dimensional coordinate positions of the palm fruits on the right and left, guiding the robotic arm to automatically grab the palm fruits on both sides of the transporter. The display screen 2 can display the RGB image and depth information obtained by the camera 3 and processed by the built-in algorithm of the visual computing main control unit 1 on the screen for the driver to view;
[0083] Please refer to Figure 2 , in the picture of the display screen 2, the fruit-grabbing range of the robotic arm is marked with a line for the driver's reference, so as to drive the automatic fruit-grabbing transporter close to the palm fruit pile and make the palm fruits within the line of the fruit-grabbing range.
[0084] In some embodiments, the selected model of the visual computing main control unit 1 is the Jetson Orin Nano SUB development kit (4GB), the selected model of the camera 3 is the Intel Realsense D455f depth camera, and the USB to CAN module used is the USB-CAN of ViteSmart.
[0085] On the basis of referring to the above exemplary scenarios, please refer to Figure 3 , in one embodiment of the palm fruit grabbing and loading control method of the present application, it includes:
[0086] Step S10: Respond to the palm fruit grabbing and loading control instruction, and obtain the RGB image of the target working area containing the palm fruits to be loaded and / or pedestrians in the palm fruit transporter and its corresponding depth image of the target working area;
[0087] The visual computing master control 1 in the palm fruit transporter can respond to the palm fruit grasping and loading control instruction, and obtain the RGB image of the target working area containing the palm fruits to be loaded and / or pedestrians in the palm fruit transporter and its corresponding depth image of the target working area through the camera 3 in the palm fruit transporter; the type and source of the RGB image of the target working area are determined according to the actual application scenario. For example, in the application scenario of seedling tray sowing, the RGB image of the target working area can be a static picture specified by the user, or the RGB image of the target working area containing the palm fruits to be loaded and / or pedestrians submitted by the camera 3 in the palm fruit transporter to the visual computing master control 1. And so on, depending on the specific application scenario, the RGB image of the target working area can be determined as needed.
[0088] In some embodiments, before the step of obtaining the RGB image of the target working area containing the palm fruits to be loaded in the palm fruit transporter, it includes:
[0089] Step S101, in response to the image preprocessing instruction, perform image scaling, data augmentation, and normalization on the RGB image of the target working area containing the palm fruits to be loaded and / or pedestrians;
[0090] Step S102, adjust the RGB image of the target working area to a preset image size for image scaling;
[0091] Step S103, perform data augmentation on the RGB image of the target working area after image scaling by using random flipping, rotation, and brightness adjustment;
[0092] Step S104, perform normalization on the RGB image of the target working area after data augmentation.
[0093] In some embodiments, pictures of palm fruits are taken at the job site, and the dataset consists of these taken pictures. The detection model is trained with the dataset, and the performance of the model is highly related to the dataset. The pictures in the dataset should cover different working hours of the fruit-grabbing transfer vehicle and different lighting conditions; pictures are taken respectively during the morning, afternoon, and evening working hours of the fruit-grabbing transfer vehicle, and it is necessary to cover the weather and lighting conditions that may occur during the fruit-grabbing operation to take the dataset (for example, pictures need to be taken under strong sunlight on sunny days, cloudy days, and weak light in the evening); scenes with large background changes need to be covered, such as the color of the road, whether it is placed on a haystack, etc. A certain number of pictures need to be taken for each situation, for example, 100 pictures are taken for each situation. The camera 3 is fixed on the fruit-grabbing transfer vehicle. When the driver parks the fruit-grabbing transfer vehicle, the angle of the palm fruit photographed by the camera is roughly unchanged. Therefore, when taking pictures, the pictures can be taken at a fixed angle at the fixed height of the camera 3. At the same time, in order to implement the function of stopping the robotic arm after detecting a person, when taking pictures of all the above situations, pictures of a person appearing in the picture should also be taken, and the proportion of pictures with a person appearing is about 10% of each situation.
[0094] This part mainly makes the dataset by taking pictures at different times, different angles, different backgrounds, and at the height and angle of the camera 3. The dataset made in this way can enhance the model's adaptability to complex lighting environments, that is, effectively enhance the generalization of the model. (It is equivalent to letting the model see palm fruits in various environments and learn the characteristics of palm fruits in all working environments).
[0095] Furthermore, the open-source image annotation tool labelimg is used to annotate all the palm fruits and people in the dataset images. The annotation categories are set to two types, named Palm fruit and person respectively, and the label files are stored in XML format.
[0096] Furthermore, in order to improve the generalization of the dataset for recognizing palm fruits and people under different light intensity changes and different angles, the dataset is enhanced. The images in the dataset are processed by three different data enhancement methods: flipping, brightness change, and adding noise. Each picture is operated four times, and at least one data enhancement method takes effect each time. The number of dataset images is increased to 5 times the original. Then the label files in XML format are converted into TXT format. The training set, validation set, and test set are divided according to 7:2:1.
[0097] Step S20: Use the target detection model that has been trained to a converged state to perform target detection on the RGB image of the target working area, so as to determine the central two-dimensional pixel coordinates corresponding to the palm fruits and / or pedestrians to be loaded. Among them, the basic network architecture of the target detection model is an improved Yolov8 network model. The improved Yolov8 network model uses the DWConv module to replace all Conv Module modules in the original Yolov8 network model, and uses the Faster-PGLU module to replace all C2F modules in the original Yolov8 network model. Among them, the Faster-PGLU module is constructed by combining a convolutional gated linear unit with a FasterNet structure introducing the PConv operator;
[0098] Obtain the RGB image of the target working area containing the palm fruits and / or pedestrians to be loaded in the palm fruit transporter and its corresponding depth image of the target working area. Use the target detection model that has been trained to a converged state to perform target detection on the RGB image of the target working area, so as to determine the central two-dimensional pixel coordinates corresponding to the palm fruits and / or pedestrians to be loaded. Among them, the basic network architecture of the target detection model is an improved Yolov8 network model. The improved Yolov8 network model uses the DWConv module to replace all Conv Module modules in the original Yolov8 network model, and uses the Faster-PGLU module to replace all C2F modules in the original Yolov8 network model. Among them, the Faster-PGLU module is constructed by combining a convolutional gated linear unit with a FasterNet structure introducing the PConv operator;
[0099] In a specific embodiment, in order to run a complex detection algorithm on the vision computing master 1 with limited computing resources and reduce the computing burden on the vision computing master 1, such as Figure 4 shown, based on the YOLOv8 detection model, a new lightweight network structure is designed to construct an improved Yolov8 network model. The improved Yolov8 network model mainly includes:
[0100] 1. Adjust the model compound scaling constant: The Yolov8 model provides 5 models with different sizes and complexities, namely n, s, m, l, and x. The scaling constant used by the smallest-scale yolov8 is: the depth factor is set to 0.33, the width factor is set to 0.25, and the maximum number of channels is set to 1024. We set the depth factor of the improved Yolov8 network model to 0.165, the width factor to 0.125, and the maximum number of channels to 512.
[0101] 2. Use DWConv to replace all Conv Module modules in the original Yolov8 network model;
[0102] 3. Replace all C2F modules in the original Yolov8 network model with Faster-PGLU. As shown in Figure 5 , introduce the Convolutional Gated Linear Unit (CGLU), and combine the FasterNet structure using the PConv operator to form the Faster-PGLU module (Faster Partial Gated Linear Unit Model), and replace the original C2F module with the Faster-PGLU module.
[0103] 4. To compensate for the decrease in detection accuracy caused by the lightweight network, test the impact of different IOU on the detection accuracy, and select the IOU with the best performance. For the detection task of palm fruits, we replaced the initial CIOU with WIOU and achieved good detection results.
[0104] 5. Model training: Set the size of the image to 640, the training batch to 300, and the optimizer of the experiment to auto, allowing the model to automatically find the optimal optimizer and learning rate. After the trained model converges, a pt file result is obtained.
[0105] 6. Real-time inference of RGB images: Use the pt file obtained after model training to infer the video stream real-time collected by the camera (3). If palm fruits or people appear in the image, use a detection box to frame the palm fruits or people in the image.
[0106] 7. Output the two-dimensional pixel coordinates of the palm fruits to be loaded: Divide each frame of the image into a two-dimensional coordinate system according to h×w (height×width) pixels, and take the midpoints of all detection boxes of the palm fruits to be loaded to obtain the two-dimensional pixel coordinates of the centers of the palm fruits to be loaded.
[0107] The steps of using the object detection model trained to the convergence state to perform object detection on the RGB image of the target working area to determine the two-dimensional pixel coordinates of the centers corresponding to the palm fruits to be loaded and / or pedestrians include:
[0108] Step S201: Call the improved Yolov8 network model trained to the convergence state, and input the RGB image of the target working area containing the palm fruits to be loaded and / or pedestrians into the improved Yolov8 network model;
[0109] Step S202: Perform feature extraction through the DWConv module in the improved Yolov8 network model to obtain feature maps of different scales;
[0110] Step S203: Use the Faster-PGLU module to process the features of different scales to strengthen the correlation between features, and identify the key features of the palm fruits to be loaded and pedestrians in the feature map;
[0111] Step S204: Use the detection head network to perform target localization and classification based on the key features of the palm fruits to be loaded and the pedestrians, so as to generate the bounding box coordinates of the palm fruits to be loaded and the pedestrians;
[0112] Step S205: According to the bounding box coordinates of the palm fruits to be loaded and the pedestrians, determine the corresponding central two-dimensional pixel coordinates of the palm fruits to be loaded and / or pedestrians.
[0113] Furthermore, the step of determining the corresponding central two-dimensional pixel coordinates of the palm fruits to be loaded and / or pedestrians according to the bounding box coordinates of the palm fruits to be loaded and the pedestrians includes:
[0114] Step S2051: Determine the bounding box coordinates of the palm fruits to be loaded and the bounding box coordinates of the pedestrians, where the bounding box coordinates include the abscissa of the upper left corner of the bounding box, the ordinate of the upper left corner of the bounding box, the abscissa of the lower right corner of the bounding box, and the ordinate of the lower right corner of the bounding box;
[0115] Step S2052: Calculate and determine the first average value between the abscissa of the upper left corner of the bounding box of the palm fruits to be loaded and the abscissa of the lower right corner of the bounding box of the palm fruits to be loaded, and use the first average value as the abscissa of the central two-dimensional pixel coordinates of the palm fruits to be loaded;
[0116] Step S2053: Calculate and determine the second average value between the ordinate of the upper left corner of the bounding box of the palm fruits to be loaded and the ordinate of the lower right corner of the bounding box of the palm fruits to be loaded, and use the second average value as the ordinate of the central two-dimensional pixel coordinates of the palm fruits to be loaded;
[0117] Step S2054: Calculate and determine the third average value between the abscissa of the upper left corner of the bounding box of the pedestrians and the abscissa of the lower right corner of the bounding box of the pedestrians, and use the third average value as the abscissa of the central two-dimensional pixel coordinates of the pedestrians;
[0118] Step S2055: Calculate and determine the fourth average value between the ordinate of the upper left corner of the bounding box of the pedestrians and the ordinate of the lower right corner of the bounding box of the pedestrians, and use the fourth average value as the ordinate of the central two-dimensional pixel coordinates of the pedestrians, so as to determine the corresponding central two-dimensional pixel coordinates of the palm fruits to be loaded and / or pedestrians.
[0119] In the DWConv module of the improved Yolov8 network model, spatial features are extracted by performing individual convolution operations on each input channel, while 1x1 convolution is used for information fusion between channels. This approach significantly reduces the number of parameters and computational complexity, improving the efficiency of the network. The DWConv module can capture detailed features in the image, making the network more discriminative when detecting palm fruits and pedestrians.
[0120] In the Faster-PGLU module, by combining the convolutional gating mechanism, convolution operations are used to control the flow of information, thereby enhancing the network's ability to model complex features. By introducing the gating mechanism, the Faster-PGLU module can adaptively select important information in different parts of the feature map, thus improving the accuracy of object detection, especially in distinguishing palm fruits from pedestrians more effectively in a dense background.
[0121] The input image passes through the improved YOLOv8 model. First, it undergoes preliminary feature extraction through the DWConv module to obtain feature maps of different scales. The Faster-PGLU module is used to process the features of different scales, strengthening the correlation between features and ensuring that the key features of palm fruits and pedestrians can be effectively identified in the feature map. Through the final detection layer of the model, object localization and classification are performed to generate the bounding boxes of palm fruits and pedestrians and their corresponding confidence scores. After obtaining the predicted boxes, NMS is applied to remove boxes with high overlap and retain the boxes with the highest confidence to ensure the accuracy of the final detection results. The coordinates of the detected bounding boxes are converted into two-dimensional pixel coordinates of the original image for subsequent loading and processing operations. Finally, the model will output the two-dimensional pixel coordinates of the palm fruits and pedestrians to be loaded and transfer this information to downstream applications such as automated loading systems, monitoring systems, etc.
[0122] As can be seen from the above steps, by combining the DWConv and Faster-PGLU modules, the improved YOLOv8 model has efficient feature extraction and powerful information expression capabilities when dealing with the object detection of palm fruits and pedestrians to be loaded. This architecture can not only improve the detection accuracy but also optimize the use of computing resources, making it an ideal choice for realizing automated operations.
[0123] Step S30: Align the RGB image of the target working area and the depth image of the target working area, and match the depth value in the depth image of the target working area corresponding to the central two-dimensional pixel coordinates of the palm fruit in the RGB image of the target working area to determine the three-dimensional coordinates of the palm fruit to be loaded in the real space;
[0124] After determining the central two-dimensional pixel coordinates corresponding to the palm fruits to be loaded onto the vehicle and / or pedestrians, align the RGB image of the target working area with the depth image of the target working area, and match the depth value in the depth image of the target working area according to the central two-dimensional pixel coordinates of the palm fruits in the RGB image of the target working area, so as to determine the three-dimensional coordinates of the palm fruits to be loaded onto the vehicle in the real space;
[0125] Specifically, to obtain the internal and external parameters of the camera, the internal parameters of the camera can be obtained through the checkerboard calibration method. Some cameras come with special function packages to obtain the internal parameters of the camera; our camera is fixed on the vehicle body, and the pose of the palm fruit obtained is relative to the camera. It is necessary to convert the pose of the palm fruit relative to the camera to the world coordinate system of the robotic arm, and the coordinate conversion can be completed through the homogeneous transformation matrix.
[0126] Furthermore, the depth frame data can be directly called through the depth camera to determine the depth image of the target working area. The RGB image of the target working area and the depth image of the target working area can be aligned by calling the pyrealsense2 toolkit. Among them, pyrealsense2 is a Python library for interacting with the Intel RealSense camera series, allowing developers to use Python to obtain and process data from RealSense cameras. The RealSense camera is provided by Intel and is mainly used in applications such as depth perception, computer vision, and augmented reality. Through pyrealsense2, developers can easily access data such as the video stream, depth image, and point cloud captured by the camera and perform further processing and analysis.
[0127] Even further, map the two-dimensional pixel coordinates to the three-dimensional coordinates of the palm fruits to be loaded onto the vehicle in the real space, and perform normalization processing on the two-dimensional pixel coordinates, which includes:
[0128]
[0129] Among them, x p and y p are the two-dimensional pixel coordinates; c p is the center of the image, which is expressed as (c x , c y ); f x , f y are the focal lengths of the image plane in the x and y directions respectively; x n , y n are the normalized image coordinates;
[0130] Convert the normalized image coordinates and the depth value to the camera coordinates, which includes:
[0131] Xc = x n * depth,
[0132] Y c = x n * depth,
[0133] Z c = depth,
[0134] wherein, X c , Y c , Z c are the target three - dimensional coordinates in the camera coordinate system, and depth is the depth value obtained from the depth image.
[0135] Through the above steps, the central two - dimensional pixel coordinates of the to - be - loaded palm fruits obtained previously can be mapped and converted into the three - dimensional coordinates of the to - be - loaded palm fruits in the real space.
[0136] Step S40: Determine the safe operating space of the robotic arm of the palm fruit transporter, and detect whether the central two - dimensional pixel coordinates corresponding to the pedestrian are within the safe operating space. If the central two - dimensional pixel coordinates corresponding to the pedestrian are outside the safe operating space, control the robotic arm to perform the loading operation according to the three - dimensional coordinates of the to - be - loaded palm fruits in the real space, wherein the safe operating space is larger than the operable space of the robotic arm;
[0137] After the step of determining the three - dimensional coordinates of the to - be - loaded palm fruits in the real space, determine the safe operating space of the robotic arm of the palm fruit transporter, and detect whether the central two - dimensional pixel coordinates corresponding to the pedestrian are within the safe operating space. If the central two - dimensional pixel coordinates corresponding to the pedestrian are outside the safe operating space, control the robotic arm to perform the loading operation according to the three - dimensional coordinates of the to - be - loaded palm fruits in the real space, wherein the safe operating space is larger than the operable space of the robotic arm;
[0138] The robotic arm of the palm fruit transporter has a fixed safe operating space. To ensure the safe use of the machine, it should be avoided that the robotic arm accidentally injures people. So when we detect a person, we also record the central two - dimensional pixel coordinates of the person. When this central two - dimensional pixel coordinates approaches the safe operating space of the robotic arm, use a specific frame ID and send a message representing stopping the robotic arm through CAN communication. The effect of pedestrian detection is as Figure 6 shown.
[0139] In some embodiments, such as Figure 7As shown, the detection system interface can display the detection effect in real time, can achieve palm fruit counting, and can give a reminder when the palm fruit exceeds the grasping range. CAN communication has 8 data bits that can be used to send information. The vision computing main control 1 and the robotic arm 4 complete communication through a USB-to-CAN module. The USB-to-CAN module is set to the AT sending mode. The obtained three-dimensional coordinates (x, y, z) of the palm fruit in the real space are converted into 3 hexadecimal numbers, and then converted into the bytes format, and output in the low-order - high-order format. The low-order and high-order each account for 1 bit of 8 data bits, and 6 data bits can be used to output the coordinate information of (x, y, z).
[0140] In some embodiments, the step of detecting whether the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operation space, and if the central two-dimensional pixel coordinates corresponding to the pedestrian are outside the safe operation space, controlling the robotic arm to perform the loading operation according to the three-dimensional coordinates of the palm fruit to be loaded in the real space, includes:
[0141] Responding to the operable space monitoring instruction, judging whether the palm fruit to be loaded exceeds the operable space of the robotic arm based on the three-dimensional coordinates of the palm fruit to be loaded in the real space. If it exceeds, a warning instruction is sent to the system interface of the palm fruit transporter to display on the system interface that the current palm fruit to be loaded exceeds the operable space range.
[0142] In some embodiments, Figure 8 is the hierarchical control system framework diagram of the palm fruit detection system; Figure 9 is the electrical structure schematic diagram of the palm fruit detection system.
[0143] Step S50: If the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operation space, a robotic arm operation stop instruction is sent to control the robotic arm to pause the loading operation to complete the control of grasping and loading the palm fruit.
[0144] After controlling the robotic arm to perform the loading operation according to the three-dimensional coordinates of the palm fruit to be loaded in the real space, if the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operation space, a robotic arm operation stop instruction is sent to control the robotic arm to pause the loading operation to complete the control of grasping and loading the palm fruit.
[0145] As can be seen from the above embodiments, compared with the prior art, in view of the problems in the prior art that the loading and transportation of palm fruits are usually carried out manually, the traditional manual fruit loading operation has low efficiency and high management cost, and the complex and changing outdoor environment during the actual fruit grasping operation brings great challenges to palm fruit recognition, etc., the present application includes but is not limited to the following beneficial effects:
[0146] First, accurate identification and positioning of palm fruits have been achieved in complex lighting and homogeneous dense stacking scenarios. The feedback results of inference using the validation set and test set data based on the improved Yolov8 network model show that the accuracy of the improved Yolov8 network model reaches 90.6%; according to the comparison between the actual measurement data and the algorithm positioning data, the positioning accuracy is less than 2 cm;
[0147] Second, a new lightweight convolution module has been built, and the structure of the detection model has been lightweight designed, enabling the model to run on devices with low computing power. While maintaining the detection accuracy, the number of parameters of our lightweight model is only 12.94% of the original model, and the computational volume is only 22.22% of the original model;
[0148] Third, many small areas have been divided in the palm garden, and each area has a dedicated management staff. The quantity of palm fruits produced in this area is directly related to the workers' wages. Therefore, we added a palm fruit counting and display function to the robot. In the densely stacked fruit piles (more than 50), the counting accuracy is ±1. Generally, the number of palm fruit piles in the palm fruit garden will not exceed 50, and accurate palm fruit counting can be achieved.
[0149] Fourth, due to its own length limitation, the fruit-gripping hydraulic manipulator has a fixed working space, and the hydraulic manipulator cannot grasp palm fruits outside the working space. We circled the fruit-gripping space of the manipulator in the real-time video detection screen to facilitate adjusting the relative position between the robot and the palm fruits. The palm fruits within the drawn line can be stably grasped.
[0150] Fifth, to improve the safety of the machine, we added a pedestrian detection function. When a pedestrian is detected approaching the working space of the manipulator, the manipulator will stop immediately. We adjusted the angle of the camera to keep the camera at an appropriate field of view for shooting, and the manipulator will stop working when the feet of a person appear at the edge of the camera screen. It can ensure that the manipulator will not accidentally injure the approaching pedestrians.
[0151] Sixth, the operating speed of the robot is closely related to the trust of the user and the operating efficiency of the robot. Currently, the average automatic fruit-gripping speed of the manipulator is comparable to that of an ordinary operator controlling the manipulator.
[0152] Seventh, a camera is arranged in the middle position, which can rotate through the camera base to enable the camera to obtain the three-dimensional coordinate positions of palm fruits on the right and left, guiding the manipulator to perform automatic grasping, without the need to arrange a camera on each side, saving expenses.
[0153] Furthermore, by combining the DWConv and Faster-PGLU modules, the improved YOLOv8 model has efficient feature extraction and powerful information expression capabilities when dealing with the detection of palm fruits to be loaded and pedestrians. This architecture can not only improve the detection accuracy but also optimize the use of computing resources, making it an ideal choice for realizing automated operations.
[0154] Please refer to Figure 10 , a palm fruit grasping and loading control device provided to meet one of the purposes of this application, includes an image acquisition module 1100, a target detection module 1200, an image matching module 1300, a loading operation module 1400, and a safety control module 1500. Among them, the image acquisition module 1100 is set to respond to the palm fruit grasping and loading control instruction to acquire the RGB image of the target working area containing the palm fruits to be loaded and / or pedestrians in the palm fruit transporter and its corresponding depth image of the target working area; the target detection module 1200 is set to use the target detection model trained to the convergence state to perform target detection on the RGB image of the target working area to determine the central two-dimensional pixel coordinates corresponding to the palm fruits to be loaded and / or pedestrians. Among them, the basic network architecture of the target detection model is an improved Yolov8 network model. The improved Yolov8 network model replaces all Conv Module modules in the original Yolov8 network model with DWConv modules and replaces all C2F modules in the original Yolov8 network model with Faster-PGLU modules. Among them, the Faster-PGLU module is constructed by combining a convolutional gated linear unit with a FasterNet structure introducing a PConv operator; the image matching module 1300 is set to align the RGB image of the target working area and the depth image of the target working area, and match the depth value in the corresponding depth image of the target working area according to the central two-dimensional pixel coordinates of the palm fruits in the RGB image of the target working area to determine the three-dimensional coordinates of the palm fruits to be loaded in the real space; the loading operation module 1400 is set to determine the safe operation space of the robotic arm of the palm fruit transporter, detect whether the central two-dimensional pixel coordinates corresponding to the pedestrians are within the safe operation space. If the central two-dimensional pixel coordinates corresponding to the pedestrians are outside the safe operation space, control the robotic arm to perform the loading operation according to the three-dimensional coordinates of the palm fruits to be loaded in the real space, where the safe operation space is larger than the operable space of the robotic arm; the safety control module 1500 is set to send a robotic arm operation stop instruction to control the robotic arm to pause the loading operation if the central two-dimensional pixel coordinates corresponding to the pedestrians are within the safe operation space, so as to complete the control of palm fruit grasping and loading.
[0155] Based on any embodiment of this application, please refer toFigure 11 , another embodiment of the present application further provides an electronic device, which can be implemented by a computer device. As shown in Figure 11 , a schematic internal structure diagram of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The control information sequence can be stored in the database. When the computer-readable instructions are executed by the processor, the processor can implement a palm fruit grasping and loading control method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. Computer-readable instructions can be stored in the memory of the computer device. When the computer-readable instructions are executed by the processor, the processor can execute the palm fruit grasping and loading control method of the present application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand that Figure 11 , the structure shown in
[0156] In this embodiment, the processor is used to execute Figure 10 , the specific functions of each module and its sub-modules. The memory stores the program code and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program code and data required to execute all modules / sub-modules in the palm fruit grasping and loading control device of the present application. The server can call the program code and data of the server to execute the functions of all sub-modules.
[0157] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the palm fruit grasping and loading control method described in any embodiment of the present application.
[0158] The present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the palm fruit grasping and loading control method described in any embodiment of the present application are implemented.
[0159] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments of the method of this application can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of each method. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0160] The above are only some embodiments of this application. It should be noted that for those of ordinary skill in the technical field, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
[0161] In summary, by combining the DWConv and Faster-PGLU modules, the improved YOLOv8 model in this application has efficient feature extraction and powerful information expression capabilities when processing palm fruits and pedestrian target detections to be loaded onto the vehicle. This architecture can not only improve the detection accuracy but also optimize the use of computing resources, making it an ideal choice for realizing automated operations.
Claims
1. A palm fruit grabbing and loading control method, characterized in that: include: In response to the palm fruit grabbing and loading control instruction, an RGB image of a target working area of the palm fruit transfer vehicle containing palm fruits to be loaded and / or pedestrians and a corresponding depth image of the target working area are obtained; Using a target detection model that has been trained to a convergent state to perform target detection on the RGB image of the target working area to determine the central two-dimensional pixel coordinates corresponding to the palm fruits and / or pedestrians to be loaded, wherein the basic network architecture of the target detection model is an improved Yolov8 network model, and the improved Yolov8 network model uses a DWConv module to replace all Conv Module modules in the original Yolov8 network model, and uses a Faster-PGLU module to replace all C2F modules in the original Yolov8 network model, wherein the Faster-PGLU module is constructed by combining a convolutional gated linear unit with a FasterNet structure that introduces a PConv operator; Aligning the target working area RGB image and the target working area depth image, matching the central two-dimensional pixel coordinates of the palm fruits in the target working area RGB image with the depth values in the target working area depth image, so as to determine the three-dimensional coordinates of the palm fruits to be loaded in the real space; Determine the safe operating space of the mechanical arm of the palm fruit transporter, detect whether the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operating space, and if the central two-dimensional pixel coordinates corresponding to the pedestrian are outside the safe operating space, control the mechanical arm to perform loading operations according to the three-dimensional coordinates of the palm fruits to be loaded in the real space, wherein the safe operating space is larger than the operable space of the mechanical arm; If the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operating space, a robot arm operation stop instruction is sent to control the robot arm to suspend the loading operation to complete the control of palm fruit grabbing and loading.
2. The palm fruit grabbing and loading control method according to claim 1 is characterized in that: The improved Yolov8 network model has a depth factor of 0.165, a width factor of 0.125, and a maximum number of channels of 512.
3. The palm fruit grabbing and loading control method according to claim 1 is characterized in that: The step of obtaining an RGB image of a target working area of a palm fruit transfer vehicle containing palm fruits to be loaded includes: In response to the image preprocessing instruction, image scaling, data enhancement and normalization processing are performed on the RGB image of the target working area including the palm fruits to be loaded and / or pedestrians; Adjusting the target working area RGB image to a preset image size for image scaling; The data of the RGB image of the target working area after image scaling is enhanced by random flipping, rotation and brightness adjustment; The RGB image of the target working area after data enhancement is normalized.
4. The palm fruit grabbing and loading control method according to claim 1 is characterized in that: The step of using the target detection model that has been trained to a convergent state to perform target detection on the RGB image of the target working area to determine the central two-dimensional pixel coordinates corresponding to the palm fruits to be loaded and / or pedestrians includes: Calling an improved Yolov8 network model that has been trained to a convergence state, and inputting the RGB image of the target working area including the palm fruits to be loaded and / or pedestrians into the improved Yolov8 network model; Feature extraction is performed through the DWConv module in the improved Yolov8 network model to obtain feature maps of different scales; The Faster-PGLU module is used to process the features of different scales to strengthen the correlation between the features, and the key features of the palm fruits to be loaded and the pedestrians are identified in the feature map; Using a detection head network to perform target positioning and classification according to key features of the palm fruits to be loaded and the pedestrian, so as to generate bounding box coordinates of the palm fruits to be loaded and the pedestrian; The central two-dimensional pixel coordinates corresponding to the palm fruits to be loaded and / or the pedestrian are determined according to the boundary box coordinates of the palm fruits to be loaded and the pedestrian.
5. The palm fruit grabbing and loading control method according to claim 4 is characterized in that: The step of determining the central two-dimensional pixel coordinates corresponding to the palm fruits to be loaded and / or the pedestrian according to the bounding box coordinates of the palm fruits to be loaded and the pedestrian comprises: Determine the bounding box coordinates of the palm fruits to be loaded and the bounding box coordinates of the pedestrian, wherein the bounding box coordinates include the upper left corner horizontal coordinate of the bounding box, the upper left corner vertical coordinate of the bounding box, the lower right corner horizontal coordinate of the bounding box, and the lower right corner vertical coordinate of the bounding box; Calculate and determine a first average value between the abscissa of the upper left corner of the bounding box of the palm fruits to be loaded and the abscissa of the lower right corner of the bounding box of the palm fruits to be loaded, and use the first average value as the abscissa of the central two-dimensional pixel coordinate of the palm fruits to be loaded; Calculate and determine a second average value between the ordinate of the upper left corner of the bounding box of the palm fruits to be loaded and the ordinate of the lower right corner of the bounding box of the palm fruits to be loaded, and use the second average value as the ordinate of the central two-dimensional pixel coordinate of the palm fruits to be loaded; Calculate and determine a third average value between the abscissa of the upper left corner of the pedestrian's bounding box and the abscissa of the lower right corner of the pedestrian's bounding box, and use the third average value as the abscissa of the central two-dimensional pixel coordinate of the pedestrian; Calculate and determine the fourth average value between the upper left corner ordinate of the pedestrian's bounding box and the lower right corner ordinate of the pedestrian's bounding box, and use the fourth average value as the ordinate of the central two-dimensional pixel coordinate of the pedestrian to determine the central two-dimensional pixel coordinate corresponding to the palm fruits to be loaded and / or the pedestrian.
6. The palm fruit grabbing and loading control method according to claim 1 is characterized in that: The step of aligning the target working area RGB image and the target working area depth image, and matching the central two-dimensional pixel coordinates of the palm fruits in the target working area RGB image with the depth values in the target working area depth image to determine the three-dimensional coordinates of the palm fruits to be loaded in the real space comprises: Mapping the two-dimensional pixel coordinates into the three-dimensional coordinates of the palm fruit to be loaded in the real space, and normalizing the two-dimensional pixel coordinates, which includes: Among them, x p and p is the two-dimensional pixel coordinate; c p is the center of the image, which is represented by (c x , c y );f x 、f y are the focal lengths of the image plane in the x and y directions respectively; n ,y n are the normalized image coordinates; Convert the normalized image coordinates and depth values to camera coordinates, which includes: X c =x n *depth, Y c =x n *depth, WITH c =depth, Among them, X c , Y c , Z c is the target 3D coordinate in the camera coordinate system, and depth is the depth value obtained from the depth image.
7. The palm fruit grabbing and loading control method according to claim 1, characterized in that: The step of detecting whether the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operating space, and if the central two-dimensional pixel coordinates corresponding to the pedestrian are outside the safe operating space, controlling the robot arm to perform loading operations according to the three-dimensional coordinates of the palm fruits to be loaded in the real space, comprises: In response to the operating space monitoring instruction, it is determined whether the palm fruits to be loaded exceed the operating space of the robotic arm based on the three-dimensional coordinates of the palm fruits to be loaded in the real space. If exceeded, a warning instruction is sent to the system interface of the palm fruit transfer vehicle to display on the system interface that the palm fruits to be loaded currently exceed the operating space.
8. A palm fruit grabbing and loading control device, characterized in that: include: An image acquisition module is configured to respond to a palm fruit grabbing and loading control instruction to acquire an RGB image of a target working area of a palm fruit transfer vehicle containing palm fruits to be loaded and / or pedestrians and a corresponding depth image of the target working area; A target detection module is configured to use a target detection model that has been trained to a convergence state to perform target detection on the RGB image of the target working area to determine the central two-dimensional pixel coordinates corresponding to the palm fruits to be loaded and / or pedestrians, wherein the basic network architecture of the target detection model is an improved Yolov8 network model, and the improved Yolov8 network model uses a DWConv module to replace all Conv Module modules in the original Yolov8 network model, and uses a Faster-PGLU module to replace all C2F modules in the original Yolov8 network model, wherein the Faster-PGLU module is constructed by combining a convolutional gated linear unit with a FasterNet structure that introduces a PConv operator; An image matching module is configured to align the target working area RGB image and the target working area depth image, and match the central two-dimensional pixel coordinates of the palm fruits in the target working area RGB image with the depth values in the target working area depth image to determine the three-dimensional coordinates of the palm fruits to be loaded in real space; a loading operation module, configured to determine a safe operating space of a mechanical arm of the palm fruit transporter, detect whether the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operating space, and if the central two-dimensional pixel coordinates corresponding to the pedestrian are outside the safe operating space, control the mechanical arm to perform a loading operation according to the three-dimensional coordinates of the palm fruit to be loaded in real space, wherein the safe operating space is larger than the operable space of the mechanical arm; The safety control module is configured to send a robot arm operation stop instruction to control the robot arm to suspend the loading operation if the central two-dimensional pixel coordinates corresponding to the pedestrian are within the safe operation space, so as to complete the control of palm fruit grabbing and loading.
9. An electronic device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
Citation Information
Patent Citations
Loading and unloading position detection method, device and system based on binocular depth camera
CN112614191A
Excavator bucket position determining method and device and excavator
CN118485719A