Intensive scene mechanical arm grabbing method, system and device based on machine vision

Through the improved YOLOv8 model and multi-scale attention mechanism, the problem of uncertainty in traditional robotic arms grabbing in dense scenes is solved, and efficient and safe target recognition and grabbing operations are achieved.

CN120259840APending Publication Date: 2025-07-04QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510335141.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional robotic arms lack flexible adaptability to environmental changes in dense scenarios, resulting in failed grabbing or operating errors, which may collide important equipment or samples, causing losses.

Method used

Using a machine vision-based method, the improved YOLOv8 model is used to obtain target segmentation and pose data, combined with multi-scale attention module, separable nuclear attention mechanism and multi-scale expansion attention mechanism, to improve the accuracy of target recognition and capture.

Benefits of technology

It improves the operating efficiency and safety of the robotic arms in complex environments, reduces the risk of collision, and ensures the accuracy and stability of the grasping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259840A_ABST
    Figure CN120259840A_ABST
Patent Text Reader

Abstract

The invention provides a dense scene mechanical arm grabbing method, system and device based on machine vision, and relates to the technical field of machine vision, and the grabbing method comprises the steps: obtaining a target image and target geometric feature data; segmenting the target image based on an improved YOLOv8 model to obtain a target area, extracting geometric center point coordinates of the target area and performing mean value processing to obtain target center point coordinates; obtaining target pose data based on the target center point coordinate and the target geometric feature data; and target grabbing pose information is determined based on the target pose data, and a mechanical arm is controlled to execute grabbing operation according to the target grabbing pose information. The grabbing device integrates an image processing technology, a machine learning algorithm and precise mechanical control, and can be widely applied to scenes such as warehouse logistics, intelligent manufacturing and automatic production lines so as to improve the production efficiency and the operation safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision technology, and particularly to a dense-scene robotic arm grasping method, system, and device based on machine vision. Background Art

[0002] The statements in this section merely provide background technologies related to the present disclosure and do not necessarily constitute prior art.

[0003] Robotic arm grasping systems may operate in various dense scenes, such as automated warehouses, precision manufacturing factories, complex assembly lines, and scientific research laboratories, etc. These places have extremely high requirements for the precise grasping and placement of items. Especially in industrial environments, in order to ensure work safety and precise operation and increase work efficiency, highly precise and reliable grasping means are required.

[0004] However, when traditional robotic arms perform grasping tasks, they often lack flexible adaptability to environmental changes and may cause grasping failures or operation errors due to unexpected obstacles or the movement of target objects. This uncertainty may cause the robotic arm to collide with important equipment or samples during operation in dense scenes, resulting in unnecessary losses. Summary of the Invention

[0005] To solve the above problems, the present invention provides a dense-scene robotic arm grasping method, system, and device based on machine vision, which improve the efficiency and accuracy of automated operations and reduce the risks during operation in complex environments by integrating machine vision, image processing, and precision motion control technologies.

[0006] The first aspect of the present invention provides a dense-scene grasping method based on machine vision, including:

[0007] S101: Obtain a target image and target geometric feature data;

[0008] S102: Segment the target image based on an improved YOLOv8 model to obtain a target region, extract the geometric center point coordinates of the target region and perform mean processing to obtain target center point coordinates; in the improved YOLOv8 model, a multi-scale attention module C2f_EMA is added after the second convolution in the C2f module in YOLOv8 to replace the original C2f module;

[0009] S103: Obtain target pose data based on the target center point coordinates and target geometric feature data;

[0010] S104: Determine target grasping pose information based on the target pose data, and control the robotic arm to perform a grasping operation according to the target grasping pose information.

[0011] Further, the multi-scale attention module C2f_EMA includes: a CBS module, a splitter, n Bottleneck modules, a first connector, a CBS module, and an EMA module connected in sequence;

[0012] Among them, the output ends of the n Bottleneck modules are all connected to the input end of the first connector.

[0013] Further, in the improved YOLOv8 model, the original SPPF module is replaced by the SPPF_LSKA module.

[0014] Further, the SPPF_LSKA module includes: an input layer, a CBS module, a first max pooling layer, a second max pooling layer, a third max pooling layer, a second connector, an LSKA module, a CBS module, and an output layer connected in sequence;

[0015] Among them, the output ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer are all connected to the input end of the second connector.

[0016] Further, in the improved YOLOv8 model, an MSDA module is added before each of the four detection heads of the YOLOv8 model.

[0017] Further, the MSDA module divides the channels of the feature map into multiple independent heads, and each head is assigned a specific task to perform self-attention operations at different dilation rates.

[0018] The second aspect of the present invention provides a robotic arm grasping system for dense scenes based on machine vision, including:

[0019] A data acquisition module for acquiring target images and target geometric feature data;

[0020] A target recognition module for segmenting the target image based on the improved YOLOv8 model to obtain a target area, extracting the geometric center point coordinates of the target area and performing mean processing to obtain target center point coordinates; in the improved YOLOv8 model, a multi-scale attention module C2f_EMA is added after the second convolution in the C2f module in YOLOv8 to replace the original C2f module;

[0021] A positioning module for obtaining target pose data based on the target center point coordinates and target geometric feature data;

[0022] A grasping module for determining target grasping pose information based on the target pose data and controlling the robotic arm to perform a grasping operation according to the target pose data.

[0023] The third aspect of the present invention provides a robotic arm grasping device for dense scenes based on machine vision. The device includes a memory and a processor; the memory is used to store computer programs; the processor is used to implement the above-mentioned grasping method for dense scenes based on machine vision when executing the computer programs.

[0024] Furthermore, the device further includes an object monitoring device for real-time monitoring of the position and state of objects in the working area, a robotic arm device for performing grasping actions, a power supply device, and a human-machine interaction interface for visualization.

[0025] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned grasping method for robotic arms in dense scenes based on machine vision is implemented.

[0026] Compared with the prior art, the grasping method, system, and device for robotic arms in dense scenes based on machine vision provided by the present invention have the following beneficial effects:

[0027] 1. Adding the EMA attention mechanism behind the second convolution of the C2f backbone is more conducive to globally adjusting feature information, with lower computational costs and higher training stability.

[0028] 2. To address the problem of target occlusion in dense scenes, a large separable kernel attention mechanism (SPPF_LSKA) is integrated into the SPPF of YOLOv8. This module uses convolutional kernels of different sizes to create different receptive fields. It enhances the extraction of local spatial information in the image and effectively improves the feature extraction of occluded objects.

[0029] 3. To address the problem of object scale variation in dense scenes, an MSDA module is added before the detection head of YOLOv8. This addition enables the model to focus more on important features during feature integration and reduces the influence of the background. In addition, the fourth detection head enhances the ability to extract information from the detected targets and optimizes the recognition ability of the network. Description of the Drawings

[0030] The specification drawings constituting a part of this disclosure are used to provide a further understanding of this disclosure. The schematic embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an improper limitation of this disclosure.

[0031] Figure 1 It is a flowchart of the method for Embodiment 1;

[0032] Figure 2 It is a schematic example diagram of the YOLOv8 model;

[0033] Figure 3Schematic diagram of the improved YOLOv8 model provided for Example 1;

[0034] Figure 4 Schematic diagram of the C2f_EMA module in the improved YOLOv8 model provided for Example 1;

[0035] Figure 5 Schematic diagram of the SPPF_LSKA module in the improved YOLOv8 model provided for Example 1;

[0036] Figure 6 Schematic diagram of the MSDA in the improved YOLOv8 model provided for Example 1;

[0037] Figure 7 Diagram of the device connection relationship provided for Example 3;

[0038] Figure 8 Schematic diagram of the robotic arm device provided for Example 3;

[0039] Figure 9 Effect of the method provided by the present invention when applied to a dense scenario Figure 1 ;

[0040] Figure 10 Effect of the method provided by the present invention when applied to a dense scenario Figure 2 ;

[0041] Figure 11 Instance ratio of each category in a training dataset provided by the present invention;

[0042] Among them, 1. Industrial control computer, 2. Depth camera, 3. Image processing unit, 4. Improved YOLOv8 model, 5. Inference and decision-making unit, 6. Control unit, 7. Robotic arm device, 8. Communication module, 9. Other expansion modules, 10. Human-machine interaction interface, 11. Training dataset, 12. Power supply device, 13. Robotiq85 gripper, 14. RealsenseD435 camera, 15. Aobo I5 robotic arm. Detailed implementation manners

[0043] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0044] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0045] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0046] All data acquisition in this embodiment is based on compliance with laws and regulations and user consent, and is a legal application of the data.

[0047] Embodiment 1

[0048] Please refer to the attached instruction Figures 1-6 , Figure 1 which shows a flowchart of a method for a robotic arm to grasp in a dense scene based on machine vision provided by the present invention:

[0049] A method for a robotic arm to grasp in a dense scene based on machine vision provided by the present invention includes:

[0050] S101: Obtain a target image and target geometric feature data;

[0051] S102: Segment the target image based on an improved YOLOv8 model to obtain a target region, extract the geometric center point coordinates of the target region and perform mean processing to obtain target center point coordinates; in the improved YOLOv8 model, a multi-scale attention module C2f_EMA is added after the second convolution in the C2f module in YOLOv8 to replace the original C2f module;

[0052] S103: Obtain target pose data based on the target center point coordinates and target geometric feature data;

[0053] S104: Determine target grasping pose information based on the target pose data, and control the robotic arm to perform a grasping operation according to the target grasping pose information.

[0054] Figure 2It shows the overall network structure of YOLOv8. Similar to YOLOv5, it adopts the design idea of PAN. First, it performs preliminary feature extraction through the Backbone. Then, it enhances the features through a Neck part that goes from bottom to top and from top to bottom. Finally, it outputs feature maps of three sizes and then makes predictions on these three-sized feature maps respectively.

[0055] Specifically, first is the Backbone part, including: input an RGB image of 640*640, reduce its size by half through a CBS convolution, and then change the number of channels from 3 to 64; then through a combination of a CBS and a C2f, reduce the size of the picture from 320 to 160, and then double the number of channels. Among them, the role of CBS is to reduce the size to half of the original and double the number of channels, while C2f keeps the picture size and the number of channels unchanged; successively pass through 4 combined modules of CBS and C2f, and finally the size of the picture becomes 20*20 and the number of channels becomes 1024; next, pass through the SPPF module for a series of pooling connection operations.

[0056] Then, it is the Neck part, including performing an upsampling operation on the features extracted from the top layer, expanding the picture from 20*20 to 40*40 to facilitate a Concat operation with the 6th layer in the Backbone part because the Concat operation requires their sizes to be the same. In this way, a feature map of 40*40*(1024 + 512 = 1536 channels) is obtained, and then it undergoes feature enhancement through C2f and reduces the number of channels to 512. Then, it performs another upsampling, from 40 to 80, and conducts a Concat with the 4th layer of the Backbone to obtain a feature map of 80*80*768, and then undergoes feature enhancement through a C2f, which completes the feature fusion from bottom to top. Next, following the same principle, starting from the topmost feature map and going down, it is scaled to 40*40 to perform a Concat with the result of the just-mentioned C2f; then it undergoes scaling and feature enhancement through C2f and CBS.

[0057] Finally, it obtains the output of the feature maps of three sizes in the Head part, which are 20*20, 40*40, and 80*80 respectively. Among them, the small feature map of 20*20 is used to predict large targets, the medium feature map of 40*40 is used to predict medium targets, and the large feature map of 80*80 is used to predict small targets.

[0058] Figure 3The overall network structure of the improved YOLOv8 is shown. In the improved YOLOv8 model, a multi-scale attention module C2f_EMA is added after the second convolution in the C2f module in YOLOv8 to replace the original C2f module.

[0059] The multi-scale attention module C2f_EMA includes: a CBS module, a splitter, n Bottleneck modules, a first connector, a CBS module, and an EMA module connected in sequence;

[0060] Among them, the output ends of the n Bottleneck modules are all connected to the input end of the first connector.

[0061] To improve the feature extraction ability of the model, we added an efficient multi-scale attention module (EMA) after the second convolution of the C2f module. This method reshapes a part of the channel dimension into the batch dimension and performs grouped processing, effectively avoiding the side effects that may be introduced in the process of channel dimensionality reduction by the traditional attention mechanism. This design retains the key information of each channel, ensures that important data will not be lost during the feature extraction process of the model, and greatly improves the feature extraction ability of the model.

[0062] By adding the EMA attention mechanism behind the second convolution of the C2f backbone, it is more conducive to globally adjusting the feature information, with lower computational cost and higher training stability.

[0063] In the improved YOLOv8 model, the original SPPF module is replaced by the SPPF_LSKA module.

[0064] The SPPF_LSKA module includes: an input layer, a CBS module, a first max pooling layer, a second max pooling layer, a third max pooling layer, a second connector, an LSKA module, a CBS module, and an output layer connected in sequence;

[0065] Among them, the output ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer are all connected to the input end of the second connector.

[0066] In dense scenarios, object occlusion poses challenges to detection, while the LSKA (Large Separable Kernel Attention) mechanism cleverly combines the design of large separable convolutional kernels with the characteristics of spatially dilated convolutions. This process generates detailed attention maps and intelligently weights the original features, thereby significantly enhancing the network's attention to key features. Consequently, the model exhibits obvious advantages in handling complex visual tasks, especially in cases of dense occlusion. Given the remarkable advantages of the LSKA module, we place it before the second convolutional layer (Conv2) of the SPPF module, immediately following the operation of the max pooling layer (MaxPool2d). During the forward propagation process, the data stream first undergoes feature extraction through the first convolutional layer (cv1), followed by three consecutive max pooling layers, each responsible for reducing the spatial dimension and extracting key features. To make full use of the information from each pooling level, we fuse the outputs of these three max pooling layers through a concatenation operation to form a richer and more comprehensive feature representation. This step effectively promotes cross-scale integration of information. Subsequently, the fused feature map is sent to the LSKA module for processing. It further explores the relationships and importance among features, adaptively adjusts the feature weights, and enhances the model's ability to capture important information while suppressing unnecessary noise. Finally, the optimized feature map processed by the LSKA module is sent to the second convolutional layer (cv2) for deeper feature extraction and transformation, generating the final feature representation for the detection task.

[0067] To address the problem of object occlusion in dense scenarios, the Large Separable Kernel Attention mechanism (SPPF_LSKA) is integrated into the SPPF of YOLOv8. This module uses convolutional kernels of different sizes to create different receptive fields. It enhances the extraction of local spatial information in the image and effectively improves the feature extraction of occluded objects.

[0068] For the improved YOLOv8 model, the MSDA module is added before each of the four detection heads of the YOLOv8 model.

[0069] To improve the multi-scale feature integration of the model, we apply the Multi-Scale Dilated Attention mechanism (MSDA) before the detection heads of the YOLOv8 model. This method aggregates semantic information at different scales and captures multi-scale features. It is crucial for understanding various abstract levels and details of the image and helps the model process complex data. The outputs of MSDA are merged through concatenation and aggregated through a linear layer. This aggregation method integrates the information learned by each head, thereby generating a richer feature representation and improving the model's expressive ability and accuracy.

[0070] MSDA aims to address the redundancy problem in the global dependency modeling of shallow features in ViT models. Its mechanism combines the essence of scale diversity with the attention mechanism. In this mechanism, the channels of the feature map are divided into multiple independent heads, and each head is assigned a specific task to perform self-attention operations at different dilation rates. Specifically, in each head, the self-attention operation is carried out in a dynamically resized window around the red query block, defining different receptive fields (e.g., 3×3, 5×5, 7×7). This strategy enables each head to focus on a specific scale of image features, from fine local textures to broad context information. Then, the features captured at different dilation rates are concatenated to form a cross-scale feature representation. This feature not only contains details at all levels of the image but also rich context relationships. Finally, this information is sent to a linear layer for further aggregation and refinement, so as to gain a deeper and more comprehensive understanding of the image content. Based on the advantages of MSDA, we added MSDA before the 4 detection heads to enhance the model's multi-scale information analysis ability.

[0071] To address the problem of object scale variation in dense scenes, an MSDA module was added before the detection head of YOLOv8. This addition enables the model to focus more on important features during the feature integration process and reduces the influence of the background. In addition, the fourth detection head enhances the ability to extract information from the detected targets and optimizes the network's recognition ability.

[0072] Embodiment 2

[0073] This embodiment provides a dense scene grasping system based on machine vision, including:

[0074] A data acquisition module for acquiring target images and target geometric feature data;

[0075] A target recognition module for segmenting the target image based on an improved YOLOv8 model to obtain a target region, extracting the geometric center point coordinates of the target region and performing mean processing to obtain target center point coordinates; in the improved YOLOv8 model, a multi-scale attention module C2f_EMA is added after the second convolution in the C2f module in YOLOv8 to replace the original C2f module;

[0076] A positioning module for obtaining target pose data based on the target center point coordinates and target geometric feature data;

[0077] A grasping module for determining target grasping pose information based on the target pose data and controlling the robotic arm to perform a grasping operation according to the target pose data.

[0078] A dense scene grasping system based on machine vision provided by this embodiment, the recognition effect of its target recognition module in the application scenario is as Figures 9-10 shown.

[0079] Embodiment III

[0080] This embodiment provides a dense scene grasping device based on machine vision. The device includes a memory and a processor; the memory is used to store computer programs; the processor is used to implement the above-mentioned dense scene grasping method based on machine vision when executing the computer programs.

[0081] Wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in Embodiment I above.

[0082] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0083] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0084] In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.

[0085] The method in Embodiment I can be directly embodied as being executed by the hardware processor, or completed by a combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0086] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0087] The device further includes an object monitoring device for real-time monitoring of the position and state of objects in the working area, a robotic arm device for performing grasping actions, a power supply device, and a human-machine interaction interface for visualization.

[0088] In one embodiment, as shown in the attached Figure 7 and 8 The device includes an industrial control computer (1), a depth camera (2) for real-time monitoring of the position and state of objects (especially target objects to be grasped) in the working area, a robotic arm device (7) for performing planned grasping operations, a communication module (8) for connecting with other expansion modules, and a human-machine interaction module (10) for facilitating user monitoring and custom adjustment.

[0089] The depth camera (2) uses a high-resolution camera to capture images of the surrounding environment, and is connected to the industrial control computer (1) through a communication signal to transmit data.

[0090] The industrial control computer (1) stores a machine learning model, namely a monitoring model (4), and the monitoring model (4) adopts the improved YOLOv8 model provided in Embodiment 1. The improved YOLOv8 model has a highly parallel characteristic, allowing the entire image to be processed in a single forward pass. It realizes that in a lower inference time, the system provided in Embodiment 2 can quickly and effectively detect objects in an environment with high real-time requirements. The monitoring accuracy of the model is further improved through a training dataset (11). After the monitoring model finishes reasoning and making decisions on the current situation, it formulates an intelligent planning strategy and controls the robotic arm device (7) through a control unit (6).

[0091] In one embodiment, the training dataset (11) selects 17 object categories from the COCO dataset, with a total of 12,807 images. The instance ratio of each category is shown as Figure 11 Each image is selected according to the requirement of having more than three target objects. Finally, these are compiled into a dataset named D-COCO. In the experiment, the dataset is randomly divided into a training set and a test set, with a ratio of 8:2. The training set contains 10,246 images, and the test set contains 2,561 images. In addition, the image size is also adjusted to 640×640 to meet the input size.

[0092] After receiving the communication signal from the industrial control computer (1), the robotic arm device (7) immediately plans tasks for the located target. After reaching the target point, the gripper on the robotic arm will perform the grasping task.

[0093] The human-machine interaction (10) is used for debugging the system, manually adjusting various system parameters, or manual operation. The system parameters that can be manually adjusted include model selection, planning speed, delay events, grasping accuracy, and so on. It can also manually operate the robotic arm device for planned grasping. At the same time, it can also be used for debugging other expansion modules.

[0094] Embodiment Four

[0095] A computer-readable storage medium provided by another embodiment of the present invention stores a computer program. When the computer program is executed by a processor, the method for robotic arm grasping in a dense scene based on machine vision as described above is implemented.

[0096] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiment of the present invention. In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0097] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the protection scope of the present invention.

Claims

1. A dense scene grasping method based on machine vision, characterized in that, Including: S101: Obtain a target image and target geometric feature data; S102: Segment the target image based on the improved YOLOv8 model to obtain a target region, extract the geometric center point coordinates of the target region and perform mean processing to obtain target center point coordinates; In the improved YOLOv8 model, a multi-scale attention module C2f_EMA is added after the second convolution in the C2f module in YOLOv8 to replace the original C2f module; S103: Obtain target pose data based on the target center point coordinates and target geometric feature data; S104: Determine target grasping pose information based on the target pose data, and control the robotic arm to perform a grasping operation according to the target grasping pose information.

2. The method for grasping in a dense scene based on machine vision according to claim 1, wherein The multi-scale attention module C2f_EMA includes: a CBS module, a splitter, n Bottleneck modules, a first connector, a CBS module, and an EMA module connected in sequence; Among them, the output ends of the n Bottleneck modules are all connected to the input end of the first connector.

3. The machine vision-based dense scene grasping method according to claim 1, wherein In the improved YOLOv8 model, the original SPPF module is replaced by the SPPF_LSKA module.

4. The machine vision-based dense scene grasping method according to claim 3, wherein The SPPF_LSKA module includes: an input layer, a CBS module, a first max pooling layer, a second max pooling layer, a third max pooling layer, a second connector, an LSKA module, a CBS module, and an output layer connected in sequence; Among them, the output ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer are all connected to the input end of the second connector.

5. The method for grasping in a dense scene based on machine vision according to claim 1, wherein In the improved YOLOv8 model, an MSDA module is added before each of the four detection heads of the YOLOv8 model.

6. The method for grasping in a dense scene based on machine vision according to claim 5, wherein The MSDA module divides the channels of the feature map into multiple independent heads, and each head is assigned a specific task to perform self-attention operations at different dilation rates.

7. The machine vision-based dense scene grasping system according to claim 1, characterized in that Including: A data acquisition module for obtaining a target image and target geometric feature data; A target recognition module for segmenting the target image based on the improved YOLOv8 model to obtain a target region, extracting the geometric center point coordinates of the target region and performing mean processing to obtain target center point coordinates; In the improved YOLOv8 model, a multi-scale attention module C2f_EMA is added after the second convolution in the C2f module in YOLOv8 to replace the original C2f module; A positioning module for obtaining target pose data based on the target center point coordinates and target geometric feature data; A grasping module for determining target grasping pose information based on the target pose data and controlling the robotic arm to perform a grasping operation according to the target pose data.

8. A dense scene grasping device based on machine vision, characterized in that, The device includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the machine vision-based dense scene grasping method according to any one of claims 1 to 6 when executing the computer program.

9. The machine vision-based dense scene grasping device according to claim 8, wherein, The device further includes an object monitoring device for real-time monitoring of the position and state of objects in the working area, a robotic arm device for performing grasping actions, a power supply device, and a human-machine interface for visualization.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, the machine vision-based dense scene grasping method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Luggage case pull rod and RFID tag identification method based on deep network vision, terminal equipment and storage medium

    CN120823484A

  • Luggage handle and RFID tag recognition method based on deep network vision, terminal device and storage medium

    CN120823484B