DCEMA-YOLO clamp key point detection method
By introducing the DCEMA attention mechanism and the DCEMA-C3K2 module in the YOLO model, the problem of insufficient detection accuracy of fixture key points in the existing technology is solved, and higher detection accuracy and robustness are achieved, which is suitable for key points detection of special-shaped fixtures and complex scenarios.
Patent Information
- Application Number
- CN202510195744.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-23
AI Technical Summary
When dealing with special-shaped fixtures and complex scenarios, the key point detection accuracy is insufficient, and it is difficult to capture subtle feature changes. The loss function attaches importance to target box regression and ignores the key point positioning accuracy.
A DCEMA-YOLO fixture key point detection method is proposed. By introducing an improved DCEMA attention mechanism and DCEMA-C3K2 module, the feature extraction and feature fusion capabilities are enhanced, and the key point deviation loss calculation is strengthened during the detection stage.
It improves the accuracy and robustness of key points detection of fixtures, can better capture subtle feature changes in key points of fixtures, enhances the ability to identify local key points features, and improves detection accuracy and application quality.
Smart Images

Figure CN120032140A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of target detection and relates to a DCEMA-YOLO fixture key point detection method. Background Art
[0002] The present invention is aimed at the key point detection technology of the fixture based on machine vision, which aims to solve the problem of insufficient accuracy of the existing technology when dealing with special-shaped fixtures and complex scenes, thereby improving the quality and efficiency of applications such as fixture positioning, grasping and posture detection.
[0003] Challenges faced by existing technologies:
[0004] There are many types of fixtures with different shapes, which brings challenges to key point detection.
[0005] The fixture often works in a complex environment, and factors such as lighting conditions and background interference will affect the detection accuracy.
[0006] Existing deep learning methods, such as the YOLO series of models, perform well in object detection, but the accuracy of key point detection still needs to be improved.
[0007] The existing models have limited feature extraction capabilities for the key points of the fixture and are unable to capture subtle feature changes.
[0008] In the process of feature fusion, existing models tend to ignore local key point features, resulting in decreased detection accuracy.
[0009] The loss functions of existing models often only focus on the regression of the target box and pay insufficient attention to the positioning accuracy of key points.
[0010] In view of the above problems, the present invention proposes a fixture key point detection method based on DCEMA-YOLO, aiming to improve the detection accuracy and robustness. Summary of the invention
[0011] In view of this, the object of the present invention is to provide a DCEMA-YOLO fixture key point detection method to improve the fixture key point detection performance based on machine vision.
[0012] In order to achieve the above object, the present invention provides the following technical solutions:
[0013] A DCEMA-YOLO fixture key point detection method comprises the following steps:
[0014] Step 1: Collect image data of fixtures in different scenes, annotate them, and generate training and test sets containing fixture bounding boxes and key point information;
[0015] Step 2: Build a network model containing the improved DCEMA attention mechanism, where:
[0016] The DCEMA attention mechanism module includes a dual C3 convolution structure for enhancing feature extraction capabilities;
[0017] The network model introduces a DCEMA attention mechanism module in the feature extraction stage to extract the features of the key points of the fixture;
[0018] The network model introduces the DCEMA-C3K2 module in the neck stage to fuse feature maps at different levels and strengthen local key point features;
[0019] The network model strengthens the calculation of key point deviation loss in the detection phase to improve the key point positioning accuracy.
[0020] Furthermore, in the step 1, data enhancement operations of random scaling and rotation are performed on the training set images, and the fixture bounding box and key point information are updated synchronously.
[0021] Furthermore, the DCEMA attention mechanism module includes:
[0022] The first channel is used to extract channel features;
[0023] The second channel contains two 3×3 convolutional layers to extract more abstract and essential features;
[0024] The matrix multiplication unit is used to fuse the outputs of the first channel and the second channel.
[0025] Furthermore, the network model is a YOLOv11 network model.
[0026] Further, the DCEMA-C3K2 module comprises:
[0027] DCEMA attention mechanism module;
[0028] C3K2 module, used to fuse feature maps at different levels.
[0029] Furthermore, the key point deviation loss calculation formula is:
[0030]
[0031] Among them, d is the Euclidean distance between the predicted key point and the real key point, Sigmas is the smoothing parameter, and area is the area where the key point is located. This makes the error proportional to the size of the object. The influence of d is strengthened in the exponential calculation to improve the accuracy of fixture key point detection.
[0032] A fixture key point detection system based on the fixture key point detection method includes:
[0033] Image acquisition module, used to collect image data of fixtures in different scenes;
[0034] The image annotation module is used to annotate the collected images and generate training and test sets containing the fixture bounding box and key point information;
[0035] The network model has the same structure and function as the aforementioned network model.
[0036] A computer-readable storage medium stores a computer program, which implements the fixture key point detection method when executed by a processor.
[0037] A computer device comprises a processor and a memory, wherein the computer program is stored in the memory.
[0038] A YOLOv11 network model for fixture key point detection, the structure of the network model is the same as the network model.
[0039] The beneficial effects of the present invention are:
[0040] The present invention provides a method for detecting key points of fixtures based on DCEMA-YOLO, which effectively solves the problem of insufficient accuracy of the prior art when processing special-shaped fixtures and complex scenes, and achieves higher detection accuracy and robustness.
[0041] By introducing the DCEMA attention mechanism, the features of the key points of the fixture can be better captured, especially some subtle feature changes, thereby improving the detection accuracy.
[0042] Through the DCEMA-C3K2 module, feature maps at different levels can be effectively integrated, global features can be weakened, and local key point features can be strengthened, thereby improving the positioning accuracy of key points.
[0043] By strengthening the calculation of key point deviation loss, the model's positioning accuracy of key points can be more accurately evaluated, and the model can be guided for optimization, thereby further improving detection accuracy.
[0044] The present invention can effectively cope with the influence of factors such as illumination changes and background interference, and can maintain high detection accuracy in different scenarios.
[0045] The technical effect of the present invention makes the DCEMA-YOLO model more suitable for the task of fixture key point detection, and provides reliable technical support for the intelligent application of fixtures.
[0046] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0048] Figure 1 is a flow chart of the present invention;
[0049] Figure 2 This is a schematic diagram of the structure of the EMA improved DCEMA module of the present invention;
[0050] Figure 3 Schematic diagram of the C3K2 improved DCEMA-C3K2 module invented in this paper;
[0051] Figure 4 It is a schematic diagram of the network model structure of DCEMA-YOLO of the present invention;
[0052] Figure 5 This is the schematic diagram of the neck stage;
[0053] Figure 6 For the complete improved DCEMA-YOLO model. DETAILED DESCRIPTION
[0054] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0055] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0056] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0057] See also Figure 1 to Figure 4 The present invention provides a DCEMA-YOLO fixture key point detection method, which is mainly divided into five parts. The first part is to collect and preprocess the fixture data set, divide it into a training set and a test set, and expand the training image; the second part is to establish an improved DCEMA attention mechanism to solve the key point feature extraction problem; the third part is to set DCEMA in the backbone network to realize the extraction of key point features in the feature extraction stage; the fourth part is to set the DCEMA-C3K2 module in the neck stage to weaken the global features and strengthen the local key point features; the fifth part is to strengthen the key point deviation loss calculation in the detection part to improve the key point detection accuracy.
[0058] The flowchart of a DCEMA-YOLO fixture key point detection method provided by the present invention is as follows: Figure 1 As shown, the specific steps include: S1: collect the fixture data set and preprocess it, divide it into training set and test set, and expand the training images; S2: establish an improved DCEMA attention mechanism to solve the problem of key point feature extraction; S3: in the feature extraction stage, set DCEMA in the backbone network to realize the extraction of key point features; S4: set the DCEMA-C3K2 module in the neck stage to weaken the global features and strengthen the local key point features; S5: in the detection part, strengthen the key point deviation loss calculation to improve the key point detection accuracy.
[0059] Specifically, in an embodiment of the present invention, S1 collects fixture data sets and preprocesses them, collects fixture images under different backgrounds and lighting, and then screens and names the collected image data. Labelme software is used to annotate the fixture images, frame the fixture target area, and use the pose task annotation method in YOLO11 posture detection to annotate key points, and convert the annotated images and jeson format annotation files into corresponding yolo format txt tags, and divide the converted yolo format fixture data sets into training sets, verification sets, and test sets. Randomly scale and rotate the fixture training set images, and correct the bounding box and key point coordinates, and expand the training images.
[0060] like Figure 2 As shown in the figure, in order to solve the problem of fixture key point feature extraction, the EMA attention mechanism is improved, the DCEMA attention mechanism module is constructed, Conv3*3 is added to channel 2, and a secondary convolution operation with a convolution kernel of 3×3 is performed on the channel 2 feature. The output size remains unchanged from the input size, and the output size is calculated as shown in the formula.
[0061]
[0062] Hin and Win are the height and width of the input, Hout and Wout are the height and width of the output feature map, kernel_size and stride correspond to the convolution kernel size and step size respectively, and padding is used to indicate the number of rows and columns filled.
[0063] Through two Conv3*3, convolution kernel kernel_size=3, stride=1, padding=1, the depth of the network is improved under the condition of the same receptive field, thereby enhancing more abstract and essential features.
[0064] In this example, the double convolution is established, and the output feature size remains unchanged, Hout = Hin, Wout = Win.
[0065] After the double convolution of channel 2, pooling is performed all the way to reduce the number of network parameters and reduce the computational complexity. On the other hand, it retains key feature information while avoiding problems such as overfitting.
[0066] After the double convolution of channel 2, another path performs matrix multiplication with the output of channel 1 to preserve the cross-dimensional interactions of EMA and the dependencies between different dimensions.
[0067] like Figure 3 As shown in the figure, in the backbone stage, the DCEMA attention mechanism is introduced in the 11th layer. The channel information and spatial information are combined through the DCEMA module, and the feature extraction of the key points of the fixture is improved through the multi-scale dual-channel structure.
[0068] like Figure 4 As shown, DCEMA is combined with C3K2, C3K2 is improved, and the DCEMA-C3K2 module is constructed. Due to the introduction of DCEMA, the model further extracts the key point features of the fixture.
[0069] like Figure 5As shown in the figure, the C3K2 in the neck stage is removed and the DCEMA-C3K2 module is inserted to improve the fusion of feature maps at different levels and the accuracy of fixture key point detection. The DCEMA-C3K2 insertion points are 17, 20, and 23 layers, and the corresponding three layers are output to the head stage respectively.
[0070] like Figure 6 The figure shows the complete improved DCEMA-YOLO model. In the detection part, the key point deviation loss calculation is strengthened to improve the key point detection accuracy.
[0071] The keypoint loss is normalized and calculated as the exponent of the keypoint loss:
[0072]
[0073] d is the Euclidean distance between the predicted key point and the real key point, Sigmas is the smoothing parameter, and area is the area where the key point is located. The error is made proportional to the size of the object. The influence of d is strengthened in the exponential calculation to improve the accuracy of fixture key point detection.
[0074] The fixture key point detection method of DCEMA-YOLO designed by the present invention mainly includes two stages: training and testing. In the training stage, feature extraction is performed and model weight parameters are obtained, various loss values are calculated, and the weight parameters of the model are further updated to finally obtain the optimal model weight. In the testing stage, the optimal model weight needs to be loaded, the fixture images of the fixture test set are tested, and the information of the bounding box and key points of the fixture is directly output.
[0075] This embodiment aims to demonstrate the application process of the fixture key point detection method of the present invention through specific numerical parameters and verify its effectiveness.
[0076] 1. Data preparation
[0077] Data collection: Collect 1,000 images of fixtures in different scenarios, such as factory workshops, laboratories, and other environments, ensuring that the images contain a variety of fixture types and complex backgrounds.
[0078] Data annotation: Use labelme software to annotate the collected images, including the bounding box and key point information of the fixture. In this example, the fixture type is a U-shaped fixture, and the key points include two endpoints and the center point, a total of 3 key points.
[0079] Data division: The labeled data set is divided into a training set of 800 images, a validation set of 100 images, and a test set of 100 images.
[0080] Data enhancement: Perform random scaling (0.8-1.2 times), rotation (-10°-10°) and other data enhancement operations on the training set images, and simultaneously update the fixture bounding box and key point information.
[0081] 2. Network Model
[0082] Model structure: This embodiment adopts the YOLOv11 network model as the basic model, and improves it by introducing the DCEMA attention mechanism and DCEMA-C3K2 module.
[0083] DCEMA attention mechanism: In the backbone network of YOLOv11, the DCEMA attention mechanism module is introduced in the 11th layer to extract the features of the key points of the fixture.
[0084] DCEMA-C3K2 module: In the neck stage of YOLOv11, the C3K2 module is replaced by the DCEMA-C3K2 module to fuse feature maps at different levels and strengthen local key point features.
[0085] Parameter settings: learning rate is set to 0.001, batch size is set to 16, and number of training rounds is set to 100.
[0086] 3. Model training
[0087] Loss function: A multi-task loss function consisting of target box regression loss, confidence loss, and keypoint deviation loss is used for training.
[0088] Key point deviation loss: The calculation formula described in claim 6 is used to strengthen the key point deviation loss calculation and improve the key point positioning accuracy. The smoothing parameter Sigmas is set to 0.05.
[0089] Training process: Use the training set to train the model and use the validation set to evaluate it. Adjust the model parameters based on the evaluation results until satisfactory performance is achieved. In this example, after 100 rounds of training, the average key point positioning error of the model on the test set is 2.1 pixels, which meets the requirements of applications such as fixture positioning, grasping, and posture detection.
[0090] 4. Application scenarios
[0091] Fixture positioning: Match the detected fixture key points with the pre-set fixture model to achieve accurate fixture positioning. In this example, the model can accurately identify the three key points of the U-shaped fixture and match them with the preset model, achieving a fixture positioning accuracy of 98%.
[0092] Fixture grabbing: Based on the detected fixture key point information, the robot arm is controlled to perform the fixture grabbing operation. In this example, the model can identify the center point of the U-shaped fixture and use it as the target point for the robot arm to grab the fixture accurately.
[0093] Fixture posture detection: Analyze the relationship between the key points of the fixture to determine the fixture posture state. In this example, the model can determine the opening direction and rotation angle of the fixture based on the relative positions of the two endpoints of the U-shaped fixture to achieve fixture posture detection.
[0094] This embodiment demonstrates the application process of the fixture key point detection method of the present invention through specific numerical parameters, and verifies its effectiveness in application scenarios such as fixture positioning, grasping and posture detection.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A DCEMA-YOLO fixture key point detection method, characterized in that: The following steps are involved: Step 1: Collect image data of fixtures in different scenes, annotate them, and generate training and test sets containing fixture bounding boxes and key point information; Step 2: Build a network model containing the improved DCEMA attention mechanism, where: The DCEMA attention mechanism module includes a dual C3 convolution structure for enhancing feature extraction capabilities; The network model introduces a DCEMA attention mechanism module in the feature extraction stage to extract the features of the key points of the fixture; The network model introduces the DCEMA-C3K2 module in the neck stage to fuse feature maps at different levels and strengthen local key point features; The network model strengthens the calculation of key point deviation loss in the detection phase to improve the key point positioning accuracy.
2. The DCEMA-YOLO fixture key point detection method according to claim 1, characterized in that: In the step 1, data augmentation operations of random scaling and rotation are performed on the training set images, and the fixture bounding box and key point information are updated synchronously.
3. The fixture key point detection method of DCEMA-YOLO according to claim 1, characterized in that: The DCEMA attention mechanism module includes: The first channel is used to extract channel features; The second channel contains two 3×3 convolutional layers to extract more abstract and essential features; The matrix multiplication unit is used to fuse the outputs of the first channel and the second channel.
4. The DCEMA-YOLO fixture key point detection method according to claim 1, characterized in that: The network model is a YOLOv11 network model.
5. The DCEMA-YOLO fixture key point detection method according to claim 1, characterized in that: The DCEMA-C3K2 module includes: DCEMA attention mechanism module; C3K2 module, used to fuse feature maps at different levels.
6. The DCEMA-YOLO fixture key point detection method according to claim 1, characterized in that: The key point deviation loss calculation formula is: Among them, d is the Euclidean distance between the predicted key point and the real key point, Sigmas is the smoothing parameter, and area is the area where the key point is located. This makes the error proportional to the size of the object. The influence of d is strengthened in the exponential calculation to improve the accuracy of fixture key point detection.
7. A fixture key point detection system based on the fixture key point detection method according to claim 1, characterized in that: include: Image acquisition module, used to collect image data of fixtures in different scenes; The image annotation module is used to annotate the collected images and generate training and test sets containing the fixture bounding box and key point information; The network model has the same structure and function as the aforementioned network model.
8. A computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method for detecting key points of a fixture as claimed in claim 1.
9. A computer device, characterized in that: The device comprises a processor and a memory, wherein the computer program according to claim 8 is stored in the memory.
10. A YOLOv11 network model for fixture key point detection, characterized in that: The structure of the network model is the same as the network model described in claim 1.
Citation Information
Cited By
Water surface garbage identification method and system based on unmanned aerial vehicle
CN120708103A
An unmanned aerial vehicle-based water surface garbage identification method and system
CN120708103B