Method and system for correcting position of metal diaphragm piece based on YOLOv8 improvement

By preprocessing the metal diaphragm position adjustment, using the feature-level super-resolution fusion FSR mechanism, ECA attention mechanism, and IoU-Guided NMS mechanism, the problems of low manual efficiency and poor robustness of traditional algorithms in metal diaphragm position adjustment are solved, achieving high-precision and high-real-time automated positioning and correction.

CN121883789APending Publication Date: 2026-04-17NANJING YUNTONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the position adjustment of metal diaphragm parts relies on manual operation, which is inefficient and has unstable accuracy. Traditional machine vision algorithms have poor robustness in scenarios with surface reflection, lighting changes, and background interference on metal diaphragm parts. The YOLOv8 model has insufficient positioning accuracy for small key points and lacks an integrated solution for model inference and robotic arm execution, resulting in large positioning errors and inability to adapt to high-speed automated production lines.

Method used

A dataset is constructed using preprocessing techniques including contrast enhancement, illumination intensity adjustment, and geometric transformation enhancement. Feature-level super-resolution fusion (FSR) and ECA attention mechanisms are introduced into YOLOv8, and a hybrid loss function combining Focal Loss and gradient equalization (GHM) is designed to replace the traditional YOLOv8 loss function. An IoU-Guided NMS mechanism is introduced to optimize the selection of keypoint candidate boxes, and a lightweight inference engine is encapsulated to drive the robotic arm adjustment.

Benefits of technology

It significantly improves the model's adaptability to changes in the production environment, increases the accuracy of key point positioning and angle calculation, reduces positioning errors, achieves fully automatic position correction, and is compatible with high-speed automated production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883789A_ABST
    Figure CN121883789A_ABST
Patent Text Reader

Abstract

The invention relates to a metal diaphragm piece position correction method and system based on YOLOv8 improvement, and the method comprises the steps: designing a conventional YOLOv8 target detection model structure which is improved by fusing feature-level super-resolution with an FSR mechanism, and is specially used for recognizing feature key points in a metal diaphragm piece image so as to determine a placement posture; a loss function of a traditional YOLOv8 is improved based on a gradient equalization mechanism GHM, contribution of difficult-to-separate samples and easy-to-separate samples to model training is further balanced, and extreme sample interference is reduced; an IoU-Guidded NMS mechanism is designed to optimize screening of key point candidate frames, and the redundancy suppression effect of key point detection results is optimized; various monitoring indexes are designed to train a metal diaphragm piece position correction network, lightweight packaging and deployment are carried out, mechanical arm linkage control is carried out, and metal diaphragm piece position correction is achieved. The method can be widely applied to the fields of electronic manufacturing, automatic assembly and the like, and the problems of low manual adjustment efficiency and large positioning deviation of the metal diaphragm piece are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and deep learning technology, and in particular to a method and system for correcting the position of metal diaphragms based on YOLOv8. Background Technology

[0002] Currently, the adjustment of metal diaphragm components is still mainly done manually, relying on workers to visually observe and judge the posture of the metal diaphragm components, and manually push the fixture or use simple tools to adjust the position. This method has obvious drawbacks: First, it is inefficient, with each metal diaphragm component requiring a lot of time to adjust, making it difficult to adapt to high-speed automated production lines; second, the accuracy is unstable, affected by subjective factors such as worker experience and visual fatigue, and the angle adjustment error and position deviation can easily exceed production requirements, failing to meet the needs of high-precision assembly; third, labor costs are high, requiring a large number of quality inspection and adjustment personnel in large-scale production scenarios, increasing the company's operating costs.

[0003] Although some companies have attempted to introduce traditional machine vision technology for the localization of metal diaphragm parts, traditional algorithms (such as template matching and edge detection) are prone to misjudgment or omission when locating key points due to factors such as surface reflection of the metal diaphragm parts, changes in ambient lighting, and interference from background clutter. For example, when there are slight scratches or stains on the surface of the metal diaphragm parts, the template matching algorithm may mistake the defect for a key point, leading to an error in angle calculation; while changes in lighting intensity will reduce the feature extraction capability of the edge detection algorithm, further affecting the localization accuracy.

[0004] In recent years, deep learning object detection technology based on the YOLO series has shown potential in industrial positioning scenarios. YOLOv8, as the current mainstream model, has the advantages of being lightweight and having high inference speed. However, the default network structure of YOLOv8 is designed for general targets and lacks sufficient positioning accuracy for targets with "small key points and low contrast," such as metal diaphragms. At the same time, existing technologies lack an integrated solution for "model inference - robotic arm execution," meaning that the model output cannot directly drive hardware devices to complete position correction, resulting in a gap in the technology's practical application.

[0005] In summary, existing metal diaphragm component position correction technologies have the following problems: Manual adjustments are inefficient and their accuracy is greatly affected by subjective factors, making them unsuitable for high-speed automated production lines. Traditional machine vision algorithms have poor robustness in scenarios involving surface reflection, lighting changes, and background interference on metal films, and their key point localization accuracy is insufficient. The localization scheme based on the general YOLOv8 model has weak recognition ability for small key points of metal film parts, and the traditional "upsampling + convolution" fusion method of the Neck layer leads to fuzzy semantic features of P3 and low matching degree with P2 detailed features, which further amplifies the angle calculation error. The lack of an integrated solution for "model reasoning and robotic arm execution" makes it impossible to achieve fully automated position correction. Traditional NMS (Non-maximum suppression mechanism) has problems such as "missed detection due to fixed threshold" or "inability to remove redundant boxes" in the detection of key points of metal diaphragms: When the density of key points of metal diaphragms is high (such as multiple adjacent positioning holes), the fixed threshold of traditional NMS is prone to misjudging adjacent small key points as redundant boxes and suppressing them (missed detection); when there are slight stains on the surface of metal diaphragms, traditional NMS has difficulty removing interference boxes with low IoU (redundancy), resulting in deviation in angle calculation.

[0006] Therefore, there is an urgent need to propose a position correction system that integrates and improves the YOLOv8 model, adapts to the key point positioning requirements of metal diaphragm components, and has the full-process capability of "data processing-model training-inference control-robotic arm linkage" to solve the problems of low efficiency of manual adjustment, poor robustness of traditional algorithms, and difficulty in technology implementation, so as to meet the high precision and high real-time requirements of industrial automated production. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a method and system for correcting the position of metal diaphragms based on an improved version of YOLOv8. This method alleviates the problem of misjudgment or omission when locating key points in traditional metal diaphragm position correction methods. It also solves the problems of insufficient adaptability of the default network structure of YOLOv8 to the task of correcting the position of metal diaphragms and the lack of an integrated solution in existing technologies.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for correcting the position of a metal diaphragm based on YOLOv8, comprising the following steps: S1. Obtain the original images of the metal diaphragm at different placement angles, and perform preprocessing operations such as contrast enhancement, illumination intensity adjustment, and geometric transformation enhancement to label the real key points and their real category labels, and construct a dataset. S2. Based on the lightweight ESPCN design, a feature-level super-resolution fusion FSR mechanism is added to the YOLOv8 structure and an ECA attention mechanism is introduced to optimize feature extraction and localization accuracy. A metal film component position correction network is constructed, and a special monitoring index for the feature-level super-resolution fusion FSR mechanism is designed. S3. Design a hybrid loss that combines Focal Loss and gradient equalization mechanism GHM to replace the traditional YOLOv8 classification loss. At the same time, introduce the gradient equalization mechanism GHM into the traditional YOLOv8 bounding box regression loss to obtain the GHM-Reg loss. Design a special monitoring index for the gradient equalization mechanism GHM. S4. Design an IoU-Guided NMS mechanism to replace the traditional YOLOv8 nonmaximum suppression mechanism in the inference stage of the metal film device position correction network, optimize the screening of key point candidate boxes, obtain a list of screened key point boxes to guide the generation of predicted key points, and design specific monitoring indicators for the IoU-Guided NMS mechanism. S5. Using key point localization accuracy and angle calculation error as core monitoring indicators, and combining specific monitoring indicators for feature-level super-resolution fusion FSR mechanism, gradient equalization mechanism GHM and IoU-Guided NMS mechanism, the dataset constructed in step S1 is divided into training set, validation set and test set, and a metal diaphragm position correction network is trained; the metal diaphragm position correction network outputs predicted key points and calculates the predicted deflection angle of the metal diaphragm based on the coordinates of the predicted key points. S6. Encapsulate and deploy the trained metal diaphragm position correction network, design an acceleration optimization strategy for the feature-level super-resolution fusion FSR mechanism, acquire metal diaphragm images and use the deployed metal diaphragm position correction network to predict key points and the predicted deflection angle of the metal diaphragm. Before the robotic arm linkage control, the IoU-Guided NMS mechanism is used again to screen candidate boxes for key points. Based on the acquired predicted key point coordinates and predicted deflection angle of the metal diaphragm, the robotic arm linkage control is performed, and the position correction effect of the metal diaphragm is verified based on the core monitoring indicators until the verification result meets the standard and the position correction of the metal diaphragm is completed.

[0009] Furthermore, in step S1, the specific process includes the following steps: S11. Use an industrial camera to capture images of the original metal film parts under different lighting conditions. The shooting process covers multiple placement angles to ensure that the dataset covers the posture deviations that may occur in actual production. S12. Perform contrast enhancement processing on the original metal diaphragm image by linearly stretching the grayscale range to highlight the feature areas of the metal diaphragm, which facilitates the identification of key points. S13. By randomly adjusting the light intensity, the brightness channel of the metal film image after contrast enhancement is adjusted to simulate the light changes during different production periods. S14. Apply horizontal flipping and random rotation geometric transformations to the images of metal film components after adjusting a certain proportion of the illumination intensity to complete the preprocessing operation and expand the dataset size. S15. For each preprocessed metal diaphragm image, label the centers of three non-collinear key positioning holes as real key points, record the coordinates of the real key points and their corresponding real category labels, and together with the preprocessed metal diaphragm images, form a dataset.

[0010] Furthermore, in step S2, the specific process includes the following steps: S21. A feature-level super-resolution fusion (FSR) mechanism is added to the Neck feature fusion layer of the YOLOv8 structure. The feature-level super-resolution fusion (FSR) mechanism is based on lightweight ESPCN for super-resolution reconstruction, extracts P3-FSR features, and then performs dynamic weighted fusion and smoothing processing to obtain integrated features. S22. An ECA attention mechanism is introduced after each C2f module of the YOLOv8 structure to complete the construction of the metal diaphragm position correction network. S23. The design includes specific monitoring indicators for the feature-level super-resolution fusion FSR mechanism, calculating the structural similarity between the P3-FSR features and the P2 features extracted by the metal diaphragm position correction network in the key point region. ,Require Simultaneously, the changes in localization accuracy of small keypoints before and after the intervention of the feature-level super-resolution fusion (FSR) mechanism are statistically analyzed, with an accuracy improvement of ≥3%. If the target is not met, the number of convolution kernels or the magnification factor of the sub-pixel convolutional layer in the feature-level super-resolution fusion (FSR) mechanism are adjusted to ensure that the feature-level super-resolution fusion (FSR) mechanism can play an effective role.

[0011] Furthermore, in step S21, the specific process includes the following steps: S211. Extract P3 features based on the feature extraction layer of lightweight ESPCN; S212. Based on the lightweight ESPCN subpixel convolutional layer, the resolution difference between P2 and P3 features is matched, and the resolution of P3 features is improved to be consistent with that of P2 features to obtain P3-FSR features. S213. The number of channels of the P3-FSR feature is unified with that of the P2 feature through a convolutional layer, and then the following dynamic weighted fusion is performed: ; in , These represent the P2 features in the dynamic weighted fusion process. and P3-FSR characteristics The weight, For fusion features; after dynamic weighted fusion, the fusion features The image is then smoothed again by a convolutional layer to obtain integrated features.

[0012] Furthermore, in step S3, the specific process includes the following steps: S31. Introduce dynamic weight coefficients based on Focal loss. ,in Indicates sample The true category label, Follow The indicated category is dynamically adjusted; S32. Based on the gradient equalization mechanism (GHM), the classification gradient magnitude of each training sample is statistically analyzed, and a classification gradient density histogram is constructed. The classification gradient density histogram divides the classification gradient magnitude into 10-15 gradient intervals, and the classification gradient density weight of each gradient interval is calculated. ; In the formula, Indicates the first The classification gradient density weights for each gradient interval. The total number of samples, This represents the number of gradient intervals. For the first The number of samples in each gradient interval; S33. Construct a hybrid loss by combining dynamic weight coefficients and classification gradient density weights. as follows: ; In the formula, For the first A set of samples in gradient intervals For the sample The predicted probability, FocalLoss is the easy / difficult sample adjustment coefficient. S34. Using Smooth L1 loss as the basic bounding box regression loss, calculate the regression gradient density weights for the regression gradient magnitude of each training sample in the same manner as in step S32. Construct GHM-Reg loss as follows: ; In the formula, Indicates Smooth L1 loss, For the sample The true bounding box parameters, For the sample Predicted bounding box parameters; S35. The design includes specific monitoring metrics for the gradient equalization mechanism GHM. After each training round, the distribution of classification gradient magnitude and regression gradient magnitude of samples in the training and validation sets is statistically analyzed to ensure that the gradients are concentrated within a preset effective interval. If the gradients exceed this effective interval, the corresponding gradient interval division or loss function parameters are adjusted. Simultaneously, the classification gradient density weights of each gradient interval are monitored. and regression gradient density weights The fluctuation range is required to be ≤15% for three consecutive rounds to ensure the training stability of the metal diaphragm position correction network.

[0013] Furthermore, in step S4, the specific process includes the following steps: S41. Sort the candidate keypoint boxes of the same category from highest to lowest confidence level, and select the box with the highest confidence level as the baseline box. ; S42, Calculate the baseline frame Compared with other keypoint candidate boxes The intersection and union ratio; S43. Introducing distance weighting coefficient Define the guidance inhibition threshold : ; in The basic suppression threshold is set, with a value ranging from 0.3 to 0.5. This indicates the intersection-union ratio (IU) calculation operation; distance weighting coefficient. Candidate boxes for each key point With reference frame center distance Average size of key points The ratio is determined by the formula: ; in The attenuation coefficient ranges from 1.2 to 1.5 and is dynamically adjusted based on the density of key points in the metal diaphragm. The higher the density of key points, the larger the attenuation coefficient. It is the base of the natural logarithm; S44. If key point candidate boxes With reference frame The intersection-union ratio ≥ Then inhibit If the intersection-union ratio is < Then keep As a new baseline box, repeat the above process until all candidate boxes for key points have been processed and the filtering is complete. S45. The design includes specific monitoring indicators for the IoU-Guided NMS mechanism. The proportion of duplicate points in the predicted keypoints generated from the test set to the total predicted keypoints is required to have a duplication rate ≤ 1%. Simultaneously, the proportion of undetected real keypoints in the test set to the total real keypoints is required to have a false negative rate ≤ 0.5%. If the above indicators are not met, return to step S43 to adjust the attenuation coefficient. or baseline inhibition threshold Then, the inference and verification of the metal diaphragm position correction network were repeated until the target was met.

[0014] Furthermore, in step S5, the process of using key point positioning accuracy and angle calculation error as core monitoring indicators includes the following steps: S51. The metal diaphragm piece position correction network detects and outputs the coordinates of three non-collinear prediction key points. , , ,by Construct a vector for the origin ,calculate The perpendicular vector is used as the normal vector. Calculate the angle between the normal vector and the horizontal reference direction. This is the prediction of the deflection angle of the metal diaphragm component; S52. The accuracy of key point localization is calculated by determining the Euclidean distance between the predicted key point coordinates and the actual key point coordinates. As a key point positioning deviation, when When the pixel count is used, it is considered accurate. S53, the aforementioned angle calculation error is for predicting the deflection angle of the metal diaphragm. The difference between the actual deflection angle of the metal diaphragm and the actual deflection angle of the metal diaphragm. S54. Stop training when the validation set metrics do not improve for several consecutive rounds.

[0015] Furthermore, in step S6, the specific process includes the following steps: S61. The trained metal diaphragm position correction network is packaged into a lightweight inference engine and deployed to an edge computing device; S62. For the feature-level super-resolution fusion FSR mechanism, inference acceleration optimization is performed. TensorRT is used to fuse the feature extraction layer and sub-pixel convolution layer of the lightweight ESPCN into a single custom operator to reduce memory access overhead. At the same time, the weight precision of the feature-level super-resolution fusion FSR mechanism is quantized from FP32 to FP16. While ensuring the quality of feature reconstruction, the inference speed of the metal film piece position correction network meets the real-time detection requirements. S63. The edge computing device acquires images of the metal diaphragm in real time using an industrial camera, calculates and outputs a predicted deflection angle of the metal diaphragm using a metal diaphragm position correction network. ; S64. Robotic arm linkage control: Edge computing device deflects the metal diaphragm component by an angle. Deviation from key point positioning The data is sent to the robotic arm controller, which then drives the robotic arm to... To correct the angle, To compensate for positional deviations, rotation and translation operations are performed; S65. Perform effect verification. After the robotic arm is adjusted, re-acquire the image of the metal diaphragm. Verify that the key point positioning accuracy is ≥98% and the angle calculation error is ≤0.5°. If the verification result is satisfactory, return to step S63 and repeat steps S63-S65 with the metal diaphragm after the robotic arm is adjusted until the verification result is satisfactory, thus completing the position correction of the metal diaphragm.

[0016] Furthermore, before the robotic arm performs the position correction of the metal diaphragm component, the IoU-Guided NMS mechanism is applied again to receive the filtered list of keypoint boxes sent by the edge device. The list contains verification flags indicating that the predicted keypoints are free of duplicates and omissions. If duplicates or omissions are found in the list, an abnormal verification flag is generated, and the robotic arm controller triggers a pause command. At the same time, the edge computing device automatically readjusts the attenuation coefficient. or baseline inhibition threshold Continue until a list of qualified keypoint boxes is output, and then determine the keypoints.

[0017] Furthermore, the present invention also provides a metal diaphragm position correction system based on YOLOv8, for applying the aforementioned metal diaphragm position correction method based on YOLOv8, including: an image acquisition and preprocessing module, a model optimization module, and an inference and control module; The image acquisition and preprocessing module is used to acquire original images of metal film components at different placement angles and to perform preprocessing operations such as contrast enhancement, light intensity adjustment, and geometric transformation enhancement. The model optimization module is used to execute the feature-level super-resolution fusion FSR mechanism, the ECA attention mechanism, the calculation of the hybrid loss and GHM-Reg loss combined with the gradient equalization mechanism GHM, and the parameter optimization of the IoU-Guided NMS mechanism. The model optimization module has a built-in FSR parameter configuration submodule, which dynamically adjusts the magnification factor and convolution kernel parameters of the super-resolution reconstruction according to the key point size of the metal film in the feature-level super-resolution fusion FSR mechanism, so as to adapt to the feature fusion requirements of metal film with different specifications. The inference and control module is used to deploy the trained metal diaphragm position correction network to the edge computing device, and output the coordinates of predicted key points and the predicted deflection angle of the metal diaphragm in real time, driving the robotic arm to perform the position correction operation of the metal diaphragm. The inference and control module includes an FSR inference acceleration unit, which ensures the efficient operation of the feature-level super-resolution fusion FSR mechanism on the edge computing device through operator fusion and precision quantization.

[0018] By employing the above technical solution, the present invention provides a method and system for correcting the position of metal diaphragm components based on YOLOv8, which has at least the following beneficial effects: (1) This invention significantly improves the model’s adaptability to changes in the production environment by employing a multi-dimensional data enhancement strategy that includes contrast enhancement, light intensity adjustment, and geometric transformation enhancement, thereby solving the positioning deviation problem caused by surface reflection and light interference of metal film parts. (2) This invention enhances YOLOv8’s ability to identify small key points of metal film components by adding a feature-level super-resolution fusion FSR mechanism and an ECA attention mechanism, thereby improving the accuracy of key point positioning and reducing angle calculation error; (3) The present invention improves the loss function of traditional YOLOv8 based on the gradient equalization mechanism GHM. The constructed hybrid loss balances the ratio of positive and negative samples through Focal Loss on the one hand, and suppresses the overtraining of extremely easy-to-distinguish samples (such as pure background) and the gradient interference of extremely difficult-to-distinguish samples (such as stain misjudgment) through GHM on the other hand, so that the model training focuses more on effective metal film feature samples, further improving the accuracy of key point localization and reducing the angle calculation error. (4) This invention solves the problems of missed detection and redundancy in key point detection of metal film components by introducing the IoU-Guided NMS mechanism. By designing a dynamic suppression threshold to adapt to the difference in key point density, the repetition rate and the missed detection rate are reduced. At the same time, the processing time on the edge device is reduced, which does not affect the real-time performance of model inference and ensures the efficiency of robotic arm adjustment. (5) This invention improves the localization accuracy of small key points by lightweight design and inference optimization of feature-level super-resolution fusion FSR mechanism. At the same time, it controls the inference time of feature-level super-resolution fusion FSR mechanism to within 2ms by TensorRT quantization and operator fusion, ensuring that the real-time performance of the overall model meets the needs of industrial scenarios. Compared with the solution without this mechanism, the localization accuracy of small key points is improved. (6) By constructing an integrated process of “model reasoning-robotic arm control”, the model output directly drives the hardware to perform adjustments without human intervention, shortening the time required to correct the position of the metal diaphragm and adapting it to high-speed automated production lines. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a method for correcting the position of a metal diaphragm based on YOLOv8, according to the present invention. Figure 2 This is an example image illustrating the preprocessing and enhancement of the original image of the same metal diaphragm piece according to the present invention; Figure 3 This is a detailed structural diagram of the metal diaphragm position correction network constructed in this invention. Detailed Implementation

[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding and implementation of how the present application uses technical means to solve technical problems and achieve technical effects.

[0021] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0022] Please refer to Figures 1-3 This illustration demonstrates a specific implementation of this embodiment. This embodiment improves the traditional YOLOv8 structure by designing a feature-level super-resolution fusion FSR mechanism to construct a metal diaphragm position correction network, improves the loss function of the traditional YOLOv8 based on the gradient equalization mechanism GHM, optimizes the selection of key point candidate boxes by designing an IoU-Guided NMS mechanism, designs core monitoring indicators and adds various special monitoring indicators, realizes the lightweight deployment of the metal diaphragm position correction network, completes the metal diaphragm position correction task, adapts to the key point positioning requirements of metal diaphragms, solves the problems of low efficiency of manual adjustment, poor robustness of traditional algorithms, and difficulty in technology implementation, and meets the high precision and high real-time requirements of industrial automated production.

[0023] Please refer to Figure 1 This embodiment proposes a method for correcting the position of a metal diaphragm based on YOLOv8, which includes the following steps: S1. Obtain the original images of the metal diaphragm at different placement angles, and perform preprocessing operations such as contrast enhancement, illumination intensity adjustment, and geometric transformation enhancement to label the real key points and their real category labels, and construct a dataset.

[0024] As a preferred embodiment of step S1, the specific process includes the following steps: S11. Using an industrial camera, photograph the metal diaphragm (100mm×150mm in size, including 3 positioning holes) under typical lighting conditions in the workshop (9:00-11:00 strong light, 14:00-16:00 side light, 18:00-20:00 weak light). Fix the position of the metal diaphragm using a conveyor belt positioning device, and control the distance between the industrial camera and the metal diaphragm to be 50cm to ensure that the image resolution is stable at 2448×2048 pixels. At the same time, adjust the placement angle of the metal diaphragm by rotating the fixture, covering one posture every 15° within the range of 0°-360°, and collect a total of 5000 original images of the metal diaphragm to ensure that the dataset covers all possible posture deviations that may occur in production.

[0025] S12. Perform contrast enhancement processing on the original metal diaphragm image by linearly stretching the grayscale range to highlight the feature regions of the metal diaphragm, facilitating key point recognition by the model (i.e., the constructed metal diaphragm position correction network); the linear stretching formula is: ; in, For contrast gain, a value of 1.2-1.5 is used (1.3 in this embodiment) to balance feature prominence and noise suppression; The value is set to brightness offset, ranging from 0 to 20 (15 in this embodiment), to highlight the characteristic areas of the metal diaphragm, such as the edges and positioning holes, and to avoid the loss of dark features. and These represent the original grayscale value of the metal diaphragm image and the grayscale value of the metal diaphragm image after contrast enhancement processing, respectively; the grayscale difference between the positioning hole and the edge of the metal diaphragm image is increased by more than 30% after processing.

[0026] S13. By randomly adjusting the light intensity, the luminance channel (V channel) of the contrast-enhanced metal film image is adjusted. Specifically, in the HSV color space, the luminance channel is randomly adjusted by ±10% to ±30% to simulate light changes during different production periods, enhancing the model's adaptability to light interference. The formula for this step is: ; in, This indicates that a random number between 0 and 1 is generated, which enhances the model's adaptability to lighting interference. , These represent the V channel brightness values ​​of the metal film image after contrast enhancement processing and the V channel brightness values ​​of the metal film image after illumination intensity adjustment, respectively.

[0027] S14. Apply horizontal flipping to the image of the metal diaphragm after adjusting the illumination intensity by 30% to eliminate left and right pose bias, and apply a random rotation of ±10° to the image of the metal diaphragm after adjusting the illumination intensity by 20% to simulate the offset of the metal diaphragm. In this step, the synchronous transformation of the key point coordinates is maintained to avoid annotation bias. Finally, 15,000 enhanced images are generated, and the dataset size is increased by 3 times, which effectively alleviates the overfitting of the model.

[0028] An example of preprocessing and enhancing the original metal sheet image can be seen here. Figure 2 , Figure 2 From left to right, the images are: the original image of the metal diaphragm, the metal diaphragm image with brightness increased by +10°, the metal diaphragm image with brightness decreased by +10°, the metal diaphragm image with horizontal flipping, and the metal diaphragm image with decreased brightness.

[0029] S15. Using the LabelImg annotation tool, annotate the centers of three non-collinear key positioning holes in each preprocessed metal diaphragm image as real key points. Record the coordinates of the real key points and their corresponding real category labels to generate a VOC format annotation file, which together with the preprocessed metal diaphragm images constitutes the dataset.

[0030] In this embodiment, a multi-dimensional data augmentation strategy, including contrast enhancement, light intensity adjustment, and geometric transformation enhancement, is adopted to significantly improve the model's adaptability to changes in the production environment and solve the positioning deviation problem caused by surface reflection and light interference of metal diaphragm parts.

[0031] S2. Improve the traditional YOLOv8 structure and construct a metal diaphragm position correction network; design a feature-level super-resolution fusion FSR mechanism based on lightweight ESPCN, add a feature-level super-resolution fusion FSR mechanism to the YOLOv8 structure and introduce an ECA attention mechanism to optimize feature extraction and localization accuracy, construct a metal diaphragm position correction network, and design specific monitoring indicators for the feature-level super-resolution fusion FSR mechanism.

[0032] As a preferred embodiment of step S2, the specific process includes the following steps: S21. A feature-level super-resolution fusion (FSR) mechanism is added to the Neck feature fusion layer of the YOLOv8 structure. The feature-level super-resolution fusion (FSR) mechanism is based on lightweight ESPCN for super-resolution reconstruction, extracts P3-FSR features, and then performs dynamic weighted fusion and smoothing processing to obtain integrated features.

[0033] The feature-level super-resolution fusion FSR mechanism adopts a lightweight ESPCN (Efficient Sub-Pixel Convolutional Neural Network) as its core structure, removing the fully connected layers and redundant convolutional blocks in the original ESPCN, and retaining only the core link of "feature extraction - sub-pixel convolution".

[0034] More specifically, in step S21, the specific process includes the following steps: S211. Extract P3 features (40×40 resolution, 128 channels) based on the feature extraction layer of lightweight ESPCN. In this embodiment, the feature extraction layer uses two 3×3 convolutions (the first layer extracts semantic features, with the number of input channels being the same as the number of P3 feature channels, which is 128, and the number of output channels being set to 256, and ReLU activation is applied; the second layer strengthens feature association, with the number of input and output channels both being 256, and ReLU activation is applied), capturing the semantic association information in the low-resolution P3 features.

[0035] S212. Based on the lightweight ESPCN, a subpixel convolutional layer matches the resolution difference between P2 and P3 features, improving the resolution of P3 features to match the resolution of P2 features, and obtaining P3-FSR features. In this embodiment, the subpixel convolutional layer uses a single 3×3 convolution (output channel number = target resolution magnification factor² × input channel number, where the magnification factor is set to 2 to match the resolution difference between P2 and P3 features). Through pixel rearrangement, the resolution of P3 features is improved from 40×40 to 80×80, which is consistent with the resolution of shallow P2 features (80×80). This process uses PyTorch's PixelShuffle function to achieve pixel rearrangement, ensuring the quality of super-resolution reconstruction.

[0036] S213. The number of channels of the P3-FSR feature (80×80 resolution, 1024 channels) is unified with that of the P2 feature (80×80 resolution, 256 channels) through a convolutional layer (set to 1×1 convolution in this embodiment), eliminating the feature imbalance caused by the difference in channel dimensions. Then, the following dynamic weighted fusion is performed: ;

[0037] in , These represent the P2 features in the dynamic weighted fusion process. and P3-FSR characteristics The weight, For fusion features; after dynamic weighted fusion, the fusion features The image is then smoothed again by a convolutional layer (3×3 convolution with stride 1 and padding 1 in this embodiment), integrating feature information, reducing feature fluctuations after weight fusion, and preserving the details and semantic associations of subtle features of the metal membrane (such as positioning holes and edge corners) to obtain integrated features. In this embodiment, the weights... , Acquired through dynamic learning, during initialization. Set it to 0.6. Set to 0.4, the value can be dynamically adjusted during training based on the keypoint localization loss. The formula is as follows: ; in, This indicates the number of key points in a single image. , For the first The x and y coordinates of the predicted key points , For the first x and y coordinates of real key points; key point localization loss. The Euclidean distance between the predicted keypoint and the real keypoint is measured. If the localization deviation of features such as P3-FSR is smaller, its weight will be increased through backpropagation in subsequent training. Conversely, the weight of P2 features will be increased.

[0038] S22. Add an ECA attention mechanism after each of the six C2f modules in the YOLOv8 structure. Set the kernel size of the ECA attention mechanism to 3 (to adapt to the feature scale of the metal diaphragm). By adaptively calculating the feature channel weights of the metal diaphragm, enhance the model's attention to the feature region of the metal diaphragm (the weight value is increased to 2-3 times that of the background region) and suppress background interference such as workbench surface and clutter.

[0039] S23. The design includes specific monitoring indicators for the feature-level super-resolution fusion FSR mechanism, calculating the structural similarity between the P3-FSR features and the P2 features extracted by the metal diaphragm position correction network in the key point region. ,Require (This indicates that the features of the two are well consistent). At the same time, the localization accuracy of small key points (5-10 pixels) before and after the intervention of the feature-level super-resolution fusion FSR mechanism is statistically analyzed. The accuracy improvement is required to be ≥3%. If the target is not met, the number of convolution kernels or the magnification factor of the sub-pixel convolution layer of the feature-level super-resolution fusion FSR mechanism is adjusted to ensure that the feature-level super-resolution fusion FSR mechanism can play an effective role.

[0040] For the specific structure of the metal film position correction network built based on the YOLOv8 structure in this embodiment, please refer to [link / reference needed]. Figure 3This paper retains the CBS module (composed of convolutional layers Conv, batch normalization layers BN, and sigmoid linear units SiLU), C2f module, and SPPF (Spatial Pyramid Pooling Fast) module from the YOLOv8 architecture. Following the C2f module, the ECA attention mechanism designed in this invention is introduced, and a feature-level super-resolution fusion (FSR) mechanism is added to the Neck feature fusion layer. SPPF incorporates a max-pooling layer (MaxPool), whose main function is to pool information across multiple regions of different scales, thereby extracting richer features and improving detection accuracy. The feature-level super-resolution fusion (FSR) mechanism connects feature maps of different scales, helping the model better capture the details and contextual information of the target.

[0041] In this embodiment, by adding a feature-level super-resolution fusion FSR mechanism and an ECA attention mechanism, YOLOv8's ability to identify small-sized key points of metal diaphragms is enhanced, the accuracy of key point positioning is improved, and the angle calculation error is reduced.

[0042] The design combining lightweight ESPCN with YOLOv8 is a customized optimization addressing the core pain points of metal diaphragm component position correction (difficulty in locating small key points, poor feature matching, and high real-time requirements). Its special significance and effects closely align with the needs of industrial scenarios. Surface reflections on metal diaphragms and changes in ambient lighting (such as strong or weak light) can lead to low contrast between the positioning hole and the background. Traditional YOLOv8 feature extraction is susceptible to noise interference, resulting in misjudgments, such as misidentifying scratches as positioning holes. ESPCN's feature extraction layer enhances the semantic features of small key points, such as the circular outline and grayscale distribution of positioning holes, suppressing invalid information from reflections and noise. The designed feature-level super-resolution fusion (FSR) mechanism with dynamic weighted fusion adaptively balances details and semantics, stabilizing key point positioning deviations to ≤2 pixels, meeting core monitoring indicators. Precise feature fusion ensures accurate calculation of the spatial relationship between three non-collinear key points, resulting in an angle calculation error of ≤0.5°. The robotic arm can then precisely perform correction operations based on the calculated angle, thus avoiding subjective errors from manual adjustments. The ultimate goal is to achieve a key point positioning accuracy of ≥98%, an angle calculation error of ≤0.5°, and a significant reduction in correction time, thereby solving the problems of low efficiency of manual adjustment, poor robustness of traditional algorithms, and technology implementation.

[0043] S3. Improve the loss function of traditional YOLOv8; design a hybrid loss that combines Focal Loss and gradient equalization mechanism GHM to replace the classification loss of traditional YOLOv8. At the same time, introduce the gradient equalization mechanism GHM into the bounding box regression loss of traditional YOLOv8 to obtain the GHM-Reg loss, and design a special monitoring index for the gradient equalization mechanism GHM.

[0044] As a preferred embodiment of step S3, the specific process includes the following steps: S31. Introduce dynamic weight coefficients based on Focal loss. ,in Indicates sample The true category label (when calculating the loss function, "sample" refers to the network input; when training the constructed metal diaphragm position correction network of this invention, "sample" refers to the preprocessed metal diaphragm image from the dataset as input) That is, the true category label corresponding to it in the dataset). Follow The indicated category is dynamically adjusted for the characteristic region samples of the metal diaphragm. For background region samples The sample ratio of the feature area and the background area of ​​the metal film was initially balanced.

[0045] S32. Based on the gradient equalization mechanism (GHM), the classification gradient magnitude of each training sample is statistically analyzed, and a classification gradient density histogram is constructed. The classification gradient density histogram divides the classification gradient magnitude into 10-15 gradient intervals, and the classification gradient density weight of each gradient interval is calculated. ; In the formula, Indicates the first The classification gradient density weights for each gradient interval. The total number of samples, This represents the number of gradient intervals. For the first Number of samples in each gradient interval; weighted by classification gradient density. Balancing the classification gradient contributions of different samples; classification gradient magnitude The calculation method is as follows: ; in, For the sample The predicted probability.

[0046] S33. Construct a hybrid loss by combining dynamic weight coefficients and classification gradient density weights. as follows: ; In the formula, For the first A set of samples in gradient intervals The coefficient for adjusting the difficulty of Focal Loss is set to 2 in this embodiment. This hybrid loss retains Focal Loss's focus on difficult samples while suppressing the excessive contribution of extremely easy samples (such as those without interfering background) through GHM. At the same time, it avoids gradient explosion of difficult abnormal samples (such as stains being misclassified as features), thereby further improving the recall rate and localization stability of key points.

[0047] S34. Using the Smooth L1 loss as the basic bounding box regression loss (Smooth L1 loss is a common loss function used to calculate bounding box regression loss), calculate the regression gradient density weights for the regression gradient magnitude of each training sample in the same manner as in step S32. Construct GHM-Reg loss as follows: ; In the formula, Indicates Smooth L1 loss, For the sample The true bounding box parameters, For the sample Predicted bounding box parameters; by regressing gradient density weights Reduce the impact of extreme biased samples (i.e., samples with severely offset bounding boxes) on regression accuracy; regression gradient magnitude The calculation method is as follows: ; In this embodiment, the classification gradient magnitude is divided into 12 gradient intervals (interval range 0-0.05, 0.05-0.1, ..., 0.95-1.0), and the regression gradient magnitude is divided into 10 intervals (interval range 0-0.2, 0.2-0.4, ..., 1.8-2.0).

[0048] S35. The design includes specific monitoring metrics for the gradient equalization mechanism GHM. After each training round, the distribution of classification gradient magnitude and regression gradient magnitude of samples in the training and validation sets is statistically analyzed to ensure that the gradients are concentrated within a preset effective range (classification gradient magnitude 0.1-1.0, regression gradient magnitude 0.5-3.0). If the gradients exceed this effective range, the corresponding gradient interval division or loss function parameters are adjusted. Simultaneously, the classification gradient density weights of each gradient interval are monitored. and regression gradient density weights The fluctuation range is required to be ≤15% for three consecutive rounds to ensure the training stability of the metal diaphragm position correction network.

[0049] In this embodiment, the loss function of traditional YOLOv8 is improved based on the gradient equalization mechanism GHM. The constructed hybrid loss balances the ratio of positive and negative samples through Focal Loss on the one hand, and suppresses the overtraining of extremely easy-to-distinguish samples (such as pure backgrounds) and the gradient interference of extremely difficult-to-distinguish samples (such as misjudgment of stains) through GHM on the other hand. This makes the model training more focused on effective metal film feature samples, further improving the accuracy of key point localization and reducing angle calculation error.

[0050] The loss function of this invention is designed to address the unique challenges of "small key points, strong reflective interference, and extreme sample distribution" in the context of metal membrane position correction: In the classification loss, a dynamic weight coefficient is introduced that dynamically adjusts with the feature region of the metal membrane and the background region. The design assigns higher weights (0.7) to samples in feature regions, enhancing the model's focus on core targets such as positioning holes. Simultaneously, based on the gradient equalization mechanism GHM, the classification gradient magnitude is divided into 10-15 intervals to suppress gradient interference from extremely easily distinguishable samples (such as clean backgrounds) and extremely difficult samples (such as misclassified dirt), preventing the model from being biased towards invalid samples during training. In the regression loss, GHM is introduced into the bounding box regression loss to obtain the GHM-Reg loss. By statistically analyzing the regression gradient density weights, the impact of severely offset bounding box samples on positioning accuracy is reduced, adapting to the fine positioning requirements of small keypoints (5-10 pixels). Furthermore, the gradient density weight fluctuation range is required to be ≤15% for three consecutive rounds to ensure training stability. This design not only solves the dual imbalance problem of positive and negative samples (feature regions and background) and difficult and easy samples (clear positioning holes and disturbed positioning holes) in the scene, but also, through precise gradient equalization, enables the model to focus on effective feature samples, ultimately achieving high precision requirements of keypoint positioning deviation ≤2 pixels and angle calculation error ≤0.5°, providing reliable parameter support for robotic arm linkage correction.

[0051] S4. Design an IoU-Guided NMS mechanism to replace the traditional YOLOv8 nonmaximum suppression mechanism in the inference stage of the metal diaphragm position correction network, optimize the screening of key point candidate boxes, obtain a list of screened key point boxes to guide the generation of predicted key points, and design specific monitoring indicators for the IoU-Guided NMS mechanism.

[0052] As a preferred embodiment of step S4, the specific process includes the following steps: S41. Sort the candidate keypoint boxes of the same category from highest to lowest confidence level, and select the box with the highest confidence level as the baseline box. .

[0053] S42, Calculate the baseline frame Compared with other keypoint candidate boxes Intersection over Union (IOU).

[0054] S43. Introducing distance weighting coefficient Define the guidance inhibition threshold : ; in The basic suppression threshold is set to 0.3-0.5 (in this embodiment, it is set to 0.35 for small keypoint characteristics). This indicates the intersection-union ratio (IU) calculation operation; distance weighting coefficient. Candidate boxes for each key point With reference frame center distance Average size of key points The ratio is determined by the formula: ; in The attenuation coefficient ranges from 1.2 to 1.5 and is dynamically adjusted based on the density of key points in the metal diaphragm. The higher the density of key points, the larger the attenuation coefficient (in this embodiment, it is set to 1.3 to match the density of positioning holes in the metal diaphragm). The base of the natural logarithm; guiding the inhibition threshold. The above formula enables dynamic adjustment, thereby reducing both the duplicate frame rate and the false negative rate.

[0055] Distance weighting coefficient The design is based on: center distance The larger, The smaller the value, the lower the inhibition threshold. The closer This allows for more rigorous removal of long-distance low IoU interference boxes; The smaller (near the key point) The larger the value, the higher the inhibition threshold. The lower the value, the more candidate boxes for nearby key points can be retained.

[0056] S44. If key point candidate boxes With reference frame The intersection-union ratio ≥ Then inhibit If the intersection-union ratio is < Then keep As a new baseline box, repeat the above process until all keypoint candidate boxes have been processed, and the filtering is complete.

[0057] S45. The design includes specific monitoring indicators for the IoU-Guided NMS mechanism. The proportion of duplicate points (the same keypoint is marked by two or more bounding boxes) in the predicted keypoints generated from the test set is calculated, with a duplication rate ≤ 1%. Simultaneously, the proportion of undetected real keypoints in the test set is calculated, with a false negative rate ≤ 0.5%. If the above indicators are not met, return to step S43 to adjust the attenuation coefficient. or baseline inhibition threshold Then, the inference and verification of the metal diaphragm position correction network were repeated until the target was met.

[0058] In this embodiment, the IoU-Guided NMS mechanism avoids the missed detections caused by fixed thresholds in traditional NMS (such as the false suppression of nearby small keypoints) and solves the problem of ineffective removal of low IoU redundant boxes. By designing a dynamic suppression threshold to adapt to differences in keypoint density, the duplicate box rate and the missed detection rate are reduced. At the same time, the processing time on edge devices is reduced, which does not affect the real-time performance of model inference and ensures the efficiency of robotic arm adjustment.

[0059] S5. Using key point localization accuracy and angle calculation error as core monitoring indicators, and combining them with specific monitoring indicators for feature-level super-resolution fusion FSR mechanism, gradient equalization mechanism GHM and IoU-Guided NMS mechanism, the dataset constructed in step S1 is divided into training set, validation set and test set, and a metal diaphragm position correction network is trained; the metal diaphragm position correction network outputs predicted key points and calculates the predicted deflection angle of the metal diaphragm based on the coordinates of the predicted key points.

[0060] As a preferred embodiment of step S5, the process of using key point positioning accuracy and angle calculation error as core monitoring indicators includes the following steps: S51. The metal diaphragm piece position correction network detects and outputs the coordinates of three non-collinear prediction key points. , , ,by Construct a vector for the origin ,calculate The perpendicular vector is used as the normal vector. With the horizontal direction to the right as the horizontal reference direction (vector) The angle between the normal vector and the horizontal reference direction is calculated using the dot product formula. : ; In the formula For inverse cosine calculation; included angle This is the deflection angle of the metal diaphragm. If the result is 15°, then the metal diaphragm needs to be rotated 15° counterclockwise.

[0061] S52. The accuracy of key point localization is calculated by determining the Euclidean distance between the predicted key point coordinates and the actual key point coordinates. As a key point positioning deviation, when When the pixel value is specified, it is considered accurate.

[0062] S53, the aforementioned angle calculation error is for predicting the deflection angle of the metal diaphragm. The difference between the actual deflection angle of the metal diaphragm and the actual deflection angle of the metal diaphrag

[0063] S54. Stop training when the validation set metrics do not improve for several consecutive rounds.

[0064] In this embodiment, the dataset is divided into a training set, a validation set, and a test set in a 7:2:1 ratio. The training set is used for model parameter learning, the validation set is used for hyperparameter tuning, and the test set is used to evaluate the final model performance. The model training uses the AdamW optimizer with an initial learning rate of 0.001. The learning rate is dynamically adjusted using a cosine annealing strategy to avoid the model training getting stuck in local optima. For the feature-level super-resolution fusion FSR mechanism, a differentiated learning rate strategy is adopted, setting the learning rate for the feature-level super-resolution fusion FSR mechanism to 1.2 times the overall model learning rate, prioritizing the optimization of super-resolution reconstruction-related parameters, and accelerating the feature matching process.

[0065] S6. Encapsulate and deploy the trained metal diaphragm position correction network, design an acceleration optimization strategy for the feature-level super-resolution fusion FSR mechanism, acquire metal diaphragm images and use the deployed metal diaphragm position correction network to predict key points and the predicted deflection angle of the metal diaphragm. Before the robotic arm linkage control, the IoU-Guided NMS mechanism is used again to screen candidate boxes for key points. Based on the acquired predicted key point coordinates and predicted deflection angle of the metal diaphragm, the robotic arm linkage control is performed, and the position correction effect of the metal diaphragm is verified based on the core monitoring indicators until the verification result meets the standard and the position correction of the metal diaphragm is completed.

[0066] As a preferred embodiment of step S6, the specific process includes the following steps: S61. Package the trained metal film device position correction network (weight file about 200MB) into a lightweight inference engine (such as an ONNX format file), and deploy it to the edge computing device through the ONNX Runtime tool. Remove training-related layers (such as Dropout) during the deployment process.

[0067] S62. For the feature-level super-resolution fusion FSR mechanism, inference acceleration optimization is performed. TensorRT is used to fuse the feature extraction layer and sub-pixel convolution layer of the lightweight ESPCN into a single custom operator, reducing memory access overhead. At the same time, the weight precision of the feature-level super-resolution fusion FSR mechanism is quantized from FP32 to FP16. While ensuring the quality of feature reconstruction, the inference speed of the metal film component position correction network meets the real-time detection requirements. The optimized model size is compressed to 100MB, and the inference speed is increased by 2 times. The inference time on edge devices meets the real-time detection requirements.

[0068] S63. The edge computing device acquires images of the metal diaphragm in real time using an industrial camera, calculates and outputs a predicted deflection angle of the metal diaphragm using a metal diaphragm position correction network. .

[0069] S64. Robotic arm linkage control: Edge computing device deflects the metal diaphragm component by an angle. (e.g., 15°) Deviation from key point positioning (For example, an X-axis offset of 5mm and a Y-axis offset of 3mm) is sent to the robotic arm controller, which drives the robotic arm's end effector to grasp the metal diaphragm and perform rotation. Translation , operate.

[0070] S65. Perform effect verification. After the robotic arm is adjusted, re-acquire the image of the metal diaphragm. Verify that the key point positioning accuracy is ≥98% and the angle calculation error is ≤0.5°. If the verification result is satisfactory, it means that there is a physical error in the actual execution of the robotic arm. That is, there is a physical error between the actual position of the metal diaphragm after the robotic arm is adjusted and the "adjusted position" expected by the model. At this time, return to step S63, use an industrial camera to re-acquire the image of the metal diaphragm after the robotic arm is adjusted in real time, run the metal diaphragm position correction network again, re-detect the coordinates of the three non-collinear key points, recalculate the deflection angle of the metal diaphragm and the deviation of the key point positioning, until the verification result is satisfactory, and the position correction of the metal diaphragm is completed.

[0071] Repeat the adjustment steps until the verification results meet the standards, thus completing the position correction of the metal diaphragm.

[0072] More specifically, before the robotic arm performs the position correction of the metal diaphragm, the IoU-Guided NMS mechanism is applied again to receive the filtered list of keypoint bounding boxes sent by the edge device. The list contains verification flags indicating no duplicate or missed bounding boxes for the predicted keypoints. If duplicate or missed bounding boxes are found in the list, an anomaly verification flag is displayed, the robotic arm controller triggers a pause command, and the edge computing device automatically readjusts the attenuation coefficient. or baseline inhibition threshold Continue until a list of qualified keypoint boxes is output, and then determine the keypoints.

[0073] In this embodiment, to address the density differences of key points in the image of the metal diaphragm, the edge device can automatically adapt the attenuation coefficient of IoU-Guided NMS by reading the metal diaphragm model configuration file. The processed key points must meet the "single key point, single frame" principle. Then, the deflection angle of the metal diaphragm is calculated based on the coordinates of the screened key points to ensure the accuracy of the angle calculation (avoiding angle deviation caused by duplicate frames or missed frames).

[0074] This embodiment improves the localization accuracy of small-sized keypoints by employing a lightweight design and inference optimization of the feature-level super-resolution fusion (FSR) mechanism. Through TensorRT quantization and operator fusion, the inference time of the FSR mechanism is controlled to within 2ms, ensuring the overall model's real-time performance meets the demands of industrial scenarios. Compared to solutions without this mechanism, the localization accuracy of small-sized keypoints is significantly enhanced. Furthermore, by constructing an integrated workflow combining "model inference" and "robotic arm control," the model output directly drives hardware adjustments without manual intervention, shortening the time required for correcting the position of metal diaphragm components and adapting to high-speed automated production lines.

[0075] This application also provides a metal diaphragm position correction system based on YOLOv8, which is used to apply the metal diaphragm position correction method based on YOLOv8, including: an image acquisition and preprocessing module, a model optimization module, and an inference and control module. The image acquisition and preprocessing module is used to acquire original images of metal film components at different placement angles and to perform preprocessing operations such as contrast enhancement, light intensity adjustment, and geometric transformation enhancement. The model optimization module is used to execute the feature-level super-resolution fusion FSR mechanism, the ECA attention mechanism, the calculation of the hybrid loss and GHM-Reg loss combined with the gradient equalization mechanism GHM, and the parameter optimization of the IoU-Guided NMS mechanism. The model optimization module has a built-in FSR parameter configuration submodule, which dynamically adjusts the magnification factor and convolution kernel parameters of the super-resolution reconstruction according to the key point size of the metal film in the feature-level super-resolution fusion FSR mechanism, so as to adapt to the feature fusion requirements of metal film with different specifications. The inference and control module is used to deploy the trained metal diaphragm position correction network to the edge computing device, and output the coordinates of predicted key points and the predicted deflection angle of the metal diaphragm in real time, driving the robotic arm to perform the position correction operation of the metal diaphragm. The inference and control module includes an FSR inference acceleration unit, which ensures the efficient operation of the feature-level super-resolution fusion FSR mechanism on the edge computing device through operator fusion and precision quantization.

[0076] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0077] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0078] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for correcting the position of a metal diaphragm based on an improved YOLOv8, characterized in that, Includes the following steps: S1. Obtain the original images of the metal diaphragm at different placement angles, and perform preprocessing operations such as contrast enhancement, illumination intensity adjustment, and geometric transformation enhancement to label the real key points and their real category labels, and construct a dataset. S2. Based on the lightweight ESPCN design, a feature-level super-resolution fusion FSR mechanism is added to the YOLOv8 structure and an ECA attention mechanism is introduced to optimize feature extraction and localization accuracy. A metal film component position correction network is constructed, and a special monitoring index for the feature-level super-resolution fusion FSR mechanism is designed. S3. Design a hybrid loss that combines Focal Loss and gradient equalization mechanism GHM to replace the traditional YOLOv8 classification loss. At the same time, introduce the gradient equalization mechanism GHM into the traditional YOLOv8 bounding box regression loss to obtain the GHM-Reg loss. Design a special monitoring index for the gradient equalization mechanism GHM. S4. Design an IoU-Guided NMS mechanism to replace the traditional YOLOv8 nonmaximum suppression mechanism in the inference stage of the metal film device position correction network, optimize the screening of key point candidate boxes, obtain a list of screened key point boxes to guide the generation of predicted key points, and design specific monitoring indicators for the IoU-Guided NMS mechanism. S5. Using key point localization accuracy and angle calculation error as core monitoring indicators, and combining specific monitoring indicators for feature-level super-resolution fusion FSR mechanism, gradient equalization mechanism GHM and IoU-Guided NMS mechanism, the dataset constructed in step S1 is divided into training set, validation set and test set, and a metal diaphragm position correction network is trained; the metal diaphragm position correction network outputs predicted key points and calculates the predicted deflection angle of the metal diaphragm based on the coordinates of the predicted key points. S6. Encapsulate and deploy the trained metal diaphragm position correction network, design an acceleration optimization strategy for the feature-level super-resolution fusion FSR mechanism, acquire metal diaphragm images and use the deployed metal diaphragm position correction network to predict key points and the predicted deflection angle of the metal diaphragm. Before the robotic arm linkage control, the IoU-Guided NMS mechanism is used again to screen candidate boxes for key points. Based on the acquired predicted key point coordinates and predicted deflection angle of the metal diaphragm, the robotic arm linkage control is performed, and the position correction effect of the metal diaphragm is verified based on the core monitoring indicators until the verification result meets the standard and the position correction of the metal diaphragm is completed.

2. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 1, characterized in that: Step S1 specifically includes the following steps: S11. Use an industrial camera to capture images of the original metal film parts under different lighting conditions. The shooting process covers multiple placement angles to ensure that the dataset covers the posture deviations that may occur in actual production. S12. Perform contrast enhancement processing on the original metal diaphragm image by linearly stretching the grayscale range to highlight the feature areas of the metal diaphragm, which facilitates the identification of key points. S13. By randomly adjusting the light intensity, the brightness channel of the metal film image after contrast enhancement is adjusted to simulate the light changes during different production periods. S14. Apply horizontal flipping and random rotation geometric transformations to the images of metal film components after adjusting a certain proportion of the illumination intensity to complete the preprocessing operation and expand the dataset size. S15. For each preprocessed metal diaphragm image, label the centers of three non-collinear key positioning holes as real key points, record the coordinates of the real key points and their corresponding real category labels, and together with the preprocessed metal diaphragm images, form a dataset.

3. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 1, characterized in that: Step S2 specifically includes the following steps: S21. A feature-level super-resolution fusion (FSR) mechanism is added to the Neck feature fusion layer of the YOLOv8 structure. The feature-level super-resolution fusion (FSR) mechanism is based on lightweight ESPCN for super-resolution reconstruction, extracts P3-FSR features, and then performs dynamic weighted fusion and smoothing processing to obtain integrated features. S22. An ECA attention mechanism is introduced after each C2f module of the YOLOv8 structure to complete the construction of the metal diaphragm position correction network. S23. The design includes specific monitoring indicators for the feature-level super-resolution fusion FSR mechanism, calculating the structural similarity between the P3-FSR features and the P2 features extracted by the metal diaphragm position correction network in the key point region. ,Require Simultaneously, the changes in localization accuracy of small keypoints before and after the intervention of the feature-level super-resolution fusion (FSR) mechanism are statistically analyzed, with an accuracy improvement of ≥3%. If the target is not met, the number of convolution kernels or the magnification factor of the sub-pixel convolutional layer in the feature-level super-resolution fusion (FSR) mechanism are adjusted to ensure that the feature-level super-resolution fusion (FSR) mechanism can play an effective role.

4. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 3, characterized in that: Step S21 specifically includes the following steps: S211. Extract P3 features based on the feature extraction layer of lightweight ESPCN; S212. Based on the lightweight ESPCN subpixel convolutional layer, the resolution difference between P2 and P3 features is matched, and the resolution of P3 features is improved to be consistent with that of P2 features to obtain P3-FSR features. S213. The number of channels of the P3-FSR feature is unified with that of the P2 feature through a convolutional layer, and then the following dynamic weighted fusion is performed: ; in , These represent the P2 features in the dynamic weighted fusion process. and P3-FSR characteristics The weight, For fusion features; after dynamic weighted fusion, the fusion features The image is then smoothed again by a convolutional layer to obtain integrated features.

5. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 1, characterized in that: Step S3 specifically includes the following steps: S31. Introduce dynamic weight coefficients based on Focal loss. ,in Indicates sample The true category label, Follow The indicated category is dynamically adjusted; S32. Based on the gradient equalization mechanism (GHM), the classification gradient magnitude of each training sample is statistically analyzed, and a classification gradient density histogram is constructed. The classification gradient density histogram divides the classification gradient magnitude into 10-15 gradient intervals, and the classification gradient density weight of each gradient interval is calculated. ; In the formula, Indicates the first The classification gradient density weights for each gradient interval. The total number of samples, This represents the number of gradient intervals. For the first The number of samples in each gradient interval; S33. Construct a hybrid loss by combining dynamic weight coefficients and classification gradient density weights. as follows: ; In the formula, For the first A set of samples in gradient intervals For the sample The predicted probability, The easy and difficult sample adjustment coefficient for Focal Loss; S34. Using Smooth L1 loss as the basic bounding box regression loss, calculate the regression gradient density weights for the regression gradient magnitude of each training sample in the same manner as in step S32. Construct GHM-Reg loss as follows: ; In the formula, Indicates Smooth L1 loss, For the sample The true bounding box parameters, For the sample Predicted bounding box parameters; S35. The design includes specific monitoring metrics for the gradient equalization mechanism GHM. After each training round, the distribution of classification gradient magnitude and regression gradient magnitude of samples in the training and validation sets is statistically analyzed to ensure that the gradients are concentrated within a preset effective interval. If the gradients exceed this effective interval, the corresponding gradient interval division or loss function parameters are adjusted. Simultaneously, the classification gradient density weights of each gradient interval are monitored. and regression gradient density weights The fluctuation range is required to be ≤15% for three consecutive rounds to ensure the training stability of the metal diaphragm position correction network.

6. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 1, characterized in that: Step S4 specifically includes the following steps: S41. Sort the candidate keypoint boxes of the same category from highest to lowest confidence level, and select the box with the highest confidence level as the baseline box. ; S42, Calculate the baseline frame Compared with other keypoint candidate boxes The intersection and union ratio; S43. Introducing distance weighting coefficient Define the guidance inhibition threshold : ; in The basic suppression threshold is set, with a value ranging from 0.3 to 0.

5. This indicates the intersection-union ratio (IU) calculation operation; distance weighting coefficient. Candidate boxes for each key point With reference frame center distance Average size of key points The ratio is determined by the formula: ; in The attenuation coefficient ranges from 1.2 to 1.5 and is dynamically adjusted based on the density of key points in the metal diaphragm. The higher the density of key points, the larger the attenuation coefficient. It is the base of the natural logarithm; S44. If key point candidate boxes With reference frame The intersection-union ratio ≥ Then inhibit If the intersection-union ratio is < Then keep As a new baseline box, repeat the above process until all candidate boxes for key points have been processed and the filtering is complete. S45. The design includes specific monitoring indicators for the IoU-Guided NMS mechanism. The proportion of duplicate points in the predicted keypoints generated from the test set to the total predicted keypoints is required to have a duplication rate ≤ 1%. Simultaneously, the proportion of undetected real keypoints in the test set to the total real keypoints is required to have a false negative rate ≤ 0.5%. If the above indicators are not met, return to step S43 to adjust the attenuation coefficient. or baseline inhibition threshold Then, the inference and verification of the metal diaphragm position correction network were repeated until the target was met.

7. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 1, characterized in that: In step S5, the key point positioning accuracy and angle calculation error are the core monitoring indicators. The specific process includes the following steps: S51. The metal diaphragm piece position correction network detects and outputs the coordinates of three non-collinear prediction key points. , , ,by Construct a vector for the origin ,calculate The perpendicular vector is used as the normal vector. Calculate the angle between the normal vector and the horizontal reference direction. This is the prediction of the deflection angle of the metal diaphragm component; S52. The accuracy of key point localization is calculated by determining the Euclidean distance between the predicted key point coordinates and the actual key point coordinates. As a key point positioning deviation, when When the pixel count is used, it is considered accurate. S53, the aforementioned angle calculation error is for predicting the deflection angle of the metal diaphragm. The difference between the actual deflection angle of the metal diaphragm and the actual deflection angle of the metal diaphrag S54. Stop training when the validation set metrics do not improve for several consecutive rounds.

8. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 7, characterized in that: Step S6 specifically includes the following steps: S61. The trained metal diaphragm position correction network is packaged into a lightweight inference engine and deployed to an edge computing device; S62. For the feature-level super-resolution fusion FSR mechanism, inference acceleration optimization is performed. TensorRT is used to fuse the feature extraction layer and sub-pixel convolution layer of the lightweight ESPCN into a single custom operator to reduce memory access overhead. At the same time, the weight precision of the feature-level super-resolution fusion FSR mechanism is quantized from FP32 to FP16. While ensuring the quality of feature reconstruction, the inference speed of the metal film piece position correction network meets the real-time detection requirements. S63. The edge computing device acquires images of the metal diaphragm in real time using an industrial camera, calculates and outputs a predicted deflection angle of the metal diaphragm using a metal diaphragm position correction network. ; S64. Robotic arm linkage control: Edge computing device deflects the metal diaphragm component by an angle. Deviation from key point positioning The data is sent to the robotic arm controller, which then drives the robotic arm to... To correct the angle, To compensate for positional deviations, rotation and translation operations are performed; S65. Perform effect verification. After the robotic arm is adjusted, re-acquire the image of the metal diaphragm. Verify that the key point positioning accuracy is ≥98% and the angle calculation error is ≤0.5°. If the verification result is satisfactory, return to step S63 and repeat steps S63-S65 with the metal diaphragm after the robotic arm is adjusted until the verification result is satisfactory, thus completing the position correction of the metal diaphragm.

9. The method for correcting the position of a metal diaphragm based on YOLOv8 as described in claim 6, characterized in that: Before the robotic arm corrects the position of the metal diaphragm, the IoU-Guided NMS mechanism is applied again to receive the filtered list of keypoint bounding boxes sent by the edge device. The list contains verification flags indicating that the predicted keypoints are free of duplicates and omissions. If duplicates or omissions are found in the list, an anomaly verification flag is displayed, and the robotic arm controller triggers a pause command. At the same time, the edge computing device automatically readjusts the attenuation coefficient. or baseline inhibition threshold Continue until a list of qualified keypoint boxes is output, and then determine the keypoints.

10. A YOLOv8-based improved metal diaphragm position correction system, used for applying the YOLOv8-based improved metal diaphragm position correction method as described in any one of claims 1-9, characterized in that, include: Image acquisition and preprocessing module, model optimization module, inference and control module; The image acquisition and preprocessing module is used to acquire original images of metal film components at different placement angles and to perform preprocessing operations such as contrast enhancement, light intensity adjustment, and geometric transformation enhancement. The model optimization module is used to execute the feature-level super-resolution fusion FSR mechanism, the ECA attention mechanism, the calculation of the hybrid loss and GHM-Reg loss combined with the gradient equalization mechanism GHM, and the parameter optimization of the IoU-Guided NMS mechanism. The model optimization module has a built-in FSR parameter configuration submodule, which dynamically adjusts the magnification factor and convolution kernel parameters of the super-resolution reconstruction according to the key point size of the metal film in the feature-level super-resolution fusion FSR mechanism, so as to adapt to the feature fusion requirements of metal film with different specifications. The inference and control module is used to deploy the trained metal diaphragm position correction network to the edge computing device, and output the coordinates of predicted key points and the predicted deflection angle of the metal diaphragm in real time, driving the robotic arm to perform the position correction operation of the metal diaphragm. The inference and control module includes an FSR inference acceleration unit, which ensures the efficient operation of the feature-level super-resolution fusion FSR mechanism on the edge computing device through operator fusion and precision quantization.