Workpiece, workpiece key point real-time detection method and related device
By using the lightweight feature extraction network combined with the YOLOv5 full convolution network and the CoordAttention attention mechanism in the workpiece detection, and integrating the improved BiFPN module in the Neck layer, the problems of high detection error rate and poor robustness in the key point detection of workpieces and workpieces are solved, real-time detection in the CPU environment is achieved.
Patent Information
- Application Number
- CN202510117288.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems with high detection error rate and poor robustness in the detection of key points of workpieces and workpieces, and methods based on mainstream object detection algorithms are difficult to achieve real-time detection in the CPU environment.
The YOLOv5 full convolution network is used to combine the CoordAttention attention mechanism to design a lightweight feature extraction network, and the improved BiFPN module is fused on the Neck layer to add artifact key point detection tasks, and model training is performed using the Wing Loss loss function and the Hardswish activation function.
While ensuring detection accuracy, the detection speed is significantly improved, real-time inference is realized in the CPU environment, and the problems of high detection error rate and poor robustness are solved.
Smart Images

Figure CN120047406A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of industrial vision, and relates to a method and related device for real-time detection of workpieces and workpiece key points. Background Art
[0002] In recent years, deep learning technology has developed rapidly, bringing new solutions to many industrial problems. Especially in the field of industrial production, more and more machine vision-based solutions are applied to intelligent production in various industries. For example, the detection task of workpieces and workpiece key points requires visually locating the position of the workpiece and at the same time locating the preset key points in the workpiece.
[0003] When using traditional detection methods for workpiece and workpiece key point detection, there are problems such as high detection error rate and poor robustness. Using the current mainstream object detection algorithms requires a large amount of computation, and at the same time, an additional detection task for workpiece key points is required, making it difficult to achieve real-time detection in a CPU environment. Summary of the Invention
[0004] The purpose of this application is to solve the problems in the prior art and provide a method and related device for real-time detection of workpieces and workpiece key points.
[0005] To achieve the above purpose, this application adopts the following technical solutions:
[0006] In the first aspect, this application provides a method for real-time detection of workpieces and workpiece key points, including the following steps:
[0007] Obtain a workpiece image, and obtain the image data of the workpiece and the annotation data of its key points according to the workpiece image;
[0008] Input the image data of the workpiece and the annotation data of its key points into a pre-trained real-time detection model for workpieces and workpiece key points to obtain the detection result of the workpiece key points;
[0009] Among them, the pre-trained real-time detection model for workpieces and workpiece key points is trained using the image data of the original workpiece and the annotation data of its key points;
[0010] The pre-trained real-time detection model for workpieces and workpiece key points adopts the YOLOv5 fully convolutional network, including: Backbone feature extraction network, Neck layer and Prediction layer, where the Backbone feature extraction network has a CoordAttention attention mechanism.
[0011] In the second aspect, this application provides a real-time detection system for workpieces and workpiece key points, including:
[0012] An image acquisition module, configured to acquire a workpiece image and obtain the image data of the workpiece and the annotation data of its key points based on the workpiece image;
[0013] A result detection module, configured to input the image data of the workpiece and the annotation data of its key points into a pre-trained real-time detection model for workpieces and workpiece key points to obtain the detection results of the workpiece key points;
[0014] Wherein, the pre-trained real-time detection model for workpieces and workpiece key points is trained by using the image data of the original workpiece and the annotation data of its key points;
[0015] The pre-trained real-time detection model for workpieces and workpiece key points adopts the YOLOv5 fully convolutional network, including: a Backbone feature extraction network, a Neck layer and a Prediction layer, wherein the Backbone feature extraction network has a CoordAttention attention mechanism.
[0016] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0018] Compared with the prior art, the present application has the following beneficial effects:
[0019] The present application uses a depthwise separable convolution module to design a lightweight feature extraction network, combines this network with a CoordAttention attention mechanism and applies it to YOLOv5. Then, the BiFPN module is fused to improve the accuracy, and at the same time, the task of workpiece key point detection is added. Through experimental analysis, this method effectively improves the detection speed while ensuring the accuracy, realizes real-time inference in the CPU environment, and solves the above problems. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1This is the flowchart of the method of this application.
[0022] Figure 2 This is the schematic diagram of the system of this application.
[0023] Figure 3 This is the overall flowchart of the method of this application.
[0024] Figure 4 This is the structural diagram of the lightweight network model of this application.
[0025] Figure 5 This is the training flowchart of the prediction model of this application. Detailed implementation manners
[0026] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. Usually, the components of the embodiments of this application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0027] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application claimed, but merely represents the selected embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application.
[0028] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0029] In the description of the embodiments of this application, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings or the orientation or positional relationship in which the inventive product is usually placed during use. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.
[0030] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.
[0031] In the description of the embodiments of the present application, it should also be noted that unless otherwise clearly specified and limited, if the terms "set", "install", "connect", and "couple" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0032] The following further describes the present application in detail with reference to the accompanying drawings:
[0033] See Figure 1 , the embodiments of the present application disclose a workpiece and a real-time detection method for workpiece key points, including the following steps:
[0034] S1. Obtain a workpiece image, and obtain the image data of the workpiece and the annotation data of its key points according to the workpiece image.
[0035] S2. Input the image data of the workpiece and the annotation data of its key points into a pre-trained real-time detection model for the workpiece and its key points to obtain the detection results of the workpiece key points.
[0036] It should be noted that the pre-trained real-time detection model for the workpiece and its key points in the present application is trained using the image data of the original workpiece and the annotation data of its key points.
[0037] In practical applications, the pre-trained real-time detection model for the workpiece and its key points adopts the YOLOv5 fully convolutional network, including: the Backbone feature extraction network, the Neck layer, and the Prediction layer. Among them, the Backbone feature extraction network has the CoordAttention attention mechanism.
[0038] In a feasible embodiment of the present application, in step S1, the position of the workpiece and the selected key points in the workpiece are annotated according to the workpiece image to obtain the image data of the workpiece and the annotation data of the key points.
[0039] In another feasible embodiment of the present application, the real-time detection model for the workpiece and its key points is trained according to the following method:
[0040] S2.1 Divide the image data of the workpiece and the annotation data of the key points to obtain a training set and a test set; convert the image data of the workpiece and the annotation data of the key points into YOLO vThe 5 standard format is used to divide the data into a training set and a test set in a ratio of 9:1; Mixup, Cutout, CutMix, and Mosaic are used to augment the image data of the workpiece and the annotation data of the key points.
[0041] S2.2 Construct a lightweight Backbon e feature extraction network and fuse it into YOLO v 5 fully convolutional network;
[0042] S2.3 In YOLO v 5 fully convolutional network's Backbon e feature extraction network to add CoordAttentio n attention mechanism; The Backbon e feature extraction network is a lightweight feature extraction network based on depthwise separable convolution; The module structure of the depthwise separable convolution is as follows:
[0043] After the input features pass through depthwise convolution, BatchNorm processing is performed, then activated using the HardSwish activation function, followed by 1*1 pointwise convolution for channel fusion and channel expansion, and finally processed by the BatchNorm layer and the HardSwish layer;
[0044] The depthwise separable convolution consists of two parts: depthwise convolution and pointwise convolution. Each convolution kernel in the depthwise convolution performs a convolution operation on its corresponding channel; If the input image size is M_IN*N_IN*C_IN and the expected output is a feature map of M_OUT*N _ OUT*C_OUT, when the convolution kernel size is K*K:
[0045] The number of parameters of the depthwise separable convolution NUM_P_DP is as follows:
[0046] NUM_P_DP = K*K*C_IN + C_IN*C_OUT
[0047] where, C_IN represents the number of input channels, and C_OUT represents the number of output channels;
[0048] The computational amount of the depthwise separable convolution NUM_C_DP is as follows:
[0049] NUM_C_DP = K*K*C_IN*M_OUT*N_OUT + K*K*C_IN*C_OUT
[0050] where, M_OUT represents the length of the output image, and N_OUT represents the width of the output image;
[0051] The computational cost comparison between depthwise separable convolution and traditional convolution, NUM_C_T, is as follows:
[0052]
[0053] The computational cost of depthwise separable convolution is 1 / C_OUT of that of traditional convolution.
[0054] S2.4 Adopts an improved BiFPN as the structure of the Neck layer of the YOLOv5 fully convolutional network;
[0055] S2.5 Adds a key point detection task to the YOLOv5 fully convolutional network, and uses the regression branch of the workpiece key point coordinates to calculate the loss of the workpiece key points in combination with the prediction box;
[0056] It should be noted that step S2.5 is specifically as follows:
[0057] S2.5.1 Adds a regression branch of workpiece key point coordinates on the basis of workpiece detection;
[0058] S2.5.3 During the prediction process, the result of each prediction box includes the confidence of the object contained in this box, the confidence that the object contained in this box is a specific workpiece, and the bounding box coordinates (x, y, w, h). The coordinates of n key points total 6 + 2n values;
[0059] S2.5.4 Uses the Wing Loss function to calculate the workpiece key point loss.
[0060] In practical applications, the Wing Loss function uses a piecewise function, as follows:
[0061]
[0062] Among them, Wing(x) represents the WingLoss function, f(x) is used to refer to Wing(x) for convenience in the following text. w ln() represents the weighting of w and ln(), ln() refers to the natural logarithm function with base e, ∈ represents the curvature of the non-linear region, w represents the range of the non-linear part set to (-w, w), C represents the constant that smoothly connects the piecewise-defined linear part and non-linear part, and C = w - w ln(1 + w^491).
[0063] S2.6 Uses the Hardswish activation function and adopts the Cosine Annealing LR learning rate decay strategy for model training.
[0064] As Figure 2 shown, the embodiment of the present application discloses a workpiece and workpiece key point real-time detection system, including:
[0065] An image acquisition module, configured to acquire a workpiece image and obtain image data of the workpiece and annotation data of key points thereof based on the workpiece image;
[0066] A result detection module, configured to input the image data of the workpiece and the annotation data of key points thereof into a pre-trained real-time detection model for workpieces and workpiece key points, and obtain a detection result of workpiece key points;
[0067] Wherein, the pre-trained real-time detection model for workpieces and workpiece key points is obtained by training using image data of original workpieces and annotation data of key points thereof;
[0068] The pre-trained real-time detection model for workpieces and workpiece key points adopts a YOLOv5 fully convolutional network, including: a Backbone feature extraction network, a Neck layer, and a Prediction layer. Among them, the Backbone feature extraction network has a CoordAttention attention mechanism.
[0069] Embodiment
[0070] As Figure 3 shown, an embodiment of the present application provides a real-time detection method for workpieces and workpiece key points based on a lightweight network, including the following steps:
[0071] Step 1: Acquire an original workpiece image, and annotate the workpiece position and selected key points in the workpiece. Convert the image data and annotation data into the YOLOv5 standard format, and divide them into a training set and a test set at a ratio of 9:1; in order to make the method in the present application have better robustness after training, data augmentation methods such as Mixup, Cutout, CutMix, and Mosaic augmentation will be adopted. When performing model training, the size of batchsize is related to the video memory size. For example: batchsize = 1, which means 1 picture. When using Mosaic augmentation, four pictures will be stitched into a large picture for training. In this way, more pictures can be stuffed into the training each time. When the GPU resources are limited, the more pictures processed by one batchsize, the better. Of course, after the four pictures are stitched together, the corresponding processing of the annotation is required. The four pictures are mixed together, increasing the complexity of the image background.
[0072] Step 2: Build a lightweight feature extraction network based on depthwise separable convolution and fuse it into YOLOv5, and at the same time add a CoordAttention attention mechanism to the backbone feature extraction network.
[0073] As Figure 4As shown in the figure, as an example of the detection model of the present application, the detection model includes a StemConv layer and 13 consecutive 3*3 depthwise separable convolution layers. The CoordAttention attention mechanism is introduced in the last two 3*3 depthwise separable convolution layers. In addition, the image sizes output by the 13 3*3 depthwise separable convolution layers decrease sequentially, and the number of channels increases sequentially.
[0074] The present application uses a lightweight feature extraction network as the backbone feature extraction network to further improve the inference speed of the network model and achieve real-time detection. The design of the backbone feature extraction network is related to the speed and accuracy of the subsequent model. To balance speed and accuracy, the present application designs a lightweight backbone feature extraction network based on depthwise separable convolution. The structure of the depthwise separable convolution module is as follows: After the input features are subjected to depthwise convolution, BatchNorm processing is performed, then activated using the HardSwish activation function, and then 1*1 pointwise convolution is used for channel fusion and channel expansion. Finally, it is processed by the BatchNorm layer and the HardSwish layer. In traditional convolution, if the input image size is M_IN*N_IN*C_IN and it is desired to output a feature map of M_OUT*N_OUT*C_OUT, when the convolution kernel size is K*K:
[0075] The formula for calculating the number of parameters of traditional convolution is:
[0076] NUM_P_T = K*K*C_IN*C_OUT
[0077] The formula for calculating the computational complexity of traditional convolution is:
[0078] NUM_C_T = K*K*C_IN*C_OUT*M_OUT*N_OUT
[0079] Depthwise separable convolution can be divided into two parts: depthwise convolution and pointwise convolution. Each convolution kernel in depthwise convolution only performs convolution operations on its corresponding channels, does not perform cross-channel calculations, and keeps the number of channels consistent. Depthwise convolution operations mainly focus on the information within the channels and do not pay attention to the information between channels. If the input image size is M_IN*N_IN*C_IN and it is desired to output a feature map of M_OUT*N_OUT*C_OUT, when the convolution kernel size is K*K:
[0080] The formula for calculating the number of parameters of depthwise separable convolution is:
[0081] NUM_P_DP = K*K*C_IN + C_IN*C_OUT
[0082] The formula for calculating the computational complexity of depthwise separable convolution is:
[0083] NUM_C_DP = K * K * C_IN * M_OUT * N_OUT + K * K * C_IN * C_OUT
[0084] The computational cost of depthwise separable convolution compared to that of traditional convolution is as follows:
[0085]
[0086] From the above analysis, it can be obtained that the computational cost of depthwise separable convolution is approximately 1 / C_OUT of that of traditional convolution, where C_OUT is the number of output channels.
[0087] Common types of attention mechanisms include channel attention mechanisms and spatial attention mechanisms. The channel attention mechanism is represented by SENet and ECANet improved based on SENet. For most current lightweight networks, considering the limitation of computing power, they often carefully choose to use attention mechanisms or choose attention mechanisms with less impact on performance, such as the SE attention mechanism. When using the SE attention mechanism, since it only focuses on the information between channels and abandons the attention to spatial position information, for data such as workpieces where position is very important, the SE attention mechanism has great limitations. To solve the problem that adding an attention mechanism will cause performance loss, in this application, an attention mechanism that simultaneously focuses on channel information and spatial position information and is specially designed for lightweight networks is adopted:
[0088] The CoordAttention attention mechanism. CoordAttention plays a significant role in improving the performance of lightweight network models. It can not only focus on the information between channels but also simultaneously focus on the information in terms of position. Compared with other channel attention mechanisms, CoordAttention aggregates the input in two directions to output feature encoding, so that the information in the two directions can be used to complement each other and enhance the representation of features.
[0089] Step 3: Use the improved BiFPN in the Neck layer. To further improve the effect of feature fusion, this application adopts the improved BiFPN as the main structure of the Neck layer. In BiFPN, the importance of features at different scales is regarded differently from the previous feature fusion methods. BiFPN believes that different feature maps need to be treated differently and different weights are assigned to features at different scales, similar to the attention mechanism, and then added together. Compared with PANet, BiFPN can achieve certain performance improvement while reducing the number of parameters.
[0090] Step 4: Add the key point detection task to the detection network. In this application, a regression branch for the coordinates of the key points of the workpiece will be added on the basis of workpiece detection. During the prediction process, the result of each prediction box needs to include the confidence of the object contained in this box, the confidence that the object contained in this box is a specific workpiece, the bounding box coordinates (x, y, w, h), and the coordinates of n key points, a total of 6 + 2n values. When calculating the loss of the workpiece key points, the Wing Loss function is used. In the workpiece key point detection, the coordinate regression difficulty of each key point is not exactly the same. In the early stage of training, the coordinate errors of each key point are relatively large. In the middle and late stages of training, the errors of each key point have become relatively small, and there may also be a situation where there are relatively more errors in a small number of key points compared to other key points. In order to reduce this small part of the error, it can be amplified. When the training is almost completed, there may still be some points with relatively large losses, and these points are called outliers. At this time, the losses of most key points are already small enough. Therefore, if an ordinary loss function is used, then in one backpropagation process, these outliers will dominate and cause more losses to other relatively accurate key points. Therefore, in order to solve the problem that outliers affect non-outliers, a piecewise function is adopted in Wing Loss, and its expression is:
[0091]
[0092] Step 5: Train the detection model. As Figure 5 shown, in the backbone feature extraction network, the Hardswish activation function is mainly used. Its main advantage is that while improving the accuracy, no additional computational cost is added. Compared with the use of sigmoid, the cost of Hardswish is lower. When calculating the loss of the workpiece key point coordinates, the Wing Loss function is used.
[0093] During the model training process, Cosine Annealing LR is adopted as the learning rate decay strategy for model training. The initial learning rate is set to 0.01, the batchsize is set to 64, and a total of 200 rounds of training are performed.
[0094] Step 6: Deploy the detection model. In this application, the model trained in Step 5 is converted into the ONNX format, and then the model optimization tool provided in OpenVINO is used to optimize and convert the model to obtain a model format suitable for OpenVINO. After implementing the above process, the model can be deployed to the CPU for inference to achieve real-time detection of workpieces and workpiece key points based on the CPU platform.
[0095] A computer device provided by an embodiment of the present application. The computer device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0096] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present application.
[0097] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory.
[0098] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0099] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory.
[0100] If the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0101] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A method for real-time detection of workpieces and key points of workpieces, characterized in that: The following steps are involved: Acquire a workpiece image, and obtain image data of the workpiece and annotation data of key points thereof according to the workpiece image; Input the image data of the workpiece and the annotation data of its key points into the pre-trained real-time detection model of the workpiece and its key points to obtain the detection results of the key points of the workpiece; The pre-trained real-time detection model for workpieces and workpiece key points is obtained by training using the image data of the original workpiece and the annotated data of its key points; The pre-trained workpiece and workpiece key point real-time detection model adopts a YOLOv5 fully convolutional network, including: a Backbone feature extraction network, a Neck layer and a Prediction layer, wherein the Backbone feature extraction network has a CoordAttention mechanism.
2. The real-time detection method for workpieces and workpiece key points according to claim 1, characterized in that: The step of obtaining the image data of the workpiece and the annotation data of its key points according to the workpiece image includes: The workpiece position and selected key points in the workpiece are marked according to the workpiece image to obtain image data of the workpiece and marked data of the key points.
3. The real-time detection method for workpieces and workpiece key points according to claim 1, characterized in that: The workpiece and workpiece key point real-time detection model is trained according to the following method: Divide the image data of the workpiece and the annotation data of the key points to obtain a training set and a test set; Build a lightweight Backbone feature extraction network and integrate it into the YOLOv5 fully convolutional network; Add the CoordAttention mechanism to the Backbone feature extraction network of the YOLOv5 fully convolutional network; The improved BiFPN is used as the structure of the Neck layer of the YOLOv5 fully convolutional network; Add key point detection tasks to the YOLOv5 fully convolutional network, use the regression branch of the workpiece key point coordinates, and combine the prediction box to calculate the loss of the workpiece key point; The Hardswish activation function is used, and the Cosine Annealing LR learning rate decay strategy is adopted to train the model.
4. The real-time detection method for workpieces and workpiece key points according to claim 3 is characterized in that: The image data of the workpiece and the annotation data of the key points are divided to obtain a training set and a test set, including: Convert the image data of the workpiece and the annotation data of key points into the YOLOv5 standard format and divide them into training set and test set with a ratio of 9:1; Mixup, Cutout, CutMix and Mosaic are used to enhance the image data of the workpiece and the annotation data of key points.
5. The real-time detection method for workpieces and workpiece key points according to claim 3, characterized in that: The lightweight Backbone feature extraction network is constructed and integrated into the YOLOv5 full convolutional network, including: The Backbone feature extraction network is a lightweight feature extraction network based on deep separable convolution; the module structure of the deep separable convolution is: After deep convolution, the input features are processed by BatchNorm, activated by HardSwish activation function, and then 1*1 point-by-point convolution is used for channel fusion and channel expansion, and finally processed by BatchNorm layer and HardSwish layer; Depthwise separable convolution consists of depthwise convolution and pointwise convolution. Each convolution kernel in depthwise convolution performs convolution operation on its corresponding channel. If the input image size is M_IN*N_IN*C_IN, and the expected output feature map is M_OUT*N_OUT*C_OUT, when the convolution kernel size is K*K: The number of depth-separable convolution parameters NUM_P_DP is as follows: NUM_P_DP=K*K*C_IN+C_IN*C_OUT Among them, C_IN represents the number of input channels, and C_OUT represents the number of output channels; The depth-separable convolution calculation amount NUM_C_DP is as follows: NUM_C_DP=K*K*C_IN*M_OUT*N_OUT+K*K*C_IN*C_OUT Among them, M_OUT represents the length of the output image, and N_OUT represents the width of the output image; The computational complexity of depthwise separable convolution compared to traditional convolution NUM_C_T is as follows: The computational complexity of depthwise separable convolution is 1 / C_OUT of that of traditional convolution.
6. The real-time detection method for workpieces and workpiece key points according to claim 3, characterized in that: The key point detection task is added to the YOLOv5 full convolutional network, and the regression branch of the workpiece key point coordinates is used to calculate the loss of the workpiece key point in combination with the prediction box, including: Add a regression branch of the workpiece key point coordinates based on workpiece detection; During the prediction process, the result of each prediction box includes the confidence that the box contains the object, the confidence that the box contains the object as a specific workpiece, and the bounding box coordinates (x, y, w, h). The coordinates of the n key points have a total of 6+2n values; The Wing Loss loss function is used to calculate the artifact key point loss.
7. The real-time detection method for workpieces and workpiece key points according to claim 6, characterized in that: The WingLoss loss function is used to calculate the workpiece key point loss, including: The Wing Loss loss function uses a piecewise function as follows: Among them, Wing(x) represents the WingLoss loss function, f(x) refers to Wing(x) for the convenience of reference below, wln() represents the weighted sum of w and ln(), ln() refers to the logarithmic function with base e, ∈ represents the curvature of the nonlinear region, w represents the range of the nonlinear part is set to (-w, w), C represents the constant for smoothly connecting the linear part and the nonlinear part defined by the piecewise connection, C = w-wln(1+w∧491).
8. A real-time detection system for workpieces and key points of workpieces, characterized in that: include: An image acquisition module is used to acquire a workpiece image and obtain image data of the workpiece and annotation data of key points thereof according to the workpiece image; A result detection module is used to input the image data of the workpiece and the annotation data of its key points into a pre-trained real-time detection model of the workpiece and its key points to obtain the detection results of the workpiece key points; The pre-trained real-time detection model for workpieces and workpiece key points is obtained by training using the image data of the original workpiece and the annotated data of its key points; The pre-trained workpiece and workpiece key point real-time detection model adopts a YOLOv5 fully convolutional network, including: a Backbone feature extraction network, a Neck layer and a Prediction layer, wherein the Backbone feature extraction network has a CoordAttention mechanism.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.