Method, device and medium for detecting dimensions of forgings in hot forming process

By introducing deformable convolution and attention mechanisms into the forging detection model and using structured interleaving and ratio loss function, the problem that convolutional neural network cannot measure the tiny deformation of forgings online in real time is solved, and efficient and accurate forging size detection is achieved.

CN119919474BActive Publication Date: 2025-07-01WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510414401.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-01
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing convolutional neural network cannot measure the slight deformation of forgings online in real time during the thermoforming process, resulting in the forgings not meeting the size of the forgings, which in turn leads to the scrapping of forgings.

Method used

Using the neural network model based on the YOLOv5 framework, the adaptability to the shape characteristics of the forging and attention to slight deformation is enhanced by introducing deformable convolution and channel- and spatial attention mechanisms in the backbone network layer and the neck network layer, and using a structured interleaving and ratio loss function.

Benefits of technology

Real-time online measurement of forging dimensions is realized, the accuracy and stability of detection is improved, the scrapping of forgings is avoided, and the production requirements of high efficiency and high precision are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919474B_ABST
    Figure CN119919474B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device and medium for detecting the size of forgings during the hot forming process, which relates to the field of computer vision technology. The method includes: obtaining image data of a forging blank to be detected; inputting the image data into a trained forging size detection model to obtain size information and classification results; the forging size detection model includes a backbone network layer, a neck network layer and a head network layer. There are several cross-stage local network layers and several generalized sparse convolution layers in both the backbone network layer and the neck network layer. The bottleneck layer in all cross-stage local network layers uses deformable convolution and adds channel and spatial attention mechanisms. The present invention solves the problems that the current size detection of forgings during the hot forming process relies on manual operation, resulting in lack of accuracy and stability, and the problem that the forming size of forgings fails to meet the standards due to insufficient deformation accuracy of the blank during the hot forming process, leading to the scrapping of forgings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method, device and medium for detecting the size of forgings during the hot forming process. Background Art

[0002] The application of advanced scientific and technological applications such as sensor technology, computer technology, and intelligent control technology to the manufacturing industry has promoted the modern industry towards intelligence, informatization, and scale. However, at present, the size of forgings during the hot forming process still depends on manual experience for detection. That is to say, the determination of the forging size largely depends on the personal experience and visual judgment of the inspectors. Usually, the inspector needs to rely on experience to judge whether the forging has reached the expected size and then use a ruler for actual measurement. This method not only consumes a large amount of time and human resources and is difficult to ensure the high efficiency of production, but also there are deviations in the forging size, which will cause uneven internal organization of the forging and affect its mechanical properties.

[0003] Machine vision technology plays an important role in the field of target detection with its high efficiency and accuracy. Machine vision technology can automatically detect and identify targets, which not only improves the production efficiency and detection accuracy, but also helps inspectors get rid of repetitive and time-consuming tasks. The commonly used Convolutional Neural Network (CNN) in machine vision technology, as a strong target detection algorithm, can replace the feature extractor and classifier in the traditional feature learning theory, enabling the model to complete the target detection process quickly and simply without complex preprocessing of the original image. Due to the advantages of automatic feature extraction and strong generalization ability of the convolutional neural network, the convolutional neural network is also increasingly applied to the forging size detection scenario.

[0004] During the hot forming process, the blank will deform. If the deformation accuracy of the blank is not enough, it will lead to the non-compliance of the forming size of the forging, and then the forging will be scrapped. However, the current CNN model can only focus on the large deformations generated by the blank and cannot focus on the small deformations generated on the blank, and cannot fully meet the requirements of real-time online measurement of the forging blank size.

[0005] During the forging process of the forging, if the problem of insufficient deformation accuracy can be detected in time, remedial measures can be taken in time. Therefore, it is necessary to provide a method for detecting the size of forgings during the hot forming process that can solve the above problems. Summary of the Invention

[0006] In view of this, the present invention proposes a method, device and medium for detecting the size of forgings during the hot forming process to solve the problem that the current machine vision technology using convolutional neural network cannot perform real-time online measurement of forgings.

[0007] The technical solution of the present invention is realized as follows:

[0008] According to a first aspect, an embodiment of the present invention provides a method for detecting the size of a forging during a hot forming process, the method comprising:

[0009] Obtaining image data of a forging blank to be detected;

[0010] Inputting the image data into a trained forging size detection model to obtain the size information and classification result of the forging blank to be detected output by the forging size detection model;

[0011] The forging size detection model is a neural network model based on the YOLOv5 framework. The forging size detection model includes a backbone network layer, a neck network layer, and a head network layer. There are several cross-stage partial network layers and several generalized sparse convolutional layers in both the backbone network layer and the neck network layer. Moreover, deformable convolutions are used in the bottleneck layers of all cross-stage partial network layers, and the number of channels and the spatial attention mechanism are increased. The structured intersection over union loss function is used to determine the penalty exponent and use the penalty exponent to determine the vector angle between the required bounding box regressions;

[0012] The cross-stage partial network layer includes a first convolutional layer, a bottleneck layer, a first tensor splicing layer, a second convolutional layer, and a third convolutional layer connected in sequence;

[0013] The input ends of the first convolutional layer and the second convolutional layer are both used to receive the input data of the cross-stage partial network layer. The output end of the second convolutional layer is connected to the input end of the first tensor splicing layer. The output end of the first tensor splicing layer is connected to the input end of the third convolutional layer. The third convolutional layer is used to output the output data of the cross-stage partial network layer.

[0014] Combined with the first aspect, in the first implementation manner of the first aspect, the step of inputting the image data into a trained forging size detection model to obtain the size information and classification result of the forging blank to be detected output by the forging size detection model specifically includes:

[0015] Inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain a forging blank fusion feature map output by the backbone network layer;

[0016] Inputting the forging blank fusion feature map into the neck network layer for feature fusion and feature enhancement processing of feature maps at different levels to obtain three-scale forging blank scale feature maps output by the neck network layer;

[0017] Input all the forging blank dimension feature maps into the head network layer for target detection processing to obtain the size information and classification results of the forging blanks to be detected output by the head network layer.

[0018] Combined with the first implementation manner of the first aspect, in the second implementation manner of the first aspect, the backbone network layer includes a focus feature fusion layer, a depth semantic extraction layer, and a fast spatial pyramid pooling layer connected in sequence. The depth semantic extraction layer includes a number of generalized sparse convolution layers and cross-stage local network layers arranged between two adjacent generalized sparse convolution layers. The first and the last layers of the depth semantic extraction layer are both generalized sparse convolution layers. The generalized sparse convolution layers are arranged at intervals, and two adjacent generalized sparse convolution layers are connected by a cross-stage local network layer.

[0019] Combined with the second implementation manner of the first aspect, in the third implementation manner of the first aspect, inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain the forging blank fusion feature map output by the backbone network layer specifically includes:

[0020] Input the image data into the focus feature fusion layer for channel expansion and downsampling processing to obtain the downsampled feature map output by the focus feature fusion layer;

[0021] Input the downsampled feature map into the depth semantic extraction layer to extract the depth semantic information in the feature map to obtain the scale feature map output by the depth semantic extraction layer;

[0022] Input the scale feature map into the fast spatial pyramid pooling layer for scale fusion processing to obtain the forging blank fusion feature map.

[0023] Combined with the third implementation manner of the first aspect, in the fourth implementation manner of the first aspect, inputting the downsampled feature map into the depth semantic extraction layer to extract the depth semantic information in the feature map to obtain the scale feature map output by the depth semantic extraction layer specifically includes:

[0024] Input the downsampled feature map into the first generalized sparse convolution layer at the beginning of the depth semantic extraction layer for generalized sparse convolution processing to obtain the first convolution feature output by the first generalized sparse convolution layer at the beginning;

[0025] Input the first convolution feature into the cross-stage local network layer of the next layer connected to the generalized sparse convolution layer to extract multi-scale features to obtain the second convolution feature output by the cross-stage local network layer;

[0026] Input the second convolutional feature into the generalized sparse convolutional layer of the next layer connected by the cross-stage local network layer for generalized sparse convolutional processing to obtain a third convolutional feature output by the generalized sparse convolutional layer, and continue this process until the third convolutional feature is input into the last generalized sparse convolutional layer to obtain a scale feature map output by the last generalized sparse convolutional layer.

[0027] Combined with the first aspect, in the fifth implementation manner of the first aspect, the bottleneck layer includes an identity mapping layer, and a fourth convolutional layer, a deformable convolutional layer, a channel attention mechanism layer, a spatial attention mechanism layer, and a feature addition layer connected in sequence;

[0028] The input ends of the fourth convolutional layer and the identity mapping layer are both used to receive the input data of the bottleneck layer. The output end of the identity mapping layer is connected to the output end of the feature addition layer, and the feature addition layer is used to output the output data of the bottleneck layer.

[0029] Combined with the first aspect, in the sixth implementation manner of the first aspect, the generalized sparse convolutional layer includes a fifth convolutional layer, a depthwise separable convolutional layer, a second tensor splicing layer, and a shuffle layer;

[0030] The input end of the fifth convolutional layer is used to receive the input data of the generalized sparse convolutional layer. The input end of the depthwise separable convolutional layer is connected to the output end of the fifth convolutional layer, and the output ends of the depthwise separable convolutional layer and the fifth convolutional layer are both connected to the input end of the second tensor splicing layer. The output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolutional layer.

[0031] Combined with the first aspect, in the seventh implementation manner of the first aspect, the forging size detection model is trained through the following steps:

[0032] Obtain the original image data of the sample forging, and preprocess the original image data to obtain the sample image data of the sample forging;

[0033] Perform target annotation on the sample image data to obtain annotation information; the annotation information includes the size bounding box and class label of the sample forging;

[0034] Use the sample image data as the input data for training, use the annotation information as the label during training, and perform training using machine learning and adjust the network weights by means of gradient descent to obtain the forging size detection model for predicting the size information and classification result of the forging blank to be detected.

[0035] According to a second aspect, an embodiment of the present invention further provides a detection device for the size of a forging during a hot forming process based on machine vision. The device includes:

[0036] A data acquisition module for acquiring image data of a forging blank to be detected;

[0037] A size detection module for inputting the image data into a trained forging size detection model to obtain the size information and classification result of the forging blank to be detected output by the forging size detection model;

[0038] The forging size detection model is a neural network model based on the YOLOv5 framework. The forging size detection model includes a backbone network layer, a neck network layer, and a head network layer. There are several cross-stage local network layers and several generalized sparse convolution layers in both the backbone network layer and the neck network layer. Moreover, deformable convolutions are used in the bottleneck layers of all cross-stage local network layers, and the number of channels and spatial attention mechanisms are increased. The structured intersection over union loss function is used to determine the penalty exponent and use the penalty exponent to determine the vector angle between the required bounding box regressions;

[0039] The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor splicing layer, a second convolutional layer, and a third convolutional layer connected in sequence;

[0040] The input ends of the first convolutional layer and the second convolutional layer are used to receive the input data of the cross-stage local network layer. The output end of the second convolutional layer is connected to the input end of the first tensor splicing layer. The output end of the first tensor splicing layer is connected to the input end of the third convolutional layer. The third convolutional layer is used to output the output data of the cross-stage local network layer.

[0041] According to a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the detection method for the size of a forging during a hot forming process as described in any one of the above are implemented.

[0042] According to a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the detection method for the size of a forging during a hot forming process as described in any one of the above are implemented.

[0043] The detection method, device, and medium for the size of a forging during a hot forming process of the present invention have the following beneficial effects compared with the prior art:

[0044] The obtained image data of the forging blank to be detected is input into the trained forging size detection model, and the size information and classification result of the forging blank to be detected output by the forging size detection model are obtained. Among them, the forging size detection model is an improvement on the traditional YOLOv5s model. The convolutional layer in the bottleneck layer of the C3 layer in the traditional YOLOv5s model is replaced with a deformable convolutional layer, which allows the convolutional kernel to dynamically learn the offset amount, so as to adapt to the shape change of the forging during the hot forming process, thereby enhancing the adaptability of the feature map when the shape characteristics of the forging blank change greatly, and making the forging size model pay more attention to the target area, reducing the influence of noise data on the model, improving the perception ability of the model, and introducing a channel and spatial attention mechanism in the C3 layer to enhance the attention ability to the minute deformations generated at each moment during the hot forming process of the forging blank. Secondly, all convolutional layers in the backbone network layer and the neck network layer are generalized sparse convolutional layers, which reduce the computational complexity and the number of parameters of the model while maintaining the model performance, and use the structured intersection over union loss function as the bounding box loss function to improve the convergence speed of the model during training and improve the accuracy of bounding box regression. The size information and classification result finally predicted by the forging size detection model are more accurate, and because it can pay attention to various types of deformation characteristics of the forging blank, it solves the problems of lack of accuracy and stability caused by relying on manual operation in the current size detection of forgings during the hot forming process, and the problem that the forming size of forgings does not meet the standard due to insufficient blank deformation accuracy during the hot forming process, resulting in the scrapping of forgings, meeting the requirements of real-time online measurement of the forging blank size, enabling the production of parts to meet the production requirements of high efficiency and high precision, improving production efficiency and product quality, reducing production costs and resource waste, and having significant economic and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0046] Figure 1 Shows a schematic flow chart of the method for detecting the size of forgings during the hot forming process of the present invention;

[0047] Figure 2 Shows a schematic structural diagram of the existing YOLOv5s model;

[0048] Figure 3 Shows a schematic structural diagram of the model in the method for detecting the size of forgings during the hot forming process of the present invention;

[0049] Figure 4 Shows a schematic structural diagram of an improved cross-stage local network layer in the method for detecting the dimensions of forgings in the hot forming process of the present invention;

[0050] Figure 5 Shows a schematic structural diagram of an improved bottleneck layer in the method for detecting the dimensions of forgings in the hot forming process of the present invention;

[0051] Figure 6 Shows a schematic structural diagram of a generalized sparse convolutional layer in the method for detecting the dimensions of forgings in the hot forming process of the present invention;

[0052] Figure 7 Shows a schematic structural diagram of a device for detecting the dimensions of forgings in the hot forming process based on machine vision provided by the present invention;

[0053] Figure 8 Shows a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0054] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0055] The application of advanced scientific and technological applications such as sensor technology, computer technology, and intelligent control technology to the manufacturing industry has promoted the modern industry to move towards intelligence, informatization, and scale. However, at present, the dimensions of forgings in the hot forming process still rely on manual experience for detection. That is to say, the determination of the dimensions of forgings largely depends on the personal experience and visual judgment of the detection personnel. Usually, the detection personnel need to rely on experience to judge whether the forgings have reached the expected dimensions, and then use a ruler for actual measurement. This method not only consumes a large amount of time and human resources and is difficult to ensure the high efficiency of production, but also there are deviations in the dimensions of the forgings, which will cause uneven internal organization of the forgings and affect their mechanical properties.

[0056] Machine vision technology plays an important role in the field of target detection with its high efficiency and accuracy. Machine vision technology can automatically detect and identify targets, which not only improves production efficiency and detection accuracy, but also helps detection personnel get rid of repetitive and time-consuming tasks. As a strong target detection algorithm, CNN commonly used in machine vision technology can replace the feature extractor and classifier in traditional feature learning theory, enabling the model to complete the target detection process quickly and simply without complex preprocessing of the original image. Due to the advantages of automatic feature extraction and strong generalization ability of convolutional neural networks, convolutional neural networks are also increasingly applied to the scenario of forging size detection.

[0057] During the hot forming process, the blank will deform. If the deformation accuracy of the blank is insufficient, it will lead to unqualified forming dimensions of the forging, and then cause the forging to be scrapped. However, the current CNN model can only focus on the large deformations generated by the blank and cannot focus on the small deformations on the blank, which cannot fully meet the requirements of real-time online measurement of forging blank dimensions.

[0058] During the forging process of the forging, if the problem of insufficient deformation accuracy can be detected in time, remedial measures can be taken in time. Therefore, it is necessary to provide a method for detecting the size of forgings during the hot forming process that can solve the above problems.

[0059] The method, device and medium for detecting the size of forgings during the hot forming process provided in this specification aim to solve the problems of lack of accuracy and stability caused by relying on manual operation in the size detection of forgings during the hot forming process, and the problem of forging scrap caused by unqualified forming dimensions of forgings due to insufficient deformation accuracy of the blank during the hot forming process.

[0060] The method for detecting the size of forgings during the hot forming process provided in this specification can be applied to electronic devices with computer vision capabilities. The electronic device can include laptops, desktop computers, smartphones, smart wearable devices (virtual reality glasses, smart watches, etc.), tablet computers, etc. Of course, the method for detecting the size of forgings during the hot forming process provided in this specification can also be applied within an application program running on the above-mentioned electronic devices. For example, the method for detecting the size of forgings during the hot forming process can be applied to a browser with computer vision capabilities, or can also be applied within an instant software with computer vision capabilities.

[0061] Please refer to Figure 1 , Figure 1 which shows the flowchart of the method for detecting the size of forgings during the hot forming process according to an embodiment of the present invention. The method may include the following steps:

[0062] S101. Obtain the image data of the forging blank to be detected.

[0063] In this embodiment, the above image data can be stored in the electronic device in advance, or can be obtained by the electronic device from the outside world. For example, by arranging industrial cameras at the production site, the image data of the forging blank to be detected during the production process can be captured. There is no limitation on the specific form of obtaining the image data here, as long as it is ensured that the electronic device can obtain the image data.

[0064] It should be noted that the forgings obtained through the entire forging process can also be used as the forging blanks to be detected. By detecting the dimensions of the formed forgings, it can be determined whether the forgings meet the actual forging standards.

[0065] S102. Input the image data into the trained forging size detection model to obtain the size information and classification result of the forging blank to be detected output by the forging size detection model.

[0066] In this process, the image data serves as the input data of the forging size detection model. The forging size detection model processes the image data and outputs the size information and classification result of the forging.

[0067] In this embodiment, the forging size detection model is a neural network model based on the YOLOv5 framework.

[0068] The YOLO algorithm is a single-stage object detection algorithm based on deep learning. Its characteristics are fast detection speed and relatively high accuracy, and it can achieve real-time object detection. The core idea of the YOLO algorithm is to use the entire image as the input of the neural network, divide the image into multiple grids, and each grid is responsible for predicting the position and category of a target, thus transforming the object detection into a regression problem. It has the advantages of a small model, fast speed, and relatively high accuracy, and is suitable for scenarios such as forging size detection.

[0069] Specifically, the core part of the YOLO algorithm includes a backbone network layer, a neck network layer, and a head network layer. Among them, the backbone network layer is mainly responsible for feature extraction, extracting the object information in the image through a convolutional network. The neck network layer is responsible for multi-scale feature fusion of the feature map, and the head network layer performs the final regression prediction.

[0070] Please refer to Figure 2, YOLOv5s is the smallest model in the YOLOv5 series of the YOLOv algorithm, featuring high speed and accuracy. YOLOv5s is implemented using the PyTorch framework, and pre-trained models can be loaded through PyTorch Hub or trained on custom datasets. The backbone network layer of traditional YOLOv5s consists of a Focus feature fusion layer and a Cross Stage Partial Network (CSPNet). The Focus feature fusion layer is responsible for downsampling the input image and expanding the channels, while the Cross Stage Partial Network layer is responsible for extracting multi-scale features.

[0071] More specifically, the Cross Stage Partial Network layer, i.e., the C3 layer, is an important feature extraction module in the YOLOv5 algorithm. Its main feature is to extract multi-scale features by stacking multiple bottleneck layers.

[0072] Please refer to Figure 3 , similarly, the trained forging size detection model also includes a backbone network layer, a neck network layer, and a head network layer. There are several Cross Stage Partial Network layers and several convolutional layers in both the backbone network layer and the neck network layer. The forging size detection model has improved the Cross Stage Partial Network layer by using deformable convolution (Deformable Convolution Network, DCN) in all bottleneck layers of the Cross Stage Partial Network layers and adding channel and spatial attention mechanisms. Moreover, all convolutional layers in the backbone network layer and the neck network layer are generalized sparse convolutional layers, that is, the convolution of all bottleneck layers in all Cross Stage Partial Network layers in traditional YOLOv5s is replaced with deformable convolution and channel and spatial attention mechanisms are added, and all convolutional layers are replaced with generalized sparse convolutional layers. The structured intersection over union loss function is used as the bounding box loss function in the head network layer of the forging size detection model to obtain the intersection over union situation between the predicted box and the ground truth box.

[0073] In this embodiment, the forging size detection model allows the convolutional kernel to dynamically learn the offset by replacing the standard convolution (Conv) with deformable convolution, thus adapting to the shape changes of the forging.

[0074] In this embodiment, a channel and spatial attention mechanism (Squeeze-and-Excitation, SE) is also introduced into the improved cross-stage local network layer. The attention mechanism can enhance the sensitivity and importance of the convolutional neural network to different channel features, increase the sensitivity of the convolutional neural network to minute differences, and optimize the spatial dimension information to improve the performance of the convolutional neural network. In this way, the improved cross-stage local network layer not only enhances the adaptability of the feature map when the shape features of the forging billet change significantly through DCN, but also makes the entire forging size detection model pay more attention to the target area, reduce the influence of noise, and improve the perception ability of the forging size detection model. Moreover, by adding a channel and spatial attention mechanism, it enhances the ability to pay attention to the minute deformations generated at each moment during the hot forming process of the forging billet, so that the finally trained forging size detection model can meet the requirements of real-time online measurement of the forging billet size. Specifically, the channel and spatial attention mechanism includes a sequentially connected channel attention mechanism layer and a spatial attention mechanism layer. The attention mechanism mainly focuses on the feature fusion between channels in the convolutional operation of the backbone network, and then learns the weight coefficient vector corresponding to each channel in the feature map, and then uses this vector to weight the feature map. Among them, channel attention processing can identify which channels contain more important information and give these channels higher weights, so as to solve the problem of conflicting information between different scale features. Then, spatial attention processing can identify the important regions in the image data and give higher weights to the features of these regions, so as to solve the problem of lack of context information. Considering the improvement of the performance and training speed of the forging size detection model, in this embodiment, the original bounding box loss function of the traditional YOLOv5s is replaced with a structured intersection over union loss function (SIoU-Loss), which further improves the accuracy and training speed of the forging size detection model. The penalty exponent is redefined by SIoU to consider the vector angle between the required regressions. The structured intersection over union loss function effectively reduces the number of degrees of freedom by adding an angle penalty term, uses the structured intersection over union loss function as the bounding box loss function to improve the convergence speed of the model during training, and improves the accuracy of bounding box regression.

[0075] Please refer to Figure 4 , the cross-stage local network layer in the forging size detection model includes a first convolutional layer, a bottleneck layer, a first tensor splicing layer, a second convolutional layer, and a third convolutional layer that are sequentially connected in order. The input ends of the first convolutional layer and the second convolutional layer are both used to receive the input data of the cross-stage local network layer. The output end of the second convolutional layer is connected to the input end of the first tensor splicing layer. The output end of the first tensor splicing layer is connected to the input end of the third convolutional layer. The third convolutional layer is used to output the output data of the cross-stage local network layer.

[0076] Specifically, the cross-stage partial network layer has two branches: the first branch is the second convolutional layer for performing convolutional processing on the input data; the second branch is the first convolutional layer, bottleneck layer, first tensor splicing layer, and third convolutional layer connected in sequence. After the input data is input into this branch, feature extraction is performed by the first convolutional layer, and then multi-scale features are extracted in the bottleneck layer. Finally, the features obtained from the bottleneck layer and the first convolutional layer are feature-spliced in the first tensor splicing layer, and then the spliced features are input into the third convolutional layer for further convolutional processing to obtain the output data of the cross-stage partial network layer.

[0077] Since deformable convolution requires additional learning of offset amounts during the convolution process, a branch network is needed to predict the offset parameters. Please refer to Figure 5 , in this embodiment, the bottleneck layer in the forging size detection model adopts a residual network structure. The bottleneck layer includes an identity mapping layer for skip connection, and a fourth convolutional layer, deformable convolutional layer, channel attention mechanism layer, spatial attention mechanism layer, and feature addition layer, that is, a feature fusion layer, connected in sequence. The input ends of the fourth convolutional layer and the identity mapping layer are both used to receive the input data of the bottleneck layer. The output end of the identity mapping layer is connected to the output end of the feature addition layer, and the feature addition layer is used to output the output data of the bottleneck layer.

[0078] Specifically, the bottleneck layer has two branches: the first branch is the identity mapping layer for performing identity mapping processing on the input data, and the identity mapping layer is used to output identity mapping features; the second branch is the fourth convolutional layer, deformable convolutional layer, channel attention mechanism layer, spatial attention mechanism layer, and feature addition layer connected in sequence. After the input data is input into this branch, feature extraction is performed by the fourth convolutional layer and the deformable convolutional layer, and then the extracted features are sequentially input into the channel attention mechanism layer and the spatial attention mechanism layer for corresponding attention mechanism processing. Finally, the features obtained from the two branches are feature-fused at the feature addition layer to obtain the output data of the bottleneck layer.

[0079] Please refer to Figure 6 , the generalized sparse convolutional layer in the forging size detection model includes a fifth convolutional layer, depthwise separable convolutional layer, second tensor splicing layer, and shuffle layer. The input end of the fifth convolutional layer is used to receive the input data of the generalized sparse convolutional layer. The input end of the depthwise separable convolutional layer is connected to the output end of the fifth convolutional layer, and the output ends of the depthwise separable convolutional layer and the fifth convolutional layer are both connected to the input end of the second tensor splicing layer. The output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolutional layer.

[0080] Specifically, the input data of the generalized sparse convolution layer is first input into the fifth convolution layer for convolution processing, and then lightweight convolution operations such as depthwise separable convolution are performed on the output data of the fifth convolution layer. After that, the output data of the fifth convolution layer and the depthwise separable convolution layer are feature concatenated in the second tensor concatenation layer. Then, the concatenated features are input into the shuffle layer for shuffling processing to obtain the output data of the generalized sparse convolution layer.

[0081] In this embodiment, the standard convolution in the original traditional YOLOv5s is replaced by a generalized sparse convolution layer. The generalized sparse convolution is a lightweight convolution operation that simulates the effect of the standard convolution operation by generating a small number of redundant features, thereby reducing the computational cost. The number of output channels of the generalized sparse convolution is the same as that of the original standard convolution to ensure compatibility with other modules.

[0082] Preferably, the specific configuration of the generalized sparse convolution layer (such as the number of groups, kernel size, etc.) can be continuously adjusted during the model training process to achieve a balance between computational overhead and accuracy.

[0083] Correspondingly, the forging size detection model is trained through the following steps:

[0084] Obtain the original image data of the sample forging, and preprocess the original image data to obtain the sample image data of the sample forging. The preprocessing includes image enhancement processing and image denoising processing on the original image data.

[0085] Since the amount of sample data in the abnormal state is small, it is necessary to perform image enhancement and image denoising on the collected original image data. Among them, image enhancement can include but is not limited to: changing brightness and contrast, adding shadows and fog. By preprocessing the original image data, the data set is expanded and the robustness of the model is improved.

[0086] After that, target annotation is performed on the sample image data to obtain annotation information; the annotation information includes the size bounding box and class label of the sample forging.

[0087] In this embodiment, LabelImg is used to perform target annotation on the state of the sample forging in the sample image data.

[0088] During training, the sample image data is used as the input data for training, the annotation information is used as the label during training, and machine learning is used for training and the network weights are adjusted by gradient descent to obtain a forging size detection model for predicting the size information and classification result of the forging blank to be detected.

[0089] Deploy the trained real-time detection model to a mobile device, and by obtaining the image data of the forging billet to be detected in real time during the forging process, use the optimized weight parameters to quickly detect the sizes of forgings in different states and output the detection results.

[0090] In this embodiment, in order to update the network parameters of the forging size detection model, the training of the forging size detection model uses the gradient descent algorithm to continuously optimize the network weights. For example, the mini-batch gradient descent method is adopted, that is, a part of the data is used each time to calculate the gradient, and the historical gradient direction and the current gradient direction can be combined when updating the network parameters. At the same time, a data augmentation strategy can be introduced during the training process to ensure the generalization ability of the forging size detection model for different forging states.

[0091] To improve the computing speed and memory utilization rate of the Graphics Processing Unit (GPU), the forging size detection model training adopts a mixed-precision training method. Through mixed-precision training, the training efficiency can be improved and the batch size can be increased without reducing the accuracy of the forging size detection model.

[0092] In this embodiment, in order to evaluate the performance of the improved model, three evaluation indicators, namely accuracy, recall rate, and mean average precision, are selected to evaluate the detection effect of the model, and then judge the actual prediction ability of the improved YOLOv5s network. The corresponding evaluation indicators include:

[0093] (1)

[0094] In formula (1), represents accuracy; represents the number of positive samples correctly determined; represents the number of positive samples misjudged as positive (should actually be negative).

[0095] (2)

[0096] In formula (2), represents recall; represents the number of negative samples misjudged as negative (should actually be positive).

[0097] (3)

[0098] In formula (3), represents mean average precision, reflects the detection accuracy of the model, and the higher the value, the better the model accuracy; Indicates the category of the detected sample. In this embodiment, ; Indicates the average precision of each category.

[0099] The processing process of the trained forging size model for the image data of the forging blank to be detected is as follows:

[0100] The input image data of the forging blank to be detected undergoes feature extraction and fusion processing through the backbone network layer to obtain the forging blank fusion feature map output by the backbone network layer. Then, the forging blank fusion feature map is input into the neck network layer for feature fusion and feature enhancement processing of feature maps at different levels, further extracting features to obtain three-scale forging blank scale feature maps output by the neck network layer. Finally, all the forging blank scale feature maps are input into the head network layer for object detection processing to obtain the size information and classification result of the forging blank to be detected output by the head network layer.

[0101] Take Figure 3 as an example for illustration. The backbone network layer includes a focus feature fusion layer, a depth semantic extraction layer, and a fast spatial pyramid pooling layer connected in sequence. The depth semantic extraction layer includes a number of generalized sparse convolutional layers and a cross-stage local network layer arranged between two adjacent generalized sparse convolutional layers. The first and the last layers of the depth semantic extraction layer are both generalized sparse convolutional layers. The generalized sparse convolutional layers are arranged at intervals, and two adjacent generalized sparse convolutional layers are connected by a cross-stage local network layer.

[0102] Preferably, the deep semantic extraction layer includes a total of 4 layers of generalized sparse convolutional layers and 3 layers of cross-stage local network layers arranged between these 4 layers of generalized sparse convolutional layers. Among them, the backbone network performs slicing operations on the image data through the focus feature fusion module, expands the input channels to 4 times the original, and obtains the downsampled feature map through one convolution. The feature map is further block-extracted for the deep semantic information of the image and the computational amount is reduced through 4 generalized sparse convolutions and 3 improved C3 layers, and the scale feature map output by the deep semantic extraction layer is obtained. Finally, through the fast spatial pyramid pooling module, the feature maps of different scales are fused into a feature map of a unified scale to obtain the forging blank fusion feature map. More specifically, the downsampled feature map is input into the first generalized sparse convolutional layer at the beginning of the deep semantic extraction layer for generalized sparse convolutional processing to obtain the first convolutional feature output by the first generalized sparse convolutional layer. Then, the first convolutional feature is input into the cross-stage local network layer of the next layer connected to the generalized sparse convolutional layer to extract multi-scale features, and the second convolutional feature output by the cross-stage local network layer is obtained. Then, the second convolutional feature is input into the generalized sparse convolutional layer of the next layer connected to the cross-stage local network layer for generalized sparse convolutional processing to obtain the third convolutional feature output by the generalized sparse convolutional layer, until the third convolutional feature is input into the last generalized sparse convolutional layer to obtain the scale feature map output by the last generalized sparse convolutional layer. This process improves the robustness of the network. After that, the feature map output by the backbone network layer is input into the neck network layer, and the neck network layer includes a feature pyramid network layer and a path aggregation network layer. The feature pyramid network layer downsamples the high-level feature map output by the improved C3 layer and fuses it with the low-level feature map in a top-down manner. After that, the path aggregation network layer further enhances the feature fusion from bottom to top, and transfers the strong localization information in the feature map output by the low-level improved C3 layer to the high-level feature map neck network. By performing split fusion on multiple feature maps output by the backbone network layer, that is, obtaining three fusion features of different scales through three fusion routes respectively. Then, the neck network layer inputs the three feature maps of different scales output to the head network layer, and the head network layer performs image target detection on them respectively, and obtains the intersection-over-union situation between the predicted box and the ground truth box according to the structured intersection-over-union loss function.

[0103] The detection method for the forging size in the hot forming process of the present invention inputs the acquired image data of the forging blank to be detected into the trained forging size detection model, and obtains the size information and classification result of the forging blank to be detected output by the forging size detection model. Among them, the forging size detection model is an improvement on the traditional YOLOv5s model. The convolutional layer in the bottleneck layer of the C3 layer in the traditional YOLOv5s model is replaced with a deformable convolutional layer, which allows the convolutional kernel to dynamically learn the offset amount, so as to adapt to the shape change of the forging in the hot forming process, thereby enhancing the adaptability of the feature map when the shape characteristics of the forging blank change greatly, and making the forging size model pay more attention to the target area, reducing the influence of noise data on the model, improving the perception ability of the model, and introducing a channel and spatial attention mechanism in the C3 layer to enhance the attention ability to the minute deformations generated at each moment during the hot forming process of the forging blank. Secondly, all convolutional layers in the backbone network layer and the neck network layer are generalized sparse convolutional layers, which reduce the computational complexity and the number of parameters of the model while maintaining the model performance, and use the structured intersection over union loss function as the bounding box loss function to improve the convergence speed of the model during training and improve the accuracy of bounding box regression. The size information and classification result finally predicted by the forging size detection model are more accurate, and because it can pay attention to various types of deformation characteristics of the forging blank, it solves the problems of lack of accuracy and stability caused by relying on manual operation in the current size detection of forgings during the hot forming process, and the problem that the forming size of the forging does not meet the standard due to insufficient precision of the blank deformation during the hot forming process, resulting in the scrapping of the forging, meeting the requirements for real-time on-line measurement of the forging blank size, enabling the production of components to meet the requirements of high efficiency and high precision, improving production efficiency and product quality, reducing production costs and resource waste, and having significant economic and social benefits.

[0104] The device provided by the embodiment of the present invention will be described below. The device described below can be correspondingly referred to the method described above.

[0105] Please refer to Figure 7 , Figure 7 which shows the structural schematic diagram of the detection method for the forging size in the hot forming process of the embodiment of the present invention. The device may include:

[0106] A data acquisition module 10, configured to acquire image data of a forging blank to be detected.

[0107] In this embodiment, the above-mentioned image data may be stored in the electronic device in advance, or may be acquired by the electronic device from the outside. For example, by arranging industrial cameras at the production site, the image data of the forging blank to be detected during the production process is captured. No specific limitation is made on the specific acquisition form of the image data here, as long as it is ensured that the electronic device can acquire the image data.

[0108] The dimension detection module 20 is used to input the image data into the trained forging dimension detection model to obtain the dimension information and classification result of the to-be-detected forging blank output by the forging dimension detection model.

[0109] In this process, the image data serves as the input data of the forging dimension detection model. The forging dimension detection model processes the image data and outputs the dimension information and classification result of the forging.

[0110] In this embodiment, the forging dimension detection model is a neural network model based on the YOLOv5 framework. The YOLO algorithm is a single-stage object detection algorithm based on deep learning, which is characterized by fast detection speed and relatively high accuracy, and can achieve real-time object detection. The core idea of the YOLO algorithm is to take the entire image as the input of the neural network, divide the image into multiple grids, and each grid is responsible for predicting the position and category of a target, thus transforming the object detection into a regression problem. It has the advantages of a small model, fast speed, and relatively high accuracy, and is suitable for scenarios such as forging dimension detection.

[0111] Specifically, the core part of the YOLO algorithm includes a backbone network layer (Backbone), a neck network layer (Neck), and a head network layer (Head). Among them, the backbone network layer mainly performs feature extraction, extracts the object information in the image through a convolutional network, the neck network layer is responsible for multi-scale feature fusion of the feature map, and the head network layer performs the final regression prediction.

[0112] YOLOv5s is the smallest model in the YOLOv5 series of the YOLOv algorithm, featuring fast speed and high accuracy. YOLOv5s is implemented using the PyTorch framework, and a pre-trained model can be loaded through PyTorch Hub, or it can be trained on a custom dataset. The backbone network layer of the traditional YOLOv5s consists of a Focus feature fusion layer and a Cross Stage Partial network layer. The Focus feature fusion layer is responsible for downsampling and channel expansion of the input image, and the Cross Stage Partial network layer is responsible for extracting multi-scale features.

[0113] More specifically, the Cross Stage Partial network layer, namely the C3 layer, is an important feature extraction module in the YOLOv5 algorithm, and its main feature is to extract multi-scale features by stacking multiple bottleneck layers.

[0114] Similarly, the trained forging size detection model also includes a backbone network layer, a neck network layer, and a head network layer. The backbone network layer and the neck network layer have several cross-stage local network layers and several convolutional layers. The forging size detection model has improved the cross-stage local network layer, using deformable convolutions in the bottleneck layers of all cross-stage local network layers and adding channel and spatial attention mechanisms, and all convolutional layers in the backbone network layer and the neck network layer are generalized sparse convolutional layers, that is, the convolutions of the bottleneck layers in all cross-stage local network layers in the traditional YOLOv5s are replaced with deformable convolutions and adding channel and spatial attention mechanisms, and all convolutional layers are replaced with generalized sparse convolutional layers. In the head network layer, the structured intersection-over-union loss function is used as the bounding box loss function to obtain the intersection of the predicted box and the true box.

[0115] In this embodiment, the forging size detection model replaces the standard convolution (Conv) with a deformable convolution, thereby allowing the convolution kernel to dynamically learn the offset, thereby adapting to the shape changes of the forging.

[0116] In this embodiment, a channel and space attention mechanism is also introduced in the improved cross-stage local network layer. The attention mechanism can enhance the sensitivity and importance of the convolutional neural network to different channel features, and increase the sensitivity of the convolutional neural network to small differences. Optimizing the spatial dimension information can improve the performance of the convolutional neural network. Specifically, the channel and space attention mechanism includes a channel attention mechanism layer and a spatial attention mechanism layer connected in sequence. The attention mechanism mainly focuses on the feature fusion between channels of the convolution operation in the backbone network, and then learns the weight coefficient vector corresponding to each channel in the feature map, and then uses the vector to perform weighted processing on the feature map. Based on the consideration of improving the performance and training speed of the forging size detection model, in this embodiment, the original bounding box loss function of the traditional YOLOv5s is replaced with a structured intersection-over-union loss function to further improve the accuracy and training speed of the forging size detection model. The penalty index is redefined by SIoU to take into account the vector angle between the required regressions. The structured intersection-over-union loss function effectively reduces the number of degrees of freedom by adding an angle penalty term. The structured intersection-over-union loss function is used as the bounding box loss function to improve the convergence speed of the model during training and improve the accuracy of bounding box regression.

[0117] The cross-stage local network layer in the forging size detection model includes a first convolutional layer, a bottleneck layer, a first tensor splicing layer, a second convolutional layer, and a third convolutional layer that are sequentially connected in order. The input ends of the first convolutional layer and the second convolutional layer are both used to receive the input data of the cross-stage local network layer. The output end of the second convolutional layer is connected to the input end of the first tensor splicing layer. The output end of the first tensor splicing layer is connected to the input end of the third convolutional layer. The third convolutional layer is used to output the output data of the cross-stage local network layer.

[0118] Specifically, the cross-stage local network layer has two branches: The first branch is the second convolutional layer that performs convolutional processing on the input data; the second branch is the first convolutional layer, the bottleneck layer, the first tensor splicing layer, and the third convolutional layer that are sequentially connected in order. After the input data is input into this branch, feature extraction is performed by the first convolutional layer, and then multi-scale features are extracted in the bottleneck layer. Finally, the features obtained by the bottleneck layer and the first convolutional layer are feature-spliced in the first tensor splicing layer. After that, the spliced features are input into the third convolutional layer for further convolutional processing to obtain the output data of the cross-stage local network layer.

[0119] Since deformable convolution requires additional learning of offset amounts during the convolution process, a branch network is needed to predict the offset parameters. In this embodiment, the bottleneck layer in the forging size detection model adopts a residual network structure. The bottleneck layer includes an identity mapping layer for skip connection, and a fourth convolutional layer, a deformable convolutional layer, a channel attention mechanism layer, a spatial attention mechanism layer, and a feature addition layer, that is, a feature fusion layer, which are sequentially connected in order. The input ends of the fourth convolutional layer and the identity mapping layer are both used to receive the input data of the bottleneck layer. The output end of the identity mapping layer is connected to the output end of the feature addition layer. The feature addition layer is used to output the output data of the bottleneck layer.

[0120] Specifically, the bottleneck layer has two branches: The first branch is the identity mapping layer that performs identity mapping processing on the input data, and the identity mapping layer is used to output the identity mapping features; the second branch is the fourth convolutional layer, the deformable convolutional layer, the channel attention mechanism layer, the spatial attention mechanism layer, and the feature addition layer that are sequentially connected in order. After the input data is input into this branch, feature extraction is performed by the fourth convolutional layer and the deformable convolutional layer, and then the extracted features are sequentially input into the channel attention mechanism layer and the spatial attention mechanism layer for corresponding attention mechanism processing. Finally, the features obtained by the two branches are feature-fused at the feature addition layer to obtain the output data of the bottleneck layer.

[0121] The generalized sparse convolution layer in the forging size detection model includes a fifth convolution layer, a depthwise separable convolution layer, a second tensor splicing layer, and a shuffle layer. The input ends of the fifth convolution layer are all used to receive the input data of the generalized sparse convolution layer. The input end of the depthwise separable convolution layer is connected to the output end of the fifth convolution layer, and the output ends of the depthwise separable convolution layer and the fifth convolution layer are both connected to the input end of the second tensor splicing layer. The output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolution layer.

[0122] Specifically, the input data of the generalized sparse convolution layer is first input into the fifth convolution layer for convolution processing, and then a lightweight convolution operation such as depthwise separable convolution is performed on the output data of the fifth convolution layer. After that, the output data of the fifth convolution layer and the depthwise separable convolution layer are feature-spliced in the second tensor splicing layer. Then, the spliced features are input into the shuffle layer for shuffling processing to obtain the output data of the generalized sparse convolution layer.

[0123] In this embodiment, the traditional standard convolution in the original YOLOv5s is replaced by a generalized sparse convolution layer. The generalized sparse convolution is a lightweight convolution operation that simulates the effect of the standard convolution operation by generating a small number of redundant features, thereby reducing the computational cost. The number of output channels of the generalized sparse convolution is the same as that of the original standard convolution to ensure compatibility with other modules.

[0124] Preferably, the specific configuration (such as the number of groups, kernel size, etc.) of the generalized sparse convolution layer can be continuously adjusted during the model training process to achieve a balance between computational overhead and accuracy.

[0125] In this embodiment, in order to update the network parameters of the forging size detection model, the training of the forging size detection model uses the gradient descent algorithm to continuously optimize the network weights. For example, the mini-batch gradient descent method is adopted, that is, a part of the data is used to calculate the gradient each time, and the historical gradient direction and the current gradient direction can be combined when updating the network parameters. At the same time, a data augmentation strategy can be introduced during the training process to ensure the generalization ability of the forging size detection model for different forging states.

[0126] To improve the computing speed and memory utilization rate of the GPU, the forging size detection model training adopts a mixed-precision training method. Through mixed-precision training, the training efficiency can be improved and the batch size can be increased without reducing the accuracy of the forging size detection model.

[0127] In this embodiment, in order to evaluate the performance of the improved model, three evaluation indicators, namely accuracy, recall rate, and mean average precision, are selected to evaluate the detection effect of the model, and then to judge the actual prediction ability of the improved YOLOv5s network.

[0128] The detection device for the size of forgings in the hot forming process based on machine vision inputs the acquired image data of the forging blank to be detected into the trained forging size detection model, and obtains the size information and classification result of the forging blank to be detected output by the forging size detection model. Among them, the forging size detection model is an improvement on the traditional YOLOv5s model. The convolutional layer in the bottleneck layer of the C3 layer in the traditional YOLOv5s model is replaced with a deformable convolutional layer, which allows the convolutional kernel to dynamically learn the offset amount, so as to adapt to the shape change of the forging in the hot forming process, thereby enhancing the adaptability of the feature map when the shape features of the forging blank change greatly, and making the forging size model pay more attention to the target area, reducing the influence of noise data on the model, improving the perception ability of the model, and introducing a channel and spatial attention mechanism in the C3 layer to enhance the attention ability to the tiny deformations generated at each moment during the hot forming process of the forging blank. Secondly, all convolutional layers in the backbone network layer and the neck network layer are generalized sparse convolutional layers, which reduce the computational complexity and the number of parameters of the model while maintaining the model performance, and use the structured intersection over union loss function as the bounding box loss function to improve the convergence speed of the model during training and improve the accuracy of bounding box regression. The size information and classification result finally predicted by the forging size detection model are more accurate, and because it can pay attention to various types of deformation characteristics of the forging blank, it solves the problems of lack of accuracy and stability caused by relying on manual operation in the current size detection of forgings during the hot forming process, and the problem of forging scrap caused by insufficient deformation accuracy of the blank during the hot forming process resulting in unqualified forming size of the forging, meeting the requirements of real-time on-line measurement of the forging blank size, enabling the production of parts to meet the production requirements of high efficiency and high precision, improving production efficiency and product quality, reducing production costs and resource waste, and having significant economic and social benefits.

[0129] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810 (processor), a communication interface 820 (Communications Interface), a memory 830 (memory), and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the following method:

[0130] Obtain the image data of the forging blank to be detected;

[0131] Input the image data into the trained forging size detection model, and obtain the size information and classification result of the forging blank to be detected output by the forging size detection model;

[0132] The forging size detection model is a neural network model based on the YOLOv5 framework. The forging size detection model includes a backbone network layer, a neck network layer, and a head network layer. There are several cross-stage partial network layers and several generalized sparse convolutional layers in both the backbone network layer and the neck network layer. Moreover, deformable convolutions are used in the bottleneck layers of all cross-stage partial network layers, and the number of channels and the spatial attention mechanism are increased. The structured intersection over union loss function is used as the bounding box loss function in the head network layer. The structured intersection over union loss function is used to determine the penalty exponent and use the penalty exponent to determine the vector angle between the required bounding box regressions;

[0133] The cross-stage partial network layer includes a first convolutional layer, a bottleneck layer, a first tensor splicing layer, a second convolutional layer, and a third convolutional layer connected in sequence;

[0134] The input ends of the first convolutional layer and the second convolutional layer are both used to receive the input data of the cross-stage partial network layer. The output end of the second convolutional layer is connected to the input end of the first tensor splicing layer. The output end of the first tensor splicing layer is connected to the input end of the third convolutional layer. The third convolutional layer is used to output the output data of the cross-stage partial network layer.

[0135] It should be noted that the electronic device in this embodiment can be a server, a PC, or other devices when specifically implemented, as long as its structure includes, for example, Figure 8 a processor 810, a communication interface 820, a memory 830, and a communication bus 840 as shown. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840, and the processor 810 can call the logical instructions in the memory 830 to execute the above method. The specific implementation form of the electronic device in this embodiment is not limited.

[0136] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0137] Furthermore, an embodiment of the present invention discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above-mentioned method embodiments, for example, including:

[0138] Obtain image data of a forging blank to be detected;

[0139] Input the image data into a trained forging size detection model to obtain the size information and classification result of the forging blank to be detected output by the forging size detection model;

[0140] The forging size detection model is a neural network model based on the YOLOv5 framework. The forging size detection model includes a backbone network layer, a neck network layer, and a head network layer. There are several cross-stage partial network layers and several generalized sparse convolution layers in both the backbone network layer and the neck network layer. Moreover, deformable convolutions are used in the bottleneck layers of all cross-stage partial network layers, and channels and spatial attention mechanisms are added. A structured intersection over union loss function is used as the bounding box loss function in the head network layer. The structured intersection over union loss function is used to determine the penalty exponent and use the penalty exponent to determine the vector angle between the required bounding box regressions;

[0141] The cross-stage partial network layer includes a first convolutional layer, a bottleneck layer, a first tensor splicing layer, a second convolutional layer, and a third convolutional layer that are sequentially connected in sequence;

[0142] The input ends of the first convolutional layer and the second convolutional layer are both used to receive the input data of the cross-stage local network layer. The output end of the second convolutional layer is connected to the input end of the first tensor splicing layer. The output end of the first tensor splicing layer is connected to the input end of the third convolutional layer. The third convolutional layer is used to output the output data of the cross-stage local network layer.

[0143] On the other hand, an embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the methods provided in the above embodiments, for example, including:

[0144] Obtain the image data of the forging blank to be detected;

[0145] Input the image data into the trained forging size detection model to obtain the size information and classification result of the forging blank to be detected output by the forging size detection model;

[0146] The forging size detection model is a neural network model based on the YOLOv5 framework. The forging size detection model includes a backbone network layer, a neck network layer, and a head network layer. There are several cross-stage local network layers and several generalized sparse convolutional layers in both the backbone network layer and the neck network layer. Moreover, deformable convolutions are used in the bottleneck layers of all cross-stage local network layers, and the number of channels and the spatial attention mechanism are increased. The structured intersection over union loss function is used as the bounding box loss function in the head network layer. The structured intersection over union loss function is used to determine the penalty exponent and use the penalty exponent to determine the vector angle between the required bounding box regressions;

[0147] The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor splicing layer, a second convolutional layer, and a third convolutional layer that are sequentially connected in sequence;

[0148] The input ends of the first convolutional layer and the second convolutional layer are both used to receive the input data of the cross-stage local network layer. The output end of the second convolutional layer is connected to the input end of the first tensor splicing layer. The output end of the first tensor splicing layer is connected to the input end of the third convolutional layer. The third convolutional layer is used to output the output data of the cross-stage local network layer.

[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0150] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting the size of forgings in a hot forming process, characterized in that: The method comprises: Acquire image data of the forging billet to be inspected; Inputting the image data into a trained forging size detection model to obtain size information and classification results of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the head network layer uses a structured intersection-over-union loss function as a bounding box loss function, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

2. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The inputting of the image data into the trained forging size detection model to obtain the size information and classification result of the forging to be detected output by the forging size detection model specifically includes: Inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain a forging blank fusion feature map output by the backbone network layer; Inputting the forging blank fusion feature map into the neck network layer to perform feature fusion and feature enhancement processing on feature maps of different levels, and obtaining forging blank scale feature maps of three scales output by the neck network layer; All of the forging billet scale feature maps are input into the head network layer for target detection processing, and the size information and classification results of the forging billet to be detected are obtained by the head network layer.

3. The method for detecting the size of forgings in a hot forming process according to claim 2, characterized in that: The backbone network layer includes a focus feature fusion layer, a deep semantic extraction layer and a fast spatial pyramid pooling layer connected in sequence. The deep semantic extraction layer includes a plurality of generalized sparse convolutional layers and a cross-stage local network layer arranged between two adjacent generalized sparse convolutional layers. The first and last layers of the deep semantic extraction layer are both generalized sparse convolutional layers. The generalized sparse convolutional layers are arranged at intervals and two adjacent generalized sparse convolutional layers are connected through a cross-stage local network layer.

4. The method for detecting the size of a forging during hot forming process according to claim 3, characterized in that: The step of inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain a forging blank fusion feature map output by the backbone network layer specifically includes: Inputting the image data into the focus feature fusion layer for channel expansion and downsampling processing to obtain a downsampled feature map output by the focus feature fusion layer; Inputting the downsampled feature map into the deep semantic extraction layer to extract the deep semantic information in the feature map, and obtaining a scale feature map output by the deep semantic extraction layer; The scale feature map is input into the fast spatial pyramid pooling layer for scale fusion processing to obtain the forging blank fusion feature map.

5. The method for detecting the size of forgings in a hot forming process according to claim 4, characterized in that: The step of inputting the downsampled feature map into the deep semantic extraction layer to extract the deep semantic information in the feature map to obtain the scale feature map output by the deep semantic extraction layer specifically includes: Inputting the downsampled feature map into the first generalized sparse convolution layer in the deep semantic extraction layer for generalized sparse convolution processing to obtain a first convolution feature output by the first generalized sparse convolution layer; Inputting the first convolutional feature into the cross-stage local network layer of the next layer connected to the generalized sparse convolutional layer to extract multi-scale features, and obtaining a second convolutional feature output by the cross-stage local network layer; The second convolution feature is input into the generalized sparse convolution layer of the next layer connected to the cross-stage local network layer for generalized sparse convolution processing to obtain the third convolution feature output by the generalized sparse convolution layer, until the third convolution feature is input into the last generalized sparse convolution layer to obtain the scale feature map output by the last generalized sparse convolution layer.

6. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The bottleneck layer includes an identity mapping layer, and a fourth convolution layer, a variable convolution layer, a channel attention mechanism layer, a spatial attention mechanism layer and a feature addition layer connected in sequence; The input ends of the fourth convolutional layer and the identity mapping layer are both used to receive input data of the bottleneck layer, the output end of the identity mapping layer is connected to the output end of the feature addition layer, and the feature addition layer is used to output the output data of the bottleneck layer.

7. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The generalized sparse convolution layer includes a fifth convolution layer, a depth-separable convolution layer, a second tensor concatenation layer, and a shuffle layer; The input end of the fifth convolutional layer is used to receive the input data of the generalized sparse convolutional layer, the input end of the depthwise separable convolutional layer is connected to the output end of the fifth convolutional layer, and the output ends of the depthwise separable convolutional layer and the fifth convolutional layer are both connected to the input end of the second tensor splicing layer, the output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolutional layer.

8. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The forging size detection model is trained by the following steps: Acquiring original image data of a sample forging, and preprocessing the original image data to obtain sample image data of the sample forging; Performing target labeling on the sample image data to obtain labeling information; The annotation information includes a size boundary box and a category label of the sample forging; The sample image data is used as input data for training, the annotation information is used as labels for training, machine learning is used for training, and the network weights are adjusted by gradient descent to obtain the forging size detection model for predicting the size information and classification results of the forging to be detected.

9. A system for detecting the size of forgings in a hot forming process, characterized in that: The system comprises: A data acquisition module, used to acquire image data of the forging to be inspected; A size detection module, used for inputting the image data into a trained forging size detection model to obtain the size information and classification result of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the head network layer uses a structured intersection-over-union loss function as a bounding box loss function, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting the size of a forging in a hot forming process as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Cross-domain hyperspectral image classification method based on space attention guidance variable convolution

    CN116310810A

  • Remote sensing image target dynamic detection method based on multi-kernel

    CN119580116A