Method and device for detecting size of forge piece in hot forming process and medium

Through the improved YOLOv5 neural network model, combined with deformable convolution, attention mechanism and generalized sparse convolution layer, the real-time online measurement problem of forging dimension detection during thermoforming is solved, and high-precision forging dimension detection is achieved, avoiding the scrapping of forgings.

CN119919474AActive Publication Date: 2025-05-02WUHAN UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510414401.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-02
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing machine vision technology cannot measure the tiny dimension deformation of forgings online in real time during the thermoforming process, resulting in the forgings not meeting the standards and scrapping.

Method used

Using a neural network model based on the YOLOv5 framework, the forging size detection model is improved by introducing deformable convolution, channel and spatial attention mechanisms, and using generalized sparse convolution layers and structured interleaving ratio loss function, and enhancing its detection ability of forging shape changes and slight deformation.

Benefits of technology

Real-time online measurement of forging dimensions is realized, the accuracy and stability of detection is improved, the scrapping of forgings is avoided, and the production requirements of high efficiency and high precision are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919474A_ABST
    Figure CN119919474A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for detecting the size of a forge piece in the hot forming process and a medium, and relates to the technical field of computer vision, and the method comprises the steps that image data of a to-be-detected forging stock is obtained; inputting the image data into a trained forging size detection model to obtain size information and a classification result; the forging size detection model comprises a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer are each provided with a plurality of cross-stage local network layers and a plurality of generalized sparse convolution layers. A bottleneck layer in all cross-stage local network layers uses deformable convolution, and a channel and a space attention mechanism are added. According to the invention, the problem of lack of accuracy and stability caused by dependence on manual operation in size detection of the forge piece in the hot forming process at present is solved, and the problem of forge piece scrapping caused by substandard forming size of the forge piece due to insufficient blank deformation precision in the hot forming process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device and medium for detecting the size of a forging in a hot forming process. Background Art

[0002] The application of advanced science and technology such as sensor technology, computer technology, and intelligent control technology in the manufacturing industry has promoted the modern industry to move towards intelligence, informationization, and scale. However, at present, the size of forgings in the hot forming process still relies on manual experience for detection, that is, the determination of forging size depends largely on the personal experience and visual judgment of the inspectors. Usually, inspectors need to rely on experience to judge whether the forgings have reached the expected size, and then use a ruler to perform actual measurements. This method not only consumes a lot of time and human resources and is difficult to ensure high production efficiency, but also has deviations in forging size, which will cause uneven internal structure of the forgings and affect their mechanical properties.

[0003] Machine vision technology plays an important role in the field of target detection with its high efficiency and accuracy. Machine vision technology can automatically detect and identify targets, which not only improves production efficiency and detection accuracy, but also helps inspectors get rid of repetitive and time-consuming tasks. Convolutional Neural Network (CNN), commonly used in machine vision technology, is a strong target detection algorithm that can replace the feature extractor and classifier in traditional feature learning theory, so that the model does not need to perform complex preprocessing on the original image, and completes the target detection process quickly and concisely. Due to the advantages of automatic feature extraction and strong generalization ability of convolutional neural networks, convolutional neural networks are also increasingly used in forging size detection scenarios.

[0004] During the hot forming process, the billet will deform. If the deformation accuracy of the billet is not enough, the forming size of the forging will not meet the standard, which will lead to the scrapping of the forging. However, the current CNN model can only focus on the larger deformation of the billet, but not the smaller deformation of the billet, and cannot fully meet the requirements of real-time online measurement of the forging billet size.

[0005] If the problem of insufficient deformation accuracy can be discovered in time during the forging process, remedial measures can be taken in time. Therefore, it is necessary to provide a method for detecting the size of forgings in a hot forming process that can solve the above problem. Summary of the invention

[0006] In view of this, the present invention proposes a method, device and medium for detecting the size of forgings in a hot forming process to solve the problem that the current machine vision technology using convolutional neural networks cannot perform real-time online measurement of forgings.

[0007] The technical solution of the present invention is achieved in this way: According to a first aspect, an embodiment of the present invention provides a method for detecting the size of a forging in a hot forming process, the method comprising: Acquire image data of the forging billet to be inspected; Inputting the image data into a trained forging size detection model to obtain size information and classification results of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

[0008] In combination with the first aspect, in a first implementation manner of the first aspect, the inputting of the image data into a trained forging size detection model to obtain the size information and classification result of the forging to be detected output by the forging size detection model specifically includes: Inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain a forging blank fusion feature map output by the backbone network layer; Inputting the forging blank fusion feature map into the neck network layer to perform feature fusion and feature enhancement processing on feature maps of different levels, and obtaining forging blank scale feature maps of three scales output by the neck network layer; All of the forging billet scale feature maps are input into the head network layer for target detection processing, and the size information and classification results of the forging billet to be detected are obtained by the head network layer.

[0009] In combination with the first embodiment of the first aspect, in the second embodiment of the first aspect, the backbone network layer includes a focal feature fusion layer, a deep semantic extraction layer, and a fast spatial pyramid pooling layer connected in sequence, the deep semantic extraction layer includes a plurality of generalized sparse convolutional layers and a cross-stage local network layer arranged between two adjacent generalized sparse convolutional layers, the first and last layers of the deep semantic extraction layer are both generalized sparse convolutional layers, the generalized sparse convolutional layers are arranged at intervals, and two adjacent generalized sparse convolutional layers are connected through a cross-stage local network layer.

[0010] In combination with the second embodiment of the first aspect, in the third embodiment of the first aspect, the inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain the forging blank fusion feature map output by the backbone network layer specifically includes: Inputting the image data into the focus feature fusion layer for channel expansion and downsampling processing to obtain a downsampled feature map output by the focus feature fusion layer; Inputting the downsampled feature map into the deep semantic extraction layer to extract the deep semantic information in the feature map, and obtaining a scale feature map output by the deep semantic extraction layer; The scale feature map is input into the fast spatial pyramid pooling layer for scale fusion processing to obtain the forging blank fusion feature map.

[0011] In combination with the third implementation manner of the first aspect, in the fourth implementation manner of the first aspect, inputting the downsampled feature map into the deep semantic extraction layer to extract the depth semantic information in the feature map to obtain the scale feature map output by the deep semantic extraction layer specifically includes: Inputting the downsampled feature map into the first generalized sparse convolution layer in the deep semantic extraction layer for generalized sparse convolution processing to obtain a first convolution feature output by the first generalized sparse convolution layer; Inputting the first convolutional feature into the cross-stage local network layer of the next layer connected to the generalized sparse convolutional layer to extract multi-scale features, and obtaining a second convolutional feature output by the cross-stage local network layer; The second convolution feature is input into the generalized sparse convolution layer of the next layer connected to the cross-stage local network layer for generalized sparse convolution processing to obtain the third convolution feature output by the generalized sparse convolution layer, until the third convolution feature is input into the last generalized sparse convolution layer to obtain the scale feature map output by the last generalized sparse convolution layer.

[0012] In combination with the first aspect, in a fifth implementation of the first aspect, the bottleneck layer includes an identity mapping layer, and a fourth convolutional layer, a variable convolutional layer, a channel attention mechanism layer, a spatial attention mechanism layer, and a feature addition layer connected in sequence; The input ends of the fourth convolutional layer and the identity mapping layer are both used to receive input data of the bottleneck layer, the output end of the identity mapping layer is connected to the output end of the feature addition layer, and the feature addition layer is used to output the output data of the bottleneck layer.

[0013] In combination with the first aspect, in a sixth implementation of the first aspect, the generalized sparse convolutional layer includes a fifth convolutional layer, a depthwise separable convolutional layer, a second tensor concatenation layer, and a shuffle layer; The input end of the fifth convolutional layer is used to receive the input data of the generalized sparse convolutional layer, the input end of the depthwise separable convolutional layer is connected to the output end of the fifth convolutional layer, and the output ends of the depthwise separable convolutional layer and the fifth convolutional layer are both connected to the input end of the second tensor splicing layer, the output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolutional layer.

[0014] In combination with the first aspect, in a seventh implementation manner of the first aspect, the forging size detection model is trained by the following steps: Acquiring original image data of a sample forging, and preprocessing the original image data to obtain sample image data of the sample forging; Performing target labeling on the sample image data to obtain labeling information; the labeling information includes a size boundary box and a category label of the sample forging; The sample image data is used as input data for training, the annotation information is used as labels for training, machine learning is used for training, and the network weights are adjusted by gradient descent to obtain the forging size detection model for predicting the size information and classification results of the forging to be detected.

[0015] According to a second aspect, an embodiment of the present invention further provides a device for detecting the size of a forging in a hot forming process based on machine vision, the device comprising: A data acquisition module, used to acquire image data of the forging to be inspected; A size detection module, used for inputting the image data into a trained forging size detection model to obtain the size information and classification result of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

[0016] According to a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for detecting the size of forgings in a hot forming process as described in any one of the above items are implemented.

[0017] According to a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for detecting the size of forgings in a hot forming process as described in any one of the above items.

[0018] The method, device and medium for detecting the size of forgings in a hot forming process of the present invention have the following beneficial effects compared with the prior art: The acquired image data of the forging to be detected is input into the trained forging size detection model to obtain the size information and classification results of the forging to be detected output by the forging size detection model. Among them, the forging size detection model is an improvement on the traditional YOLOv5s model. The convolution layer in the bottleneck layer of the C3 layer of the traditional YOLOv5s model is replaced with a deformable convolution layer, thereby allowing the convolution kernel to dynamically learn the offset, so as to adapt to the shape change of the forging during the hot forming process, so as to enhance the adaptability of the feature map when the shape characteristics of the forging change greatly, and make the forging size model pay more attention to the target area, reduce the influence of noise data on the model, improve the perception ability of the model, and introduce the channel and spatial attention mechanism in the C3 layer to enhance the attention to the tiny deformation of the forging at every moment during the hot forming process. Secondly, all the convolution layers in the backbone network layer and the neck network layer are The generalized sparse convolutional layer reduces the computational complexity and number of parameters of the model while maintaining the model performance, and uses the structured intersection-over-union loss function as the bounding box loss function to improve the convergence speed of the model during training and the accuracy of bounding box regression. The size information and classification results finally predicted by the forging size detection model are more accurate. In addition, since it can pay attention to the various types of deformation characteristics of the forging billet, it solves the current problem of lack of accuracy and stability in the size detection of forgings in the hot forming process due to reliance on manual operation, as well as the problem of insufficient deformation accuracy of the billet during the hot forming process, resulting in the forgings being scrapped due to the non-compliance of the forming size. It meets the requirements for real-time online measurement of the forging billet size, so that the production of parts can meet the high-efficiency and high-precision production requirements, improves production efficiency and product quality, reduces production costs and resource waste, and has significant economic and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 A schematic diagram showing the flow of a method for detecting dimensions of a forging during a hot forming process of the present invention; Figure 2 A schematic diagram of the structure of the existing YOLOv5s model is shown; Figure 3 The schematic diagram of the structure of the model in the method for detecting the size of forgings in the hot forming process of the present invention is shown; Figure 4A schematic diagram showing the structure of an improved cross-stage local network layer in the method for detecting the size of a forging in a hot forming process of the present invention is shown; Figure 5 The schematic diagram of the structure of the improved bottleneck layer in the method for detecting the size of forgings in a hot forming process of the present invention is shown; Figure 6 A schematic diagram showing the structure of a generalized sparse convolutional layer in the method for detecting the size of a forging in a hot forming process of the present invention; Figure 7 The schematic diagram of the structure of the device for detecting the size of forgings in a hot forming process based on machine vision provided by the present invention is shown; Figure 8 A structural schematic diagram of an electronic device provided by the present invention is shown. DETAILED DESCRIPTION

[0021] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] The application of advanced science and technology such as sensor technology, computer technology, and intelligent control technology in the manufacturing industry has promoted the modern industry to move towards intelligence, informationization, and scale. However, at present, the size of forgings in the hot forming process still relies on manual experience for detection, that is, the determination of forging size depends largely on the personal experience and visual judgment of the inspectors. Usually, inspectors need to rely on experience to judge whether the forgings have reached the expected size, and then use a ruler to perform actual measurements. This method not only consumes a lot of time and human resources and is difficult to ensure high production efficiency, but also has deviations in forging size, which will cause uneven internal structure of the forgings and affect their mechanical properties.

[0023] Machine vision technology plays an important role in the field of target detection with its high efficiency and accuracy. Machine vision technology can automatically detect and identify targets, which not only improves production efficiency and detection accuracy, but also helps inspectors get rid of repetitive and time-consuming tasks. CNN, commonly used in machine vision technology, is a strong target detection algorithm that can replace the feature extractor and classifier in traditional feature learning theory, so that the model does not need to perform complex preprocessing on the original image and completes the target detection process quickly and concisely. Due to the advantages of automatic feature extraction and strong generalization ability of convolutional neural networks, convolutional neural networks are also increasingly used in forging size detection scenarios.

[0024] During the hot forming process, the billet will deform. If the deformation accuracy of the billet is not enough, the forming size of the forging will not meet the standard, which will lead to the scrapping of the forging. However, the current CNN model can only focus on the larger deformation of the billet, but not the smaller deformation of the billet, and cannot fully meet the requirements of real-time online measurement of the forging billet size.

[0025] If the problem of insufficient deformation accuracy can be discovered in time during the forging process, remedial measures can be taken in time. Therefore, it is necessary to provide a method for detecting the size of forgings in a hot forming process that can solve the above problem.

[0026] The method, device and medium for detecting the size of forgings in the hot forming process provided in this specification are intended to solve the problem of lack of accuracy and stability in the current dimensional detection of forgings in the hot forming process due to reliance on manual operation, as well as the problem of insufficient deformation accuracy of the blank during the hot forming process resulting in the forgings not meeting the forming size standards and causing the forgings to be scrapped.

[0027] The method for detecting the size of forgings during hot forming provided in this specification can be applied to electronic devices with computer vision capabilities. The electronic devices may include notebooks, desktop computers, smart phones, smart wearable devices (virtual reality glasses, smart watches, etc.), tablet computers, etc. Of course, the method for detecting the size of forgings during hot forming provided in this specification can also be applied to applications running in the above-mentioned electronic devices. For example, the method for detecting the size of forgings during hot forming can be applied to a browser with computer vision capabilities, or to software with computer vision capabilities that responds immediately.

[0028] See also Figure 1 , Figure 1 A flow chart showing a method for detecting dimensions of a forging during a hot forming process according to an embodiment of the present invention is provided. The method may include the following steps: S101, obtaining image data of a forging to be inspected.

[0029] In this embodiment, the above-mentioned image data may be stored in the electronic device in advance, or may be acquired by the electronic device from the outside. For example, by arranging an industrial camera at the production site, the image data of the forging to be inspected during the production process is captured. There is no restriction on the specific acquisition form of the image data, and it is only necessary to ensure that the electronic device can acquire the image data.

[0030] It should be noted that the forgings obtained through the entire forging process can also be used as forging blanks to be tested. By performing dimensional inspection on the formed forgings, it can be determined whether the forgings meet the actual forging standards.

[0031] S102, inputting the image data into the trained forging size detection model to obtain the size information and classification result of the forging to be detected output by the forging size detection model.

[0032] In this process, the image data is used as the input data of the forging size detection model. The forging size detection model processes the image data and outputs the size information and classification results of the forging.

[0033] In this embodiment, the forging size detection model is a neural network model based on the YOLOv5 framework. The YOLO algorithm is a single-stage target detection algorithm based on deep learning. It is characterized by fast detection speed, high accuracy, and the ability to achieve real-time target detection. The core idea of ​​the YOLO algorithm is to use the entire image as the input of the neural network and divide the image into multiple grids. Each grid is responsible for predicting the position and category of a target, thereby converting target detection into a regression problem. It has the advantages of small model, fast speed, and high accuracy, and is suitable for scenarios such as forging size detection.

[0034] Specifically, the core parts of the YOLO algorithm include the backbone network layer (Backbone), the neck network layer (Neck) and the head network layer (Head). Among them, the backbone network layer mainly performs feature extraction, extracting object information in the image through a convolutional network, the neck network layer is responsible for multi-scale feature fusion of feature maps, and the head network layer performs the final regression prediction.

[0035] See also Figure 2 YOLOv5s is the smallest model in the YOLOv5 series of the YOLOv algorithm, with the characteristics of fast speed and high accuracy. YOLOv5s is implemented using the PyTorch framework, and can load pre-trained models through the PyTorch Hub, or train on custom datasets. The backbone network layer of the traditional YOLOv5s consists of a focus feature fusion layer (Focus) and a cross-stage partial network layer (Cross Stage Partial Network, CSPNet). The focus feature fusion layer is responsible for downsampling and channel expansion of the input image, and the cross-stage partial network layer is responsible for extracting multi-scale features.

[0036] More specifically, the cross-stage local network layer, namely the C3 layer, is an important feature extraction module in the YOLOv5 algorithm. Its main feature is to extract multi-scale features by stacking multiple bottleneck layers.

[0037] See also Figure 3Similarly, the trained forging size detection model also includes a backbone network layer, a neck network layer, and a head network layer. The backbone network layer and the neck network layer have several cross-stage local network layers and several convolutional layers. The forging size detection model improves the cross-stage local network layer, and uses deformable convolution network (DCN) in the bottleneck layer of all cross-stage local network layers and adds channel and spatial attention mechanisms. In addition, all convolutional layers in the backbone network layer and the neck network layer are generalized sparse convolutional layers, that is, the convolution of the bottleneck layer in all cross-stage local network layers in the traditional YOLOv5s is replaced by deformable convolution and adds channel and spatial attention mechanisms, and all convolutional layers are replaced by generalized sparse convolutional layers. In the head network layer of the forging size detection model, the structured intersection-over-union loss function is used as the bounding box loss function to obtain the intersection of the predicted box and the true box.

[0038] In this embodiment, the forging size detection model replaces the standard convolution (Conv) with a deformable convolution, thereby allowing the convolution kernel to dynamically learn the offset, thereby adapting to the shape changes of the forging.

[0039] In this embodiment, a channel and space attention mechanism (Squeeze-and-Excitation, SE) is also introduced in the improved cross-stage local network layer. The attention mechanism can enhance the sensitivity and importance of the convolutional neural network to different channel features, and increase the sensitivity of the convolutional neural network to small differences. Optimizing the spatial dimension information can improve the performance of the convolutional neural network. In this way, the improved cross-stage local network layer not only enhances the adaptability of the feature map when the shape characteristics of the forging billet change significantly through DCN, but also allows the entire forging size detection model to pay more attention to the target area, reduce the impact of noise, and improve the perception ability of the forging size detection model. By adding the channel and space attention mechanism, the attention to the small deformation of the forging billet at each moment in the hot forming process is enhanced, so that the forging size detection model finally trained can meet the requirements of real-time online measurement of the forging billet size. Specifically, the channel and spatial attention mechanism includes a channel attention mechanism layer and a spatial attention mechanism layer connected in sequence. The attention mechanism mainly focuses on the feature fusion between channels of the convolution operation in the backbone network, and then learns the weight coefficient vector corresponding to each channel in the feature map, and then uses the vector to weight the feature map. Among them, channel attention processing can identify which channels contain more important information and give these channels higher weights, thereby solving the problem of conflicting information between features of different scales. Spatial attention processing can identify important areas in the image data and give higher weights to the features of these areas, thereby solving the problem of lack of context information. Based on the consideration of improving the performance and training speed of the forging size detection model, in this embodiment, the original bounding box loss function of the traditional YOLOv5s is replaced with the structured intersection over Union loss (SIoU-Loss) to further improve the accuracy and training speed of the forging size detection model. The penalty index is redefined by SIoU to take into account the vector angle between the required regressions. The structured intersection-over-union loss function effectively reduces the number of degrees of freedom by adding an angle penalty term. The structured intersection-over-union loss function is used as the bounding box loss function to improve the convergence speed of the model during training and improve the accuracy of bounding box regression.

[0040] See also Figure 4 The cross-stage local network layer in the forging size detection model includes a first convolution layer, a bottleneck layer, a first tensor splicing layer, a second convolution layer and a third convolution layer connected in sequence. The input ends of the first convolution layer and the second convolution layer are used to receive input data of the cross-stage local network layer, the output end of the second convolution layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolution layer, and the third convolution layer is used to output the output data of the cross-stage local network layer.

[0041] Specifically, the cross-stage local network layer has two branches: the first branch is the second convolution layer that performs convolution processing on the input data; the second branch is the first convolution layer, the bottleneck layer, the first tensor splicing layer and the third convolution layer connected in sequence. After the input data is input into this branch, the first convolution layer performs feature extraction, and then it is input into the bottleneck layer to extract multi-scale features. Finally, the features obtained from the bottleneck layer and the first convolution layer are respectively spliced ​​in the first tensor splicing layer, and then the spliced ​​features are input into the third convolution layer for convolution processing to obtain the output data of the cross-stage local network layer.

[0042] Since deformable convolution requires additional learning of offsets during the convolution process, a branch network is required to predict the offset parameters. Figure 5 In this embodiment, the bottleneck layer in the forging size detection model adopts a residual network structure. The bottleneck layer includes an identity mapping layer for jump connection, and a fourth convolutional layer, a variable convolutional layer, a channel attention mechanism layer, a spatial attention mechanism layer, and a feature addition layer, i.e., a feature fusion layer, which are sequentially connected. The input ends of the fourth convolutional layer and the identity mapping layer are used to receive input data of the bottleneck layer, and the output end of the identity mapping layer is connected to the output end of the feature addition layer. The feature addition layer is used to output the output data of the bottleneck layer.

[0043] Specifically, the bottleneck layer has two branches: the first branch is the identity mapping layer that performs identity mapping processing on the input data, and the identity mapping layer is used to output the identity mapping feature; the second branch is the fourth convolution layer, the variable convolution layer, the channel attention mechanism layer, the spatial attention mechanism layer and the feature addition layer connected in sequence. After the input data is input into this branch, the fourth convolution layer and the variable convolution layer perform feature extraction, and then the extracted features are input into the channel attention mechanism layer and the spatial attention mechanism layer in turn for corresponding attention mechanism processing. Finally, the features obtained from the two branches are fused at the feature addition layer to obtain the output data of the bottleneck layer.

[0044] See also Figure 6 The generalized sparse convolutional layer in the forging size detection model includes a fifth convolutional layer, a depthwise separable convolutional layer, a second tensor splicing layer, and a shuffle layer. The input end of the fifth convolutional layer is used to receive the input data of the generalized sparse convolutional layer, the input end of the depthwise separable convolutional layer is connected to the output end of the fifth convolutional layer, and the output ends of the depthwise separable convolutional layer and the fifth convolutional layer are connected to the input end of the second tensor splicing layer, the output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolutional layer.

[0045] Specifically, the input data of the generalized sparse convolutional layer is first input into the fifth convolutional layer for convolution processing, and then the output data of the fifth convolutional layer is subjected to a lightweight convolution operation such as depthwise separable convolution. After that, the output data of the fifth convolutional layer and the depthwise separable convolutional layer are feature spliced ​​in the second tensor splicing layer, and then the spliced ​​features are input into the shuffling layer for shuffling processing to obtain the output data of the generalized sparse convolutional layer.

[0046] In this embodiment, the original standard convolution of the traditional YOLOv5s is replaced by a generalized sparse convolution layer. The generalized sparse convolution is a lightweight convolution operation that simulates the effect of the standard convolution operation by generating a small amount of redundant features, thereby reducing the computational cost. The number of output channels of the generalized sparse convolution is consistent with the original standard convolution to ensure compatibility with other modules.

[0047] Preferably, the specific configuration of the generalized sparse convolutional layer (such as the number of groups, kernel size, etc.) can be continuously adjusted during the model training process to achieve a balance between computational overhead and accuracy.

[0048] Accordingly, the forging size detection model is trained through the following steps: The original image data of the sample forging is obtained, and the original image data is preprocessed to obtain the sample image data of the sample forging. The preprocessing includes image enhancement processing and image denoising processing on the original image data.

[0049] Since the amount of sample data of abnormal states is small, it is necessary to perform image enhancement and image denoising on the collected original image data. Image enhancement may include but is not limited to: changing brightness and contrast, adding shadows and mist, and by preprocessing the original image data, expanding the data set and improving the robustness of the model.

[0050] Afterwards, the sample image data is labeled to obtain the labeling information; the labeling information includes the size boundary box and category label of the sample forging.

[0051] In this embodiment, LabelImg is used to target the state of the sample forging in the sample image data.

[0052] During training, the sample image data is used as the input data for training, and the annotation information is used as the label for training. Machine learning is used for training and the network weights are adjusted by gradient descent to obtain a forging size detection model for predicting the size information and classification results of the forging to be detected.

[0053] The trained real-time detection model is deployed to the mobile device, and the image data of the forging to be inspected during the forging process is obtained in real time. The dimensions of forgings in different states are quickly detected by optimizing weight parameters, and the detection results are output.

[0054] In this embodiment, in order to update the network parameters of the forging size detection model, the training of the forging size detection model uses a gradient descent algorithm to continuously optimize the network weights, such as a small batch gradient descent method, that is, using part of the data to calculate the gradient each time, and the historical gradient direction and the current gradient direction can be combined when updating the network parameters. At the same time, a data enhancement strategy can also be introduced during the training process to ensure the generalization ability of the forging size detection model for different forging states.

[0055] In order to improve the computing speed and memory utilization of the Graphics Processing Unit (GPU), the mixed precision training method is adopted in the training of the forging size detection model. Through mixed precision training, the training efficiency can be improved and the batch size can be increased without reducing the accuracy of the forging size detection model.

[0056] In this embodiment, in order to evaluate the performance of the improved model, three evaluation indicators, namely accuracy, recall and average precision, are selected to evaluate the detection effect of the model and further determine the actual prediction ability of the improved YOLOv5s network. The corresponding evaluation indicators include: (1) In formula (1), Indicates accuracy (precision); Indicates the number of samples correctly identified as positive; Indicates the number of samples that are incorrectly identified as positive (actually should be negative).

[0057] (2) In formula (2), Represents the recall rate (recall); Indicates the number of samples that are incorrectly identified as negative (they should actually be positive).

[0058] (3) In formula (3), represents the mean average precision. It reflects the detection accuracy of the model. The higher the value, the better the model accuracy. Indicates the category of the detected sample. In this embodiment, ; Represents the average precision of each category.

[0059] The processing process of the trained forging size model for the image data of the forging to be inspected is as follows: The input image data of the forging billet to be inspected is subjected to feature extraction and fusion processing in the backbone network layer to obtain the forging billet fusion feature map output by the backbone network layer, and then the forging billet fusion feature map is input into the neck network layer for feature fusion and feature enhancement processing of feature maps of different levels, and further features are extracted to obtain forging billet scale feature maps of three scales output by the neck network layer. Finally, all the forging billet scale feature maps are input into the head network layer for target detection processing to obtain the size information and classification results of the forging billet to be inspected output by the head network layer.

[0060] by Figure 3 Taking the example, the backbone network layer includes a focal feature fusion layer, a deep semantic extraction layer and a fast spatial pyramid pooling layer connected in sequence. The deep semantic extraction layer includes several generalized sparse convolutional layers and a cross-stage local network layer arranged between two adjacent generalized sparse convolutional layers. The first and last layers of the deep semantic extraction layer are both generalized sparse convolutional layers. The generalized sparse convolutional layers are arranged at intervals and two adjacent generalized sparse convolutional layers are connected through a cross-stage local network layer.

[0061] Preferably, the deep semantic extraction layer includes a total of 4 layers of generalized sparse convolutional layers and 3 layers of cross-stage local network layers arranged between the 4 layers of generalized sparse convolutional layers. The backbone network slices the image data through the focus feature fusion module, expands the input channel to 4 times the original, and obtains a downsampled feature map after one convolution. The feature map is further block-extracted through 4 generalized sparse convolutions and 3 improved C3 layers to further extract the deep semantic information of the image and reduce the amount of calculation, thereby obtaining a scale feature map output by the deep semantic extraction layer. Finally, after the fast spatial pyramid pooling module, the feature maps of different scales are fused into a feature map of unified scale to obtain the forging fusion feature map. More specifically, the downsampled feature map is input into the first generalized sparse convolution layer in the deep semantic extraction layer for generalized sparse convolution processing to obtain the first convolution feature output by the first generalized sparse convolution layer, and then the first convolution feature is input into the next layer of the cross-stage local network layer connected to the generalized sparse convolution layer to extract multi-scale features to obtain the second convolution feature output by the cross-stage local network layer, and then the second convolution feature is input into the next layer of the generalized sparse convolution layer connected to the cross-stage local network layer for generalized sparse convolution processing to obtain the third convolution feature output by the generalized sparse convolution layer, until the third convolution feature is input into the last generalized sparse convolution layer to obtain the scale feature map output by the last generalized sparse convolution layer. This process improves the robustness of the network. Afterwards, the feature map output by the backbone network layer is input into the neck network layer, which includes a feature pyramid network layer and a path aggregation network layer. The feature pyramid network layer downsamples the high-level feature map output by the improved C3 layer in a top-down manner and fuses it with the low-level feature map. After that, the path aggregation network layer further enhances the feature fusion from the bottom to the top, and passes the strong positioning information in the feature map output by the low-level improved C3 layer to the high-level feature map neck network. By performing branch fusion on multiple feature maps output by the backbone network layer, three fusion features of different scales are obtained by three fusion routes respectively. After that, the neck network layer inputs the output feature maps of three different scales into the head network layer, and the head network layer performs image target detection on them respectively, and obtains the intersection and union of the predicted box and the real box according to the structured intersection-over-union loss function.

[0062] The method for detecting the size of forgings in a hot forming process of the present invention inputs the acquired image data of the forging to be detected into a trained forging size detection model to obtain the size information and classification results of the forging to be detected output by the forging size detection model, wherein the forging size detection model is an improvement on the traditional YOLOv5s model, and the convolution layer in the bottleneck layer in the C3 layer of the traditional YOLOv5s model is replaced with a deformable convolution layer, thereby allowing the convolution kernel to dynamically learn the offset, thereby adapting to the shape change of the forging in the hot forming process, thereby enhancing the adaptability of the feature map when the shape characteristics of the forging change greatly, and allowing the forging size model to pay more attention to the target area, reduce the impact of noise data on the model, improve the perception ability of the model, and introduce channel and spatial attention mechanisms in the C3 layer to enhance the ability to pay attention to the tiny deformation of the forging at every moment during the hot forming process. Secondly, the backbone network layer and the neck network layer All convolutional layers in the method are generalized sparse convolutional layers, which reduce the computational complexity and number of parameters of the model while maintaining the model performance. The structured intersection-over-union loss function is used as the bounding box loss function to improve the convergence speed of the model during training and the accuracy of bounding box regression. The size information and classification results finally predicted by the forging size detection model are more accurate. In addition, since it can pay attention to the various types of deformation characteristics of the forging billet, it solves the current problem of lack of accuracy and stability in the size detection of forgings in the hot forming process due to reliance on manual operation, as well as the problem of insufficient deformation accuracy of the billet in the hot forming process resulting in the forging size not meeting the standard and resulting in the scrapping of the forging. It meets the requirements for real-time online measurement of the forging billet size, so that the production of parts can meet the high-efficiency and high-precision production requirements, improves production efficiency and product quality, reduces production costs and resource waste, and has significant economic and social benefits.

[0063] The following describes an apparatus provided by an embodiment of the present invention. The apparatus described below and the method described above can refer to each other.

[0064] See also Figure 7 , Figure 7 A schematic diagram showing a method for detecting dimensions of a forging during a hot forming process according to an embodiment of the present invention is shown. The device may include: The data acquisition module 10 is used to acquire the image data of the forging blank to be inspected.

[0065] In this embodiment, the above-mentioned image data may be stored in the electronic device in advance, or may be acquired by the electronic device from the outside. For example, by arranging an industrial camera at the production site, the image data of the forging to be inspected during the production process is captured. There is no restriction on the specific acquisition form of the image data, and it is only necessary to ensure that the electronic device can acquire the image data.

[0066] The size detection module 20 is used to input the image data into the trained forging size detection model to obtain the size information and classification results of the forging to be detected output by the forging size detection model.

[0067] In this process, the image data is used as the input data of the forging size detection model. The forging size detection model processes the image data and outputs the size information and classification results of the forging.

[0068] In this embodiment, the forging size detection model is a neural network model based on the YOLOv5 framework. The YOLO algorithm is a single-stage target detection algorithm based on deep learning, which is characterized by fast detection speed, high accuracy, and the ability to achieve real-time target detection. The core idea of ​​the YOLO algorithm is to use the entire image as the input of the neural network, divide the image into multiple grids, and each grid is responsible for predicting the position and category of a target, thereby converting target detection into a regression problem. It has the advantages of small model, fast speed, and high accuracy, and is suitable for scenarios such as forging size detection.

[0069] Specifically, the core parts of the YOLO algorithm include the backbone network layer (Backbone), the neck network layer (Neck) and the head network layer (Head). Among them, the backbone network layer mainly performs feature extraction, extracting object information in the image through a convolutional network, the neck network layer is responsible for multi-scale feature fusion of feature maps, and the head network layer performs the final regression prediction.

[0070] YOLOv5s is the smallest model in the YOLOv5 series of the YOLOv algorithm, with the characteristics of fast speed and high accuracy. YOLOv5s is implemented using the PyTorch framework, and can load pre-trained models through the PyTorch Hub, or train on custom datasets. The backbone network layer of the traditional YOLOv5s consists of a focus feature fusion layer (Focus) and a cross-stage local network layer. The focus feature fusion layer is responsible for downsampling and channel expansion of the input image, and the cross-stage local network layer is responsible for extracting multi-scale features.

[0071] More specifically, the cross-stage local network layer, namely the C3 layer, is an important feature extraction module in the YOLOv5 algorithm. Its main feature is to extract multi-scale features by stacking multiple bottleneck layers.

[0072] Similarly, the trained forging size detection model also includes a backbone network layer, a neck network layer, and a head network layer. The backbone network layer and the neck network layer have several cross-stage local network layers and several convolutional layers. The forging size detection model has improved the cross-stage local network layer, using deformable convolutions in the bottleneck layers of all cross-stage local network layers and adding channel and spatial attention mechanisms, and all convolutional layers in the backbone network layer and the neck network layer are generalized sparse convolutional layers, that is, the convolutions of the bottleneck layers in all cross-stage local network layers in the traditional YOLOv5s are replaced with deformable convolutions and adding channel and spatial attention mechanisms, and all convolutional layers are replaced with generalized sparse convolutional layers. In the head network layer, the structured intersection-over-union loss function is used as the bounding box loss function to obtain the intersection of the predicted box and the true box.

[0073] In this embodiment, the forging size detection model replaces the standard convolution (Conv) with a deformable convolution, thereby allowing the convolution kernel to dynamically learn the offset, thereby adapting to the shape changes of the forging.

[0074] In this embodiment, a channel and space attention mechanism is also introduced in the improved cross-stage local network layer. The attention mechanism can enhance the sensitivity and importance of the convolutional neural network to different channel features, and increase the sensitivity of the convolutional neural network to small differences. Optimizing the spatial dimension information can improve the performance of the convolutional neural network. Specifically, the channel and space attention mechanism includes a channel attention mechanism layer and a spatial attention mechanism layer connected in sequence. The attention mechanism mainly focuses on the feature fusion between channels of the convolution operation in the backbone network, and then learns the weight coefficient vector corresponding to each channel in the feature map, and then uses the vector to perform weighted processing on the feature map. Based on the consideration of improving the performance and training speed of the forging size detection model, in this embodiment, the original bounding box loss function of the traditional YOLOv5s is replaced with a structured intersection-over-union loss function to further improve the accuracy and training speed of the forging size detection model. The penalty index is redefined by SIoU to take into account the vector angle between the required regressions. The structured intersection-over-union loss function effectively reduces the number of degrees of freedom by adding an angle penalty term. The structured intersection-over-union loss function is used as the bounding box loss function to improve the convergence speed of the model during training and improve the accuracy of bounding box regression.

[0075] The cross-stage local network layer in the forging size detection model includes a first convolution layer, a bottleneck layer, a first tensor splicing layer, a second convolution layer and a third convolution layer connected in sequence. The input ends of the first convolution layer and the second convolution layer are used to receive input data of the cross-stage local network layer, the output end of the second convolution layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolution layer, and the third convolution layer is used to output the output data of the cross-stage local network layer.

[0076] Specifically, the cross-stage local network layer has two branches: the first branch is the second convolution layer that performs convolution processing on the input data; the second branch is the first convolution layer, the bottleneck layer, the first tensor splicing layer and the third convolution layer connected in sequence. After the input data is input into this branch, the first convolution layer performs feature extraction, and then it is input into the bottleneck layer to extract multi-scale features. Finally, the features obtained from the bottleneck layer and the first convolution layer are respectively spliced ​​in the first tensor splicing layer, and then the spliced ​​features are input into the third convolution layer for convolution processing to obtain the output data of the cross-stage local network layer.

[0077] Since deformable convolution requires additional learning of offsets during the convolution process, a branch network is required to predict the offset parameters. In this embodiment, the bottleneck layer in the forging size detection model adopts a residual network structure, and the bottleneck layer includes an identity mapping layer for jump connection, and a fourth convolution layer, a variable convolution layer, a channel attention mechanism layer, a spatial attention mechanism layer, and a feature addition layer, that is, a feature fusion layer, which are sequentially connected. The input ends of the fourth convolution layer and the identity mapping layer are used to receive input data of the bottleneck layer, and the output end of the identity mapping layer is connected to the output end of the feature addition layer, and the feature addition layer is used to output the output data of the bottleneck layer.

[0078] Specifically, the bottleneck layer has two branches: the first branch is the identity mapping layer that performs identity mapping processing on the input data, and the identity mapping layer is used to output the identity mapping feature; the second branch is the fourth convolution layer, the variable convolution layer, the channel attention mechanism layer, the spatial attention mechanism layer and the feature addition layer connected in sequence. After the input data is input into this branch, the fourth convolution layer and the variable convolution layer perform feature extraction, and then the extracted features are input into the channel attention mechanism layer and the spatial attention mechanism layer in turn for corresponding attention mechanism processing. Finally, the features obtained from the two branches are fused at the feature addition layer to obtain the output data of the bottleneck layer.

[0079] The generalized sparse convolutional layer in the forging size detection model includes a fifth convolutional layer, a depthwise separable convolutional layer, a second tensor splicing layer, and a shuffle layer. The input end of the fifth convolutional layer is used to receive the input data of the generalized sparse convolutional layer, the input end of the depthwise separable convolutional layer is connected to the output end of the fifth convolutional layer, and the output ends of the depthwise separable convolutional layer and the fifth convolutional layer are connected to the input end of the second tensor splicing layer, the output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolutional layer.

[0080] Specifically, the input data of the generalized sparse convolutional layer is first input into the fifth convolutional layer for convolution processing, and then the output data of the fifth convolutional layer is subjected to a lightweight convolution operation such as depthwise separable convolution. After that, the output data of the fifth convolutional layer and the depthwise separable convolutional layer are feature spliced ​​in the second tensor splicing layer, and then the spliced ​​features are input into the shuffling layer for shuffling processing to obtain the output data of the generalized sparse convolutional layer.

[0081] In this embodiment, the original standard convolution of the traditional YOLOv5s is replaced by a generalized sparse convolution layer. The generalized sparse convolution is a lightweight convolution operation that simulates the effect of the standard convolution operation by generating a small amount of redundant features, thereby reducing the computational cost. The number of output channels of the generalized sparse convolution is consistent with the original standard convolution to ensure compatibility with other modules.

[0082] Preferably, the specific configuration of the generalized sparse convolutional layer (such as the number of groups, kernel size, etc.) can be continuously adjusted during the model training process to achieve a balance between computational overhead and accuracy.

[0083] In this embodiment, in order to update the network parameters of the forging size detection model, the training of the forging size detection model uses a gradient descent algorithm to continuously optimize the network weights, such as a small batch gradient descent method, that is, using part of the data to calculate the gradient each time, and the historical gradient direction and the current gradient direction can be combined when updating the network parameters. At the same time, a data enhancement strategy can also be introduced during the training process to ensure the generalization ability of the forging size detection model for different forging states.

[0084] In order to improve the computing speed and memory utilization of the GPU, the mixed precision training method is adopted in the training of the forging size detection model. Through mixed precision training, the training efficiency can be improved and the batch size can be increased without reducing the accuracy of the forging size detection model.

[0085] In this embodiment, in order to evaluate the performance of the improved model, three evaluation indicators, namely accuracy, recall and average precision, are selected to evaluate the detection effect of the model, and then judge the actual prediction ability of the improved YOLOv5s network.

[0086] The device for detecting the size of forgings in a hot forming process based on machine vision of the present invention inputs the acquired image data of the forging to be detected into a trained forging size detection model, and obtains the size information and classification result of the forging to be detected output by the forging size detection model, wherein the forging size detection model is an improvement on the traditional YOLOv5s model, and the convolution layer in the bottleneck layer of the C3 layer of the traditional YOLOv5s model is replaced with a deformable convolution layer, thereby allowing the convolution kernel to dynamically learn the offset, so as to adapt to the shape change of the forging in the hot forming process, so as to enhance the adaptability of the feature map when the shape characteristics of the forging change greatly, and make the forging size modeling pay more attention to the target area, reduce the influence of noise data on the model, improve the perception ability of the model, and introduce the channel and spatial attention mechanism in the C3 layer to enhance the attention to the tiny deformation of the forging at every moment in the hot forming process, and secondly, the backbone network layer and the bottleneck layer of the C3 layer are used to detect the size of the forging in the hot forming process. All convolutional layers in the network layer are generalized sparse convolutional layers, which reduce the computational complexity and number of parameters of the model while maintaining the model performance, and use the structured intersection-over-union loss function as the bounding box loss function to improve the convergence speed of the model during training and the accuracy of bounding box regression. The size information and classification results finally predicted by the forging size detection model are more accurate. In addition, since it can pay attention to the various types of deformation characteristics of the forging billet, it solves the current problem of lack of accuracy and stability in the size detection of forgings in the hot forming process due to reliance on manual operation, as well as the problem of insufficient deformation accuracy of the billet in the hot forming process resulting in the forging size not meeting the standard and resulting in the scrapping of the forging, and meets the requirements for real-time online measurement of the forging billet size, so that the production of parts can meet the high-efficiency and high-precision production requirements, improves production efficiency and product quality, reduces production costs and resource waste, and has significant economic and social benefits.

[0087] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8 As shown, the electronic device may include: a processor 810 (processor), a communication interface 820 (Communications Interface), a memory 830 (memory) and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the following method: Acquire image data of the forging billet to be inspected; Inputting the image data into a trained forging size detection model to obtain size information and classification results of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the head network layer uses a structured intersection-over-union loss function as a bounding box loss function, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

[0088] It should be noted that the electronic device in this embodiment can be a server, a PC, or other devices in specific implementation, as long as its structure includes the following: Figure 8 The processor 810, communication interface 820, memory 830 and communication bus 840 shown in the figure, wherein the processor 810, communication interface 820, memory 830 communicate with each other through the communication bus 840, and the processor 810 can call the logic instructions in the memory 830 to execute the above method. This embodiment does not limit the specific implementation form of the electronic device.

[0089] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0090] Furthermore, an embodiment of the present invention discloses a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, and when the program instructions are executed by a computer, the computer can perform the methods provided by the above-mentioned method embodiments, for example, including: Acquire image data of the forging billet to be inspected; Inputting the image data into a trained forging size detection model to obtain size information and classification results of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the head network layer uses a structured intersection-over-union loss function as a bounding box loss function, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

[0091] On the other hand, an embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including: Acquire image data of the forging billet to be inspected; Inputting the image data into a trained forging size detection model to obtain size information and classification results of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the head network layer uses a structured intersection-over-union loss function as a bounding box loss function, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

[0092] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting the size of forgings in a hot forming process, characterized in that: The method comprises: Acquire image data of the forging billet to be inspected; Inputting the image data into a trained forging size detection model to obtain size information and classification results of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the head network layer uses a structured intersection-over-union loss function as a bounding box loss function, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

2. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The inputting of the image data into the trained forging size detection model to obtain the size information and classification result of the forging to be detected output by the forging size detection model specifically includes: Inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain a forging blank fusion feature map output by the backbone network layer; Inputting the forging blank fusion feature map into the neck network layer to perform feature fusion and feature enhancement processing on feature maps of different levels, and obtaining forging blank scale feature maps of three scales output by the neck network layer; All of the forging billet scale feature maps are input into the head network layer for target detection processing, and the size information and classification results of the forging billet to be detected are obtained by the head network layer.

3. The method for detecting the size of forgings in a hot forming process according to claim 2, characterized in that: The backbone network layer includes a focus feature fusion layer, a deep semantic extraction layer and a fast spatial pyramid pooling layer connected in sequence. The deep semantic extraction layer includes a plurality of generalized sparse convolutional layers and a cross-stage local network layer arranged between two adjacent generalized sparse convolutional layers. The first and last layers of the deep semantic extraction layer are both generalized sparse convolutional layers. The generalized sparse convolutional layers are arranged at intervals and two adjacent generalized sparse convolutional layers are connected through a cross-stage local network layer.

4. The method for detecting the size of a forging during hot forming process according to claim 3, characterized in that: The step of inputting the image data into the backbone network layer for feature extraction and fusion processing to obtain a forging blank fusion feature map output by the backbone network layer specifically includes: Inputting the image data into the focus feature fusion layer for channel expansion and downsampling processing to obtain a downsampled feature map output by the focus feature fusion layer; Inputting the downsampled feature map into the deep semantic extraction layer to extract the deep semantic information in the feature map, and obtaining a scale feature map output by the deep semantic extraction layer; The scale feature map is input into the fast spatial pyramid pooling layer for scale fusion processing to obtain the forging blank fusion feature map.

5. The method for detecting the size of forgings in a hot forming process according to claim 4, characterized in that: The step of inputting the downsampled feature map into the deep semantic extraction layer to extract the deep semantic information in the feature map to obtain the scale feature map output by the deep semantic extraction layer specifically includes: Inputting the downsampled feature map into the first generalized sparse convolution layer in the deep semantic extraction layer for generalized sparse convolution processing to obtain a first convolution feature output by the first generalized sparse convolution layer; Inputting the first convolutional feature into the cross-stage local network layer of the next layer connected to the generalized sparse convolutional layer to extract multi-scale features, and obtaining a second convolutional feature output by the cross-stage local network layer; The second convolution feature is input into the generalized sparse convolution layer of the next layer connected to the cross-stage local network layer for generalized sparse convolution processing to obtain the third convolution feature output by the generalized sparse convolution layer, until the third convolution feature is input into the last generalized sparse convolution layer to obtain the scale feature map output by the last generalized sparse convolution layer.

6. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The bottleneck layer includes an identity mapping layer, and a fourth convolution layer, a variable convolution layer, a channel attention mechanism layer, a spatial attention mechanism layer and a feature addition layer connected in sequence; The input ends of the fourth convolutional layer and the identity mapping layer are both used to receive input data of the bottleneck layer, the output end of the identity mapping layer is connected to the output end of the feature addition layer, and the feature addition layer is used to output the output data of the bottleneck layer.

7. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The generalized sparse convolution layer includes a fifth convolution layer, a depth-separable convolution layer, a second tensor concatenation layer, and a shuffle layer; The input end of the fifth convolutional layer is used to receive the input data of the generalized sparse convolutional layer, the input end of the depthwise separable convolutional layer is connected to the output end of the fifth convolutional layer, and the output ends of the depthwise separable convolutional layer and the fifth convolutional layer are both connected to the input end of the second tensor splicing layer, the output end of the second tensor splicing layer is connected to the input end of the shuffle layer, and the shuffle layer is used to output the output data of the generalized sparse convolutional layer.

8. The method for detecting the size of forgings in a hot forming process according to claim 1, characterized in that: The forging size detection model is trained by the following steps: Acquiring original image data of a sample forging, and preprocessing the original image data to obtain sample image data of the sample forging; Performing target labeling on the sample image data to obtain labeling information; The annotation information includes a size boundary box and a category label of the sample forging; The sample image data is used as input data for training, the annotation information is used as labels for training, machine learning is used for training, and the network weights are adjusted by gradient descent to obtain the forging size detection model for predicting the size information and classification results of the forging to be detected.

9. A system for detecting the size of forgings in a hot forming process, characterized in that: The system comprises: A data acquisition module, used to acquire image data of the forging to be inspected; A size detection module, used for inputting the image data into a trained forging size detection model to obtain the size information and classification result of the forging to be detected output by the forging size detection model; The forging size detection model is a neural network model based on the YOLOv5 framework, and the forging size detection model includes a backbone network layer, a neck network layer and a head network layer, and the backbone network layer and the neck network layer both have a number of cross-stage local network layers and a number of generalized sparse convolution layers, and the bottleneck layers in all cross-stage local network layers use deformable convolutions and increase channels and spatial attention mechanisms, and the head network layer uses a structured intersection-over-union loss function as a bounding box loss function, and the structured intersection-over-union loss function is used to determine a penalty index and use the penalty index to determine the vector angle between the required bounding box regressions; The cross-stage local network layer includes a first convolutional layer, a bottleneck layer, a first tensor concatenation layer, a second convolutional layer and a third convolutional layer which are sequentially connected; The input ends of the first convolutional layer and the second convolutional layer are both used to receive input data of the cross-stage local network layer, the output end of the second convolutional layer is connected to the input end of the first tensor splicing layer, the output end of the first tensor splicing layer is connected to the input end of the third convolutional layer, and the third convolutional layer is used to output the output data of the cross-stage local network layer.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting the size of a forging in a hot forming process as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Cross-domain hyperspectral image classification method based on space attention guidance variable convolution

    CN116310810A

  • PCB defect detection method for improving YOLOv5s

    CN118334398A

  • Remote sensing image target dynamic detection method based on multi-kernel

    CN119580116A

  • Three-dimensional lidar point cloud semantic segmentation method and apparatus based on deep learning

    WO2024130776A1