Ai-powered scrap steel material type identification method based on machine vision

WO2026200049A1PCT designated stage Publication Date: 2026-10-01ANSTEEL ENG TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/141650
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2025-12-11
Publication Date
2026-10-01

Smart Images

  • Figure CN2025141650_01102026_PF_FP_ABST
    Figure CN2025141650_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An AI-powered scrap steel material type identification method based on machine vision, which relates to the field of metallurgy and the field of image identification. The method comprises: acquiring image frame data recording scrap steel materials in a scrap steel unloading scenario; randomly dividing the image frame data into a training set and a verification set according to a proportion, so as to form a data set; establishing a target detection model on the basis of the data set, and using the data set to train the target detection model, so as to acquire a trained detection model; acquiring a real-time video stream of a scrap steel grading system, reading an image frame and using same as an instant on-site image, and inputting the instant on-site image into the trained detection model; and sending an output result of the trained detection model to the scrap steel grading system, such that the scrap steel grading system compares the output result with a standard library of scrap steel material types, so as to obtain a grading result. The method implements automatic scrap steel grading in a scrap steel unloading scenario, and performs positioning and identification on scrap steel material types by means of a trained target detection model, thereby improving the positioning accuracy and identification accuracy, reducing the manual labor intensity, and enhancing the inspection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

A machine vision-based AI method for scrap steel shape recognition Technical Field

[0001] This invention relates to the fields of metallurgy and image recognition, and in particular to an AI-based scrap steel pattern recognition method based on machine vision. Background Technology

[0002] As a green, environmentally friendly, and recyclable resource that can replace iron ore, scrap steel is playing an increasingly important role in steel smelting due to the country's stringent requirements for energy conservation, emission reduction, and sustainable development in steel enterprises. However, scrap steel quality inspection has always been a major challenge for enterprises. Conventional scrap steel quality inspection mainly relies on manual visual inspection, which requires personnel to climb onto scrap steel trucks to visually assess the quality of the scrap steel and, based on their personal experience, manually record and photograph the results. These results then need to be uploaded to the metering system and ERP for verification by designated personnel. The main difficulties currently are as follows:

[0003] 1. Inaccurate measurement - Scrap steel comes in various forms and is often piled up in a messy and mixed manner, making it difficult to accurately determine the grade of scrap steel and deduct impurities. Furthermore, the quality inspection results themselves lack a unified standard, which can easily lead to disputes between suppliers and buyers.

[0004] 2. Slow testing - Most steel mills only unload during the day shift, and personnel need to frequently get on and off the crane to take photos, which affects the unloading progress of the overhead crane and poses safety hazards.

[0005] 3. High cost - Each unloading port requires quality inspection personnel, and these personnel are highly susceptible to various external factors;

[0006] 4. Difficult to manage, chaotic on-site situation, inability to achieve real-time on-site monitoring, and manual data recording inevitably leads to errors. Summary of the Invention

[0007] The purpose of this invention is to provide an AI-based scrap steel material type recognition method based on machine vision, which solves the problem of inconsistent standards caused by current manual grading, reduces labor intensity, and improves inspection speed.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A machine vision-based AI scrap steel pattern recognition method, including

[0010] S1. Acquire image frame data of scrap steel in the scrap steel unloading scenario using an image acquisition device;

[0011] S2. Preprocess the image frame data containing scrap steel, label the scrap steel, hazardous materials, and miscellaneous materials, and form a dataset by randomly dividing it according to the proportion.

[0012] S3. Based on the YOLOv7 network, establish an object detection model according to the dataset, train the object detection model using the dataset, and obtain the trained detection model. The trained detection model includes detection model 1, detection model 2, and detection model 3. Detection model 1 is used to identify scrap steel types, detection model 2 is used to identify hazardous materials, and detection model 3 is used to identify miscellaneous items.

[0013] S4. Use the real-time video stream image frame reading of the scrap steel grading system to obtain the instantaneous on-site image and input it into the trained detection model;

[0014] S5. Send the output of the trained detection model to the scrap steel grading system and compare it with the scrap steel material type standard library to obtain the grading result.

[0015] In S2, the dataset includes a training set and a validation set. The training set includes a first training set for identifying scrap steel types, a second training set for training on hazardous materials identification, and a third training set for training on deductible debris identification. The first detection model, the second detection model, and the third detection model are trained using the effective information in the training sets, respectively.

[0016] The valid information includes basic image attributes and annotation information. Basic image attributes include file name, width, height, and bit depth. Annotation information includes first annotation information for indicating the type of scrap steel, second annotation information for indicating the category of hazardous materials, and third annotation information for indicating the category of miscellaneous items.

[0017] In S3, the YOLOv7 network includes an input module, a backbone network module, a neck network module, and a detection head network module. An instance segmentation head is added to the detection head network module, and the IoU loss function is changed to the DIoU loss function. The instance segmentation head consists of convolutional layers and upsampling layers. The DIoU loss function is used to optimize the position of the bounding box, and the formula is as follows: L DIoU =1-DIoU ③

[0018] In formulas ①-③, A represents the predicted value, B represents the actual value, and p represents b and b'. gt The Euclidean distance between them, where b represents the parameter for predicting the center coordinates. gt The parameter represents the center of the true target bounding box, and c represents the diagonal length of the minimum bounding rectangle of the two rectangles. L DIoU Let represent the loss function of DIoU.

[0019] In S5, the scrap steel grading system includes a scrap steel category analysis system, a hazardous materials identification system, and a deductible material analysis algorithm. The scrap steel category analysis system is used to give the overall scrap steel grading based on the type and number of scrap steel. The hazardous materials identification system is used to give safety and danger signals based on the number of hazardous materials counted. The deductible material analysis algorithm is used to give the weight of deductible materials based on the quantity of deductible materials and the size of the coordinate frame.

[0020] A computer device includes: at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform any of the machine vision-based AI scrap steel pattern recognition methods of claims 1-6.

[0021] A computer-readable storage medium storing computer instructions for causing a computer to perform any of the machine vision-based AI scrap steel pattern recognition methods of claims 1-6.

[0022] Compared with the prior art, the beneficial effects of the present invention are:

[0023] By combining machine vision with deep learning convolutional neural networks (which are trained detection models, including detection model 1, detection model 2, and detection model 3; detection model 1 is used to identify scrap steel types, detection model 2 is used to identify hazardous materials, and detection model 3 is used to identify miscellaneous items), automatic scrap steel grading was achieved in the unloading scenario at the scrap steel unloading site. The improved network model addresses the need to identify the position and size of the object's coordinate frame, improving positioning accuracy and recognition rate, reducing the labor intensity of employees, increasing inspection speed, and achieving intelligent, accurate, and efficient scrap steel grading process with high reliability. Attached Figure Description

[0024] Figure 1 is a schematic diagram of the deep learning model structure in an embodiment of the present invention.

[0025] Figure 2 is a flowchart of the unloading target detection method and unloading status analysis system in an embodiment of the present invention. Detailed Implementation

[0026] The present invention will now be described in detail with reference to the accompanying drawings, but it should be noted that the implementation of the present invention is not limited to the following embodiments.

[0027] The following embodiments are implemented based on the technical solution of the present invention, providing detailed implementation methods and specific operation processes. However, the scope of protection of the present invention is not limited to the following embodiments. Unless otherwise specified, the methods used in the following embodiments are conventional methods.

[0028] Example 1

[0029] See Figures 1 and 2. This paper presents an AI-based scrap steel type recognition method based on machine vision, addressing the scrap steel classification problem in existing scrap steel unloading scenarios. Considering the varying proportions of objects in complex backgrounds and the differences in object color and shape under different conditions, the training of the recognition algorithm requires a large number of data samples. Therefore, image acquisition is performed first, using an image acquisition device to collect 2000 images, including images of scrap steel, hazardous materials, and miscellaneous items. Since the selected images need to cover as many situations as possible, random cropping, stitching, and Mosaic techniques are used to augment the data and expand the sample range. Then, the LabelMe program is used to label the dataset, converting the obtained images and labels according to the format required by YOLO v7. The dataset is then randomly divided into training and validation sets in a 7:3 ratio. Considering that the YOLO v7 network function cannot meet the instance segmentation requirements, additional data is added to the YOLO network. The YOLO v7 network model is improved by adding an instance segmentation head to the detection head. This instance segmentation head consists of convolutional layers, pooling layers, and an upsampling module, aiming to restore the high-dimensional feature map to a feature map of a certain proportion of the original image size to complete mask classification. The DIoU loss function is used to optimize the position of the bounding boxes, accelerate the model convergence speed, improve accuracy, and thus improve the generalization ability and overall performance of the network model. A machine vision-based AI scrap steel material type recognition method is proposed. This method uses machine vision algorithms to collect and label datasets, constructs a dataset, and divides it into training and validation sets proportionally. The improved YOLO v7 model is trained on the prepared dataset, tested, and deployed to a designated device. Instance segmentation is performed to determine the material type category, and the scrap steel grading system performs category statistical comparison and grading rules to obtain the results. See Figure 1 for details.

[0030] S1. Acquire image frame data of scrap steel in the scrap steel unloading scenario using an image acquisition device;

[0031] S2. Preprocess the image frame data containing scrap steel, and label the scrap steel, hazardous materials, and miscellaneous materials. After randomly dividing the data proportionally, a dataset is formed. The image frames containing scrap steel are the parts of the video stream containing scrap steel. The dataset includes a training set and a validation set. The training set includes a first training set for identifying scrap steel types, a second training set for training hazardous materials identification, and a third training set for training miscellaneous material identification. The training set is a set of images in which designated objects are marked with bounding boxes. The images themselves, along with the coordinates of these bounding boxes, constitute the training set. The effective information in the training set is used to train detection model one, detection model two, and detection model three, respectively. Detection model one is used to identify scrap steel types, detection model two is used to identify hazardous materials, and detection model three is used to identify miscellaneous materials. The effective information includes basic image attributes and annotation information. Basic image attributes include file name, width, height, and bit depth. The annotation information includes first annotation information for identifying scrap steel types, second annotation information for identifying hazardous material categories, and third annotation information for identifying miscellaneous material categories.

[0032] S3. Based on the YOLOv7 network, establish an object detection model according to the dataset, and train the object detection model to obtain the trained detection model. The trained detection model includes detection model one for identifying scrap steel types, detection model two for identifying hazardous materials, and detection model three for identifying miscellaneous items. Use the processed dataset as input and perform calculations on the models starting from the input layer. The steps are as follows:

[0033] S31. The input layer will reshape the image into a 640*640*3 shape;

[0034] S32. Convolution calculation: perform dot product based on bits;

[0035] S33. Add the results together and output the dimensions. The calculation formula is as follows:

[0036] The trained detection model updates the values ​​of the weight parameters for each layer. Saving this model results in a weight file. When using it, you can call this file using PyTorch methods to load the network model and parameters, and then feed the input into the input layer to begin calculation.

[0037] YOLOv7 includes an input module, a backbone network module, a neck network module, and a detection head network module, as shown in Figure 1. The backbone network module consists of a feature extraction layer, a feature fusion layer, and a max pooling layer, used to extract multi-level features. The neck network module consists of feature extraction, feature fusion, feature concatenation, upsampling, a pooling spatial pyramid, and a multi-branch feature layer, used to collect features at multiple scales, enrich feature information, and expand the receptive field to make the extracted features closer to the original image range. Then, the features are processed more conveniently. The detection head network module outputs the results. The detection head network module consists of the detection head output and the instance segmentation head output, which output the extracted features. Given features in different dimensions, the detection head predicts the bounding box of an object and provides its category information, thus completing the object detection task. The instance segmentation head, using the given features, outputs masks for different objects, completing the semantic segmentation task. An instance segmentation head is added to the detection head network module. This head is used to change the IOU loss function to the DIOU loss function. The instance segmentation head consists of convolutional layers and upsampling layers. The convolutional layers are used for feature extraction, and the upsampling layers, also known as deconvolutional layers, are used to reduce the dimensionality of the high-dimensional feature map to recover the original image, commonly used in semantic segmentation. The DIOU loss function is used to optimize the bounding box position, as shown in the following formula: L DIoU =1-DIoU ③

[0038] In formulas ①-③, A represents the predicted value, B represents the actual value, and p represents b and b'. gt The Euclidean distance between them, where b represents the parameter for predicting the center coordinates. gt The parameter represents the center of the true target bounding box, and c represents the diagonal length of the minimum bounding rectangle of the two rectangles.

[0039] S4. Deploy the trained model on the server. An automatic capture algorithm, combined with a target detection algorithm, can locate the suction cup's adsorption position. Adjust the camera's PETZ value centered on the frame to ensure the focal length is increased while maintaining focus on the frame. Then, control the camera to capture images at an appropriate focal length and return to the preset point. Use the real-time video stream image frames from the scrap steel grading system to obtain instantaneous images of the unloading site. These images are of the scrap steel material inside the truck bed at the unloading point, representing the images that need to be graded during the unloading scenario. Input these images into the trained detection model. Use the trained detection model for instance segmentation. Return the output results to the scrap steel grading system. First, calculate the overall scrap steel grading result; second, determine the presence of hazardous materials; and finally, calculate the weight of any miscellaneous items.

[0040] S5. Send the output of the trained detection model to the scrap steel grading system and compare it with the scrap steel material type standard library to obtain the grading result.

[0041] The machine vision-based AI scrap steel type recognition system is a scrap steel grading system. This system includes a scrap steel category analysis system, a hazardous materials identification system, and a debris analysis algorithm. The scrap steel category analysis system classifies the scrap steel in the image based on its type and quantity. The hazardous materials identification system identifies safety and danger signals based on the number of hazardous materials. The debris analysis algorithm calculates the weight of the debris based on its quantity and bounding box size. The algorithm works as follows: A detection box is returned when a debris is detected. The detection box includes width and height; multiplying the width and height gives the detection box area. The areas of detection boxes of the same type are added together, and the overlapping portion is subtracted. Dividing this by the image size (1920*1080) gives the proportion of a certain type of debris. Multiplying this by the corresponding hyperparameter gives the weight. This process is repeated for all debris, and the weights are added together to obtain the final result.

[0042] The scrap steel grading system workflow includes scrap steel overall image grading, hazardous material identification, and deduction of miscellaneous materials calculation. Scrap steel overall image grading involves statistically analyzing the quantity of different types of scrap steel returned by the identification results and comparing them with a scrap steel type standard library to obtain the overall scrap steel grading result. The scrap steel type standard library contains scrap steel grading standards. For example, to determine if the steel in an image is heavy scrap (Category 1), the grading standard is that over 80% of the scrap steel types are classified as heavy scrap (Category 1). Hazardous material identification provides feedback by statistically analyzing the number of hazardous material identifications. No hazardous materials indicate safety; any hazardous material identification result is considered dangerous, triggering a hazardous material alarm. Deduction of miscellaneous materials calculation involves multiplying the area of ​​the coordinate frame of each deduction of miscellaneous material identification result by different deduction of miscellaneous material correction coefficients, summing the results, and dividing by the overall image area to obtain the deduction of miscellaneous material percentage. The weight of the deducted miscellaneous materials is then calculated by multiplying this by the deduction of miscellaneous material coefficient according to process standards.

[0043] This invention combines machine vision with deep learning convolutional neural networks (the deep learning convolutional neural network is a trained detection model, including detection model one, detection model two, and detection model three. Detection model one is used to identify the type of scrap steel, detection model two is used to identify hazardous materials, and detection model three is used to identify miscellaneous items) to achieve automatic scrap steel grading in the unloading scenario at the scrap steel unloading site. The improved network model addresses the need to identify the position and size of the object's coordinate frame, improving positioning accuracy and recognition accuracy, reducing the labor intensity of employees, increasing inspection speed, and achieving intelligent, accurate, and efficient scrap steel grading process with high reliability.

Claims

1. A machine vision-based AI-powered scrap steel material shape recognition method, characterized in that, include S1. Acquire image frame data of scrap steel in the scrap steel unloading scenario using an image acquisition device; S2. Preprocess the image frame data containing scrap steel, label the scrap steel, hazardous materials, and miscellaneous materials, and form a dataset by randomly dividing it according to the proportion. S3. Based on the YOLOv7 network, establish an object detection model according to the dataset, train the object detection model using the dataset, and obtain the trained detection model. The trained detection model includes detection model 1, detection model 2, and detection model 3. Detection model 1 is used to identify scrap steel types, detection model 2 is used to identify hazardous materials, and detection model 3 is used to identify miscellaneous items. S4. Use the real-time video stream image frame reading of the scrap steel grading system to obtain the instantaneous on-site image and input it into the trained detection model; S5. Send the output of the trained detection model to the scrap steel grading system and compare it with the scrap steel material type standard library to obtain the grading result.

2. The AI-based scrap steel shape recognition method based on machine vision according to claim 1, characterized in that, In S2, the dataset includes a training set and a validation set. The training set includes a first training set for identifying scrap steel types, a second training set for training on hazardous materials identification, and a third training set for training on deductible debris identification. The first detection model, the second detection model, and the third detection model are trained using the effective information in the training sets, respectively.

3. The AI-based scrap steel shape recognition method based on machine vision according to claim 2, characterized in that, The effective information includes basic image attributes and annotation information. The basic image attributes include file name, width, height, and bit depth. The annotation information includes first annotation information for indicating the type of scrap steel, second annotation information for indicating the category of hazardous materials, and third annotation information for indicating the category of miscellaneous items.

4. The AI ​​scrap steel shape recognition method based on machine vision according to claim 1, characterized in that, In S3, the YOLOv7 network includes an input module, a backbone network module, a neck network module, and a detection head network module. An instance segmentation head is added to the detection head network module, and the IoU loss function is changed to the DIoU loss function. The instance segmentation head consists of convolutional layers and upsampling layers. The DIoU loss function is used to optimize the position of the bounding box, and the formula is as follows: L DIoU =1-DIoU ③ In formulas ①-③, A represents the predicted value, B represents the actual value, and p represents b and b'. gt The Euclidean distance between them, where b represents the parameter for predicting the center coordinates. gt The parameter represents the center of the true target bounding box, and c represents the diagonal length of the minimum bounding rectangle of the two rectangles. L DIoU Let represent the loss function of DIoU.

5. The AI ​​scrap steel shape recognition method based on machine vision according to claim 1, characterized in that, In S5, the scrap steel grading system includes a scrap steel category analysis system, a hazardous materials identification system, and a deductible waste analysis algorithm. The scrap steel category analysis system is used to give the overall scrap steel grading based on the type and number of scrap steel. The hazardous materials identification system is used to give safety and danger signals based on the number of hazardous materials counted. The deductible waste analysis algorithm is used to give the weight of deductible waste based on the quantity of deductible waste and the size of the coordinate frame.

6. A computer device, characterized in that, include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the machine vision-based AI scrap steel pattern recognition method according to any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the AI ​​scrap steel pattern recognition method based on machine vision as described in any one of claims 1-5.