Steel mill material taking behavior identification method and system based on target detection

By combining the mask algorithm and the YOLOv5 target detection model with the CBAM attention unit, the real-time and computing resource issues of the steel plant's material picking behavior recognition are solved, and efficient and accurate material picking behavior recognition is achieved.

CN120599531AActive Publication Date: 2025-09-05UNIV OF SCI & TECH BEIJING
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510651363.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-05
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing steel mill material handling behavior recognition relies on manual features and traditional machine learning models, which are difficult to identify complex or forged behaviors. They have large computational complexity and slow inference speed, and are not suitable for steel mill environments with high real-time requirements.

Method used

A mask algorithm is used to process historical steelmaking site videos to generate images of interest. A target detection model is established based on the YOLOv5 architecture and trained in combination with the CBAM attention unit until the loss function value is less than the preset value, and the steel plant's material taking behavior recognition results are output.

Benefits of technology

It improves recognition accuracy and robustness, reduces the burden on computing resources, and enhances processing efficiency and deployment flexibility, making it suitable for industrial scenarios such as steel mills that have high real-time and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599531A_ABST
    Figure CN120599531A_ABST
Patent Text Reader

Abstract

The invention provides a steel mill material taking behavior recognition method and system based on target detection, and relates to the technical field of computer vision, and the method comprises the steps: collecting a historical steelmaking site video; processing the historical steelmaking site video through a mask algorithm to generate an interest image; yOLOv5 is used as a framework, a target detection model is established, the target detection model comprises a backbone network module, a neck module and a detection head module, and the backbone network module and the neck module each comprise a CBAM attention unit; inputting the interest image into a target detection model for training until the loss function value of the target detection model is smaller than a preset loss function value; obtaining a to-be-identified steelmaking site video; inputting a to-be-identified steelmaking site video into the trained target detection model, and outputting a detection result; and outputting a steel mill material taking behavior identification result according to a detection result. Safety accidents in a working place can be effectively prevented and reduced, and the life safety of workers is protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for identifying steel mill material taking behavior based on target detection. Background Art

[0002] Object detection is a computer vision technology used to identify the category and location (such as a rectangular box) of a target object in an image or video. Typical algorithms include YOLO and Faster R-CNN. Steel mill retrieving behavior refers to the process of measuring the composition of molten steel in the converter after it has been refined to a certain level. At this time, the converter's entrance faces outward, and workers use a longer retrieving rod to reach into the molten steel, scoop up a certain amount of molten steel, and remove it. After the molten steel condenses, it can be knocked off and its composition tested. Object detection-based steel mill retrieving behavior recognition utilizes image recognition techniques within artificial intelligence to intelligently analyze videos of the steelmaking site to identify whether the critical action of "retrieving" has occurred.

[0003] The development of the steel industry is closely linked to a country's industrial level. With the launch of "Made in China 2025," China's steel industry has entered the era of intelligent manufacturing, urgently requiring intelligent transformation and upgrading. Converter steelmaking is the first step in steelmaking, transforming iron into steel. Within this process, workers must perform the material removal operation, which significantly impacts the quality of the resulting steel. By measuring the removed portion of molten steel, the composition of the molten steel within the converter can be verified, thereby assessing its quality and determining the conditions for tapping a furnace. Therefore, identifying the operational behavior of the material removal operation is crucial.

[0004] However, existing methods for identifying material handling behaviors in steel mills rely on handcrafted features (such as LBP, SIFT, and HOG) and traditional machine learning models (such as SVM). While robust to local information, these methods struggle to identify complex or fabricated behaviors. While deep learning methods have improved recognition capabilities, they are computationally intensive and slow to infer, making them unsuitable for scenarios requiring high real-time performance. Especially in high-load environments like steel mills, the complexity of behavior recognition models makes them unsuitable for deployment, increasing the burden on computing resources and reducing recognition accuracy, making them difficult to meet the demands of real-time scenarios. Summary of the Invention

[0005] In order to solve the technical problem that the existing steel mill material picking behavior recognition relies on manual features (such as LBP, SIFT, HOG, etc.) and traditional machine learning models, which are difficult to identify complex or forged behaviors, have large computational load and slow inference speed, and are not suitable for scenarios with high real-time requirements. Especially in high-load environments such as steel mills, the behavior recognition model is highly complex and unsuitable for deployment, resulting in increased computing resource burden, reduced recognition accuracy, and difficulty in meeting scenarios with high real-time requirements, the present invention provides a steel mill material picking behavior recognition method and system based on target detection.

[0006] The technical solutions provided by the embodiments of the present invention are as follows:

[0007] First aspect:

[0008] An embodiment of the present invention provides a method for identifying material taking behavior in a steel plant based on target detection, comprising:

[0009] S1: Collect historical steelmaking site videos;

[0010] S2: Process the historical steelmaking site video through a masking algorithm to generate images of interest;

[0011] S3: Using YOLOv5 as the architecture, we build an object detection model. The object detection model includes a backbone network module, a neck module, and a detection head module. Both the backbone network module and the neck module contain CBAM attention units.

[0012] S4: Inputting the image of interest into the target detection model for training until the loss function value of the target detection model is less than the preset loss function value;

[0013] S5: Obtain the steelmaking site video to be identified;

[0014] S6: Input the steelmaking site video to be identified into the trained target detection model and output the detection results;

[0015] S7: Based on the detection results, output the steel plant's material taking behavior identification results.

[0016] Second aspect:

[0017] An embodiment of the present invention provides a steel mill material taking behavior recognition system based on target detection, comprising:

[0018] processor;

[0019] A memory stores computer-readable instructions, which, when executed by a processor, implement the steel plant material taking behavior recognition method based on target detection as in the first aspect.

[0020] The third aspect:

[0021] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for identifying material taking behavior in a steel plant based on target detection as described in the first aspect is implemented.

[0022] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0023] In an embodiment of the present invention, a mask algorithm is used to process historical steelmaking site videos to extract regions of interest in the image, thereby reducing the interference of other irrelevant targets on the model. Then, a target detection model is established using YOLOv5 as the architecture, and the image of interest is input into the target detection model for training until the loss function value of the target detection model is less than the preset loss function value, ensuring that the model can efficiently learn the core features. Finally, the steelmaking site video to be identified is obtained, and the steelmaking site video to be identified is input into the trained target detection model, and the detection results are output to complete the identification of the steel plant's material-taking behavior. Combined with a lightweight model architecture, it not only ensures recognition accuracy and robustness, reduces the burden of computing resources, but also improves processing efficiency and deployment flexibility. It is particularly suitable for industrial scenarios such as steel plants that have high real-time and resource requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0025] Figure 1 A schematic flow chart of a method for identifying material taking behavior in a steel plant based on target detection provided by an embodiment of the present invention;

[0026] Figure 2 A schematic diagram of the structure of CABM attention provided by an embodiment of the present invention;

[0027] Figure 3 A schematic structural diagram of a steel mill material taking behavior recognition system based on target detection provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0029] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0030] In the embodiments of the present invention, the terms "image" and "picture" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same. The terms "of," "corresponding," and "corresponding" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same.

[0031] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0032] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0033] Reference Manual Figure 1 , which shows a flow chart of a method for identifying material taking behavior in a steel plant based on target detection provided by an embodiment of the present invention.

[0034] The present invention provides a method for identifying steel mill reclaiming behavior based on target detection. This method can be implemented by a device for identifying steel mill reclaiming behavior based on target detection. The device can be a terminal or a server. The process flow of the method for identifying steel mill reclaiming behavior based on target detection can include the following steps:

[0035] S1: Collect historical steelmaking site videos.

[0036] Among them, historical steelmaking on-site videos refer to steelmaking operations that have occurred in the past, and refer to the real industrial production environment in which steelmaking processes are carried out in steel mills, which may include multiple visual targets such as workers, molten steel, and material handling equipment.

[0037] It's important to note that collecting historical steelmaking videos allows us to leverage the mill's existing monitoring resources, eliminating the need for additional monitoring and reducing data acquisition costs. Furthermore, these historical videos cover a variety of operating conditions and environmental changes, making the training data more diverse and rich, helping to improve the model's generalization capabilities.

[0038] S2: Use masking algorithms to process historical steelmaking site videos and generate images of interest.

[0039] The masking algorithm is an image processing method that constructs a "mask" region to retain only the portion of interest in the image and block irrelevant background information. It is commonly used in tasks such as target extraction and image segmentation. The image of interest is an image of the key area extracted from the original video frame, containing only the parts directly related to the "retrieving behavior", such as the worker, the retrieving rod, and the molten steel.

[0040] It should be noted that the mask algorithm can accurately extract the core visual areas related to the material picking behavior, which significantly improves the focus and efficiency of subsequent model training. The generation of images of interest can also reduce the labeling workload, improve data utilization efficiency, and enhance the model's adaptability in actual scenarios.

[0041] In a possible implementation, S2 specifically includes:

[0042] S201: De-frame the historical steelmaking site video and save it in the format of an image.

[0043] Among them, frame splitting refers to dividing the video into frames of static images along the time axis to facilitate subsequent processing and analysis at the image level.

[0044] S202: Filter the images using the pre-trained weighted object detection model YOLOv5 to determine worker images.

[0045] Among them, YOLOv5 refers to a mainstream target detection algorithm with the advantages of high precision and real-time performance, and is suitable for detecting objects and their positions in images.

[0046] S203: Perform random horizontal flipping and stretching on the worker images to generate a worker image dataset.

[0047] Among them, random horizontal flipping and stretching are ways to artificially change the image content to expand the data set and improve the generalization ability of the model.

[0048] S204: Label the worker image dataset to obtain a label file corresponding to each image in the worker image dataset.

[0049] S205: Extract the center coordinates of each detection box in the label file and normalize each center coordinate:

[0050]

[0051] Among them, x ij Indicates the horizontal coordinate of the upper left corner of the jth detection box in the i-th picture, y ijrepresents the vertical coordinate of the upper left corner of the jth detection box in the i-th picture, norm represents the normalization operation, and w ij Indicates the width of the jth detection box in the i-th picture, h ij Indicates the height of the jth detection box in the i-th picture, Represents the horizontal coordinate of the normalized detection box center c, Represents the vertical coordinate of the normalized detection box center c.

[0052] Specifically, the detection box represents the location of the target in the image, usually expressed in terms of upper left corner coordinates, width, and height.

[0053] S206: Perform K-means clustering on the normalized center coordinates:

[0054]

[0055] Among them, min means taking the minimum value, μ k represents the center coordinates of the kth cluster, k = 1, 2, ..., K, K represents the total number of clusters, z ij Represents the cluster label of the center point of each detection box, S k represents the center point set of the kth cluster, and (cx, cy) represents the normalized center coordinates of the detection box.

[0056] Specifically, K-means clustering is an unsupervised learning method that divides a group of points into K categories to minimize the differences between similar points and is used for spatial region classification.

[0057] S207: Generate a rectangular mask based on the clustering results and extract the main area of ​​the image to obtain the image of interest:

[0058]

[0059] I masked (u,v)=I(u,v)⊙M(u,v)

[0060] Among them, M k (u, v) represents the mask coverage area of ​​the kth cluster, u and v both represent the coordinates of the center point of the kth cluster, μ k,x represents the horizontal coordinate of the kth cluster center, w k represents the width of the kth cluster, W represents the image width, μ k,y Indicates the ordinate of the kth cluster center, h k represents the height of the kth cluster, H represents the image height, I masked (u,v) represents the image after the extraction area, I(u,v) represents the image before the extraction area, and M(u,v) represents the mask area M k The union of (u,v), ⊙ represents element-wise multiplication.

[0061] It's important to note that using a pre-trained model to screen worker images, combined with data augmentation and precise annotation, improves data diversity and representativeness. Normalization and K-means clustering make it easier for the model to learn spatial distribution patterns, while masking further compresses irrelevant information, improving training efficiency and reducing computational costs. This approach is particularly well-suited for complex environments like steel mills, where targets are concentrated but limited in number, and is a crucial data preparation step in achieving highly robust and real-time behavior recognition models.

[0062] In a possible implementation manner, the types marked in S204 specifically include: person, taker rod, and molten steel.

[0063] Reference Manual Figure 2 , which shows a structural diagram of CABM attention provided by an embodiment of the present invention.

[0064] Figure 2 In the figure, Channel Attention Module represents the channel attention mechanism, Spatial Attention Module represents the spatial attention mechanism, Input represents the input feature map, Average Pool represents average pooling, which is used to extract the global average information of the feature map, Max Pool represents maximum pooling, which is used to extract the most significant eigenvalues ​​in the feature map, Shared MLP represents shared multi-layer perceptron (fully connected neural network), which is used to perform nonlinear transformation on the pooled features and extract channel relationships, CONV represents convolution operation, which is used to extract attention weights in the spatial dimension, and Output represents the output of the module, that is, the feature map weighted by the attention mechanism.

[0065] Specifically, the spatial attention part performs average pooling and maximum pooling (along the channel dimension) on the input feature map to generate two 2D feature maps, which are then concatenated to form a new feature map. The new feature map is then sent to the convolution operation (CONV). After the convolution result passes through the Sigmoid activation function, a spatial attention map is generated (one weight for each pixel). Finally, this attention map is used to weight the input feature map pixel by pixel to obtain the spatial attention output.

[0066] It should be noted that the channel attention part performs global average pooling and maximum pooling (along the spatial dimension) on the input feature map respectively. Each operation will obtain a channel-level description vector. Then, the two pooling results are sent to the shared MLP network (two FC layers) for nonlinear transformation, and the MLP output is weighted fused (⊕). The attention weight of each channel is obtained by the Sigmoid activation function. Finally, these weights are used to weight the original input feature map channel by channel, and the output is a feature map with channel attention.

[0067] S3: Based on the YOLOv5 architecture, an object detection model is established. The object detection model includes: a backbone network module, a neck module, and a detection head module. Both the backbone network module and the neck module contain CBAM attention units.

[0068] YOLOv5 is an advanced single-stage object detection algorithm (You Only Look Once) with fast detection speed and high accuracy, making it suitable for real-time industrial vision applications. Object detection models are neural network structures used to identify the location (usually a rectangular box) and category of objects in an image. They are widely used in scenarios such as face detection and vehicle recognition.

[0069] Among them, the backbone network module is responsible for extracting basic image features, such as edges and textures, and is often used to extract deep semantic information.

[0070] The neck module, located between the backbone and the detection head, fuses feature information from different scales to enhance the model's adaptability to large and small objects (e.g., FPN and PAN structures). The detection head module outputs the final target category and location prediction results, typically including a classifier and a bounding box regressor.

[0071] Among them, the CBAM attention unit refers to a mechanism that integrates channel attention and spatial attention, which can improve the model's ability to focus on key areas and information.

[0072] It should be noted that an efficient and intelligent target detection model was constructed by introducing the YOLOv5 architecture and the CBAM attention mechanism. As a lightweight, single-stage detection algorithm, YOLOv5 has a simple structure and fast computational speed, making it ideal for applications such as steel mills, which have extremely high real-time requirements. The layered structure of the backbone network and neck module can effectively extract and fuse multi-scale information in the image, enabling the model to recognize overall movements while also focusing on detailed features. The addition of the CBAM attention unit further enhances the model's ability to recognize key image areas (such as workers' hands, the end point of the feeding rod, molten steel spots, etc.), reduces attention to redundant areas, and thus improves model accuracy and robustness.

[0073] In one possible implementation, the CBAM attention unit includes channel attention and spatial attention.

[0074] Channel attention is used to measure and adjust the importance of each channel (i.e., a certain type of semantic information in the feature map). For example, one channel may focus on detecting edges, while another channel focuses on color. Channel attention can enhance useful channels and suppress redundant or irrelevant channels.

[0075] Spatial attention focuses on adjusting the importance of different "positions" in an image, allowing the model to automatically determine "where" in the image is most worthy of attention. For example, in an image, a worker's hands and the reclaimer may be the only key points of the action. Spatial attention helps the model focus on these areas.

[0076] It's important to note that channel attention helps the model determine "what type of features to look for," while spatial attention helps the model determine "where to look." These two mechanisms work together to improve the model's ability to recognize objects in complex images. In steel mill scenarios, this mechanism allows the model to focus more on, for example, the "area where the material is being removed" rather than on background molten steel or equipment, significantly improving detection accuracy.

[0077] In one possible implementation, channel attention is used to generate channel weights. The generation of channel weights specifically includes:

[0078] S301: Calculate the global average pooling and maximum pooling of the feature map:

[0079]

[0080] F max =max p,g X(:;p,g)

[0081] Among them, F avg Represents the global average pooling result of the feature map, X represents the input feature map, p represents the pixel row index, g represents the pixel column index, the first ":" in X(:,:,p,g) represents processing of each input image, and the second ":" represents processing of each channel of the feature map.

[0082] It should be noted that the dimension of the input feature map is (b, c, p, g), b is the number of input images, and c is the number of channels of the feature map.

[0083] Specifically, global average pooling takes the average of all pixels in each channel, compressing the spatial dimensions while preserving the global semantic information between channels. Max pooling extracts the maximum response value from each channel, retaining the strongest feature expression in the channel and emphasizing salient areas.

[0084] S302: Input the global average pooling result and the maximum pooling result into the fully connected layer, process them through the ReLU activation function, and output the channel information of the feature map:

[0085]

[0086] in, represents the channel weight of the global average pooling result, Represents the channel weight of the maximum pooling result, W1 and W2 both represent the weight matrix of the fully connected layer, and ReLU represents the ReLU activation function.

[0087] Specifically, a fully connected layer is a layer structure in a neural network where all inputs are connected to all output neurons and is used to learn high-level nonlinear feature relationships.

[0088] S303: Generate the fusion channel weight of each channel through the Sigmoid activation function according to the channel information:

[0089]

[0090] X out =X⊙M c

[0091] Among them, M c represents the fusion channel weight, X out represents the output feature map after channel attention weighting, and σ represents the Sigmoid activation function.

[0092] Specifically, the fused channel weight is the final channel importance score generated by combining the channel weight information of average pooling and maximum pooling through the Sigmoid function.

[0093] It's important to note that global average pooling and max pooling are used to extract channel-level statistical features, which are then transformed through a fully connected layer and ReLU nonlinearity to generate channel information. Finally, a sigmoid function is used to fuse channel weights, achieving "weighted attention" for each channel in the feature map. This approach enables the neural network to self-learn and adjust the importance of different semantic features, allowing the network to no longer "treat all channels equally" but instead focus on discriminative feature dimensions.

[0094] In one possible implementation, spatial attention is used to generate spatial weights, and generating spatial weights specifically includes:

[0095] S304: Calculate the average pooling and maximum pooling of the channel dimension in the channel attention output feature map:

[0096]

[0097] F′ max =max c X out (:,c,:,:)

[0098] Among them, F′ avg Represents the average pooling result of the channel dimension, F′ max Represents the maximum pooling result of the channel dimension, c represents the cth channel, and C represents the total number of channels.

[0099] S305: Concatenate the average pooling result and the maximum pooling result of the channel dimension:

[0100] M s =W s *[F′ avg ||F′ max ]

[0101] Among them, M s represents the spatial weight, W s Represents the weight matrix of the fully connected layer, * represents the matrix multiplication operation, and || represents the concatenation operation.

[0102] S306: Generate spatial weights based on the splicing results through the sigmoid activation function:

[0103] M s =σ(M s )

[0104] X final =X out ⊙M s

[0105] Among them, X final Represents the final output feature map.

[0106] It should be noted that through channel-wise average and maximum pooling operations, the model can extract the average and maximum responses of each spatial location in the image across all channels, thereby capturing both global and local salient features. The concatenation operation fuses the two pooling results, preserving spatial information from different angles. The spatial weight map generated by the sigmoid function can guide the model to focus on areas in the image related to the target behavior (such as the worker's hands, material picking movements, etc.), suppressing background interference. Compared to relying solely on channel-wise attention, spatial attention can guide the model to focus on the location of the behavior in a more fine-grained manner, thereby improving recognition accuracy and positioning precision.

[0107] S4: Input the image of interest into the target detection model for training until the loss function value of the target detection model is less than the preset loss function value.

[0108] The loss function value is an indicator that measures the error between the model's prediction and the true label. The smaller the value, the better the model performance. Common examples include classification loss, regression loss, and confidence loss.

[0109] Among them, those skilled in the art can set the size of the preset loss function value according to actual conditions, and the present invention does not limit it.

[0110] It's important to note that by feeding carefully processed "images of interest" into the object detection model for training, we ensure that the model focuses on core areas related to behavior, significantly improving training efficiency and detection accuracy. The training process is guided by a loss function, ensuring that the model continuously optimizes with each learning cycle, continuously approaching the optimal parameter state. By setting a preset loss threshold, we can avoid overfitting, improve training efficiency, and ensure the optimal use of training resources.

[0111] In one possible implementation, the loss function value is calculated as follows:

[0112] L=λ box L box +λ cls L cls +λ obj L obj

[0113] Among them, L represents the loss function value, λ box represents the weight of the detection box regression loss, L box represents the detection box regression loss, λ cls Represents the weight of classification loss, L cls represents the classification loss, λ obj Represents the weight of confidence loss, L obj represents the confidence loss.

[0114] S5: Obtain the steelmaking site video to be identified.

[0115] S6: Input the steelmaking site video to be identified into the trained target detection model and output the detection results.

[0116] It's important to note that by feeding live or archived video into a trained object detection model, the system automatically identifies material handling behaviors and rapidly outputs detection results for key objects (such as workers, reclaiming rods, and molten steel), significantly reducing the burden of manual analysis. Compared to traditional methods that rely on manual review and monitoring, this approach is not only more efficient but also provides more objective and consistent results. By utilizing an optimized model trained on images of interest, the detection process exhibits strong generalization and adaptability, maintaining high accuracy even under varying working conditions, lighting conditions, and angles.

[0117] S7: Based on the detection results, output the steel plant's material taking behavior identification results.

[0118] Among them, the detection result is the recognition information output by the target detection model, which usually includes the target position (box coordinates) in each frame image, the target category (such as worker, molten steel, and material picking rod) and its confidence.

[0119] Among them, the recognition result is the final behavior judgment output, which is used to indicate whether the "material collection" behavior has occurred in the video, and can be used as the basis for monitoring alarms, recording or analysis.

[0120] It should be noted that by conducting logical analysis of the spatiotemporal relationship between multiple detection targets (such as whether the worker is close to the molten steel, whether he is holding the material-removing rod, etc.), the transition from "image-level recognition" to "behavior-level recognition" is achieved, which greatly improves the intelligence level and practical application value of the system.

[0121] In a possible implementation, S7 specifically includes:

[0122] Through the post-processing module, based on the class and location information in the detection results, it is determined whether a material removal operation has occurred:

[0123]

[0124] Among them, C t A flag indicating whether a material removal operation occurs at time t, D t-n Indicates whether the category and position of the tnth detection frame meet the requirements, N represents the total number of detection frames, Indicates the time of the t+1th material taking operation, represents the time of the t-th material removal operation, ΔT represents the time threshold, and π represents the product sign.

[0125] It should be noted that a combination of target detection and post-processing judgment is used to identify and judge the material taking action in the converter steelmaking process, and a simple model is used instead of the complex model of the behavior recognition algorithm to reduce the burden of on-site computing resources.

[0126] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0127] In an embodiment of the present invention, a mask algorithm is used to process historical steelmaking site videos to extract regions of interest in the image, thereby reducing the interference of other irrelevant targets on the model. Then, a target detection model is established using YOLOv5 as the architecture, and the image of interest is input into the target detection model for training until the loss function value of the target detection model is less than the preset loss function value, ensuring that the model can efficiently learn the core features. Finally, the steelmaking site video to be identified is obtained, and the steelmaking site video to be identified is input into the trained target detection model, and the detection results are output to complete the identification of the steel plant's material-taking behavior. Combined with a lightweight model architecture, it not only ensures recognition accuracy and robustness, reduces the burden of computing resources, but also improves processing efficiency and deployment flexibility. It is particularly suitable for industrial scenarios such as steel plants that have high real-time and resource requirements.

[0128] Reference Manual Figure 3 , which shows a structural schematic diagram of a steel plant material taking behavior recognition system based on target detection provided by the present invention.

[0129] The present invention further provides a steel mill material taking behavior recognition system 20 based on target detection, which is applied to the above-mentioned steel mill material taking behavior recognition method based on target detection, comprising:

[0130] Processor 201.

[0131] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the method for identifying material taking behavior in a steel plant based on target detection as described in the method embodiment is implemented.

[0132] The target detection-based steel mill material taking behavior identification system 20 provided by the present invention can execute the above-mentioned target detection-based steel mill material taking behavior identification method and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.

[0133] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0134] In an embodiment of the present invention, a mask algorithm is used to process historical steelmaking site videos to extract regions of interest in the image, thereby reducing the interference of other irrelevant targets on the model. Then, a target detection model is established using YOLOv5 as the architecture, and the image of interest is input into the target detection model for training until the loss function value of the target detection model is less than the preset loss function value, ensuring that the model can efficiently learn the core features. Finally, the steelmaking site video to be identified is obtained, and the steelmaking site video to be identified is input into the trained target detection model, and the detection results are output to complete the identification of the steel plant's material-taking behavior. Combined with a lightweight model architecture, it not only ensures recognition accuracy and robustness, reduces the burden of computing resources, but also improves processing efficiency and deployment flexibility. It is particularly suitable for industrial scenarios such as steel plants that have high real-time and resource requirements.

[0135] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), but may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0136] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0137] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function according to the embodiments of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available media can be magnetic media (such as floppy disks, hard disks, tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0138] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0139] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0140] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0141] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0142] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0143] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms.

[0144] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0145] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0146] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program codes.

[0147] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for identifying material taking behavior in a steel plant based on target detection as described in the method embodiment is implemented.

[0148] The computer-readable storage medium provided by the present invention can realize the steps and effects of the steel plant material taking behavior identification method based on target detection in the above method embodiment. To avoid repetition, the present invention will not go into details.

[0149] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0150] In an embodiment of the present invention, a mask algorithm is used to process historical steelmaking site videos to extract regions of interest in the image, thereby reducing the interference of other irrelevant targets on the model. Then, a target detection model is established using YOLOv5 as the architecture, and the image of interest is input into the target detection model for training until the loss function value of the target detection model is less than the preset loss function value, ensuring that the model can efficiently learn the core features. Finally, the steelmaking site video to be identified is obtained, and the steelmaking site video to be identified is input into the trained target detection model, and the detection results are output to complete the identification of the steel plant's material-taking behavior. Combined with a lightweight model architecture, it not only ensures recognition accuracy and robustness, reduces the burden of computing resources, but also improves processing efficiency and deployment flexibility. It is particularly suitable for industrial scenarios such as steel plants that have high real-time and resource requirements.

[0151] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

[0152] There are a few points to note:

[0153] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention. Other structures may refer to conventional designs.

[0154] (2) For the sake of clarity, the thickness of layers or regions in the drawings used to describe the embodiments of the present invention are exaggerated or reduced, that is, these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or intervening elements may be present.

[0155] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to form new embodiments.

[0156] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A method for identifying steel mill material taking behavior based on target detection, characterized in that: include: S1: Collect historical steelmaking site videos; S2: Processing the historical steelmaking scene video through a mask algorithm to generate an image of interest; S3: Using YOLOv5 as the architecture, establish a target detection model, which includes: a backbone network module, a neck module, and a detection head module, wherein the backbone network module and the neck module both include a CBAM attention unit; S4: inputting the image of interest into the target detection model for training until the loss function value of the target detection model is less than a preset loss function value; S5: Obtain the steelmaking site video to be identified; S6: Inputting the steelmaking site video to be identified into the trained target detection model and outputting the detection result; S7: Outputting the steel mill's material taking behavior identification result according to the detection result.

2. The method for identifying steel mill material taking behavior based on target detection according to claim 1, characterized in that: The S2 specifically includes: S201: De-frame the historical steelmaking site video and save it in a picture format; S202: Using the pre-trained weighted object detection model YOLOv5, the images are screened to determine worker images. S203: Performing random horizontal flipping and stretching processing on the worker images to generate a worker image dataset; S204: Labeling the worker image dataset to obtain a label file corresponding to each image in the worker image dataset; S205: Extract the center coordinates of each detection frame in the label file and normalize each center coordinate: Among them, x ij Indicates the horizontal coordinate of the upper left corner of the jth detection box in the i-th picture, y ij represents the vertical coordinate of the upper left corner of the jth detection box in the i-th picture, norm represents the normalization operation, and w ij Indicates the width of the jth detection box in the i-th picture, h ij Indicates the height of the jth detection box in the i-th picture, Represents the horizontal coordinate of the normalized detection box center c, Represents the vertical coordinate of the normalized detection box center c; S206: Perform K-means clustering on the normalized center coordinates: Among them, min means taking the minimum value, μ k represents the center coordinates of the kth cluster, k = 1, 2, ..., K, K represents the total number of clusters, z ij Represents the cluster label of the center point of each detection box, S k represents the center point set of the kth cluster, (cx,cy) represents the normalized center coordinates of the detection box; S207: Generate a rectangular mask based on the clustering result, and extract the main area of ​​the picture to obtain the image of interest: I masked (u,v)=I(u,v)⊙M(u,v); Among them, M k (u, v) represents the mask coverage area of ​​the kth cluster, u and v both represent the coordinates of the center point of the kth cluster, μ k,x represents the horizontal coordinate of the kth cluster center, w k represents the width of the kth cluster, W represents the image width, μ k,y Indicates the ordinate of the kth cluster center, h k represents the height of the kth cluster, H represents the image height, I masked (u,v) represents the image after the extraction area, I(u,v) represents the image before the extraction area, and M(u,v) represents the mask area M k The union of (u,v), ⊙ represents element-wise multiplication.

3. The steel mill material taking behavior recognition method based on target detection according to claim 2 is characterized in that: The types marked in S204 specifically include: people, take-up rods, and molten steel.

4. The steel mill material taking behavior recognition method based on target detection according to claim 1 is characterized in that: The CBAM attention unit includes channel attention and spatial attention.

5. The method for identifying steel mill material taking behavior based on target detection according to claim 4 is characterized in that: The channel attention is used to generate channel weights, and the generation of the channel weights specifically includes: S301: Calculate the global average pooling and maximum pooling of the feature map: F max =max p,g X(:;p,g); Among them, F avg Represents the global average pooling result of the feature map, X represents the input feature map, p represents the pixel row index, g represents the pixel column index, the first ":" in X(:,:,p,g) represents processing of each input image, and the second ":" represents processing of each channel of the feature map; S302: Input the global average pooling result and the maximum pooling result into the fully connected layer, process them through the ReLU activation function, and output the channel information of the feature map: in, represents the channel weight of the global average pooling result, Represents the channel weight of the maximum pooling result, W1 and W2 both represent the weight matrix of the fully connected layer, and ReLU represents the ReLU activation function; S303: Generate the fusion channel weight of each channel according to the channel information through the Sigmoid activation function: Among them, M c represents the fusion channel weight, X out represents the output feature map after channel attention weighting, and σ represents the Sigmoid activation function.

6. The method for identifying steel mill material taking behavior based on target detection according to claim 4, characterized in that: The spatial attention is used to generate a spatial weight, and generating the spatial weight specifically includes: S304: Calculate the average pooling and maximum pooling of the channel dimension in the channel attention output feature map: Among them, F′ avg Represents the average pooling result of the channel dimension, F′ max represents the maximum pooling result of the channel dimension, c represents the cth channel, and C represents the total number of channels; S305: Concatenate the average pooling result and the maximum pooling result of the channel dimension: M s =W s *[F′ avg ||F′ max ]; Among them, M s represents the spatial weight, W s Represents the weight matrix of the fully connected layer, * represents the matrix multiplication operation, and || represents the concatenation operation; S306: Generate spatial weights based on the splicing results through the sigmoid activation function: Among them, X final Represents the final output feature map.

7. The method for identifying steel mill material taking behavior based on target detection according to claim 1, characterized in that: The calculation formula of the loss function value is specifically: L=λ box L box +λ cls L cls +λ obj L obj ; Among them, L represents the loss function value, λ box represents the weight of the detection box regression loss, L box represents the detection box regression loss, λ cls Represents the weight of classification loss, L cls represents the classification loss, λ obj Represents the weight of confidence loss, L obj represents the confidence loss.

8. The method for identifying steel mill material taking behavior based on target detection according to claim 1, characterized in that: The S7 is specifically: Through the post-processing module, based on the information of the class and position in the detection results, it is determined whether a material removal operation has occurred: Among them, C t A flag indicating whether a material removal operation occurs at time t, D t-n Indicates whether the category and position of the tnth detection frame meet the requirements, N represents the total number of detection frames, Indicates the time of the t+1th material taking operation, represents the time of the t-th material removal operation, ΔT represents the time threshold, and π represents the product sign.

9. A steel mill material taking behavior recognition system based on target detection, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method for identifying material taking behavior in a steel plant based on target detection as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for identifying material taking behavior in a steel plant based on target detection as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Roadside construction behavior detection method based on improved YOLOv5 model

    CN115359392A

  • Video behavior recognition method and device, equipment and storage medium

    CN115761559A

  • Video pedestrian detection method based on machine learning

    CN116844017A

  • Target detection method for image collected by AR wearable device based on improved YOLOv8

    CN119942059A

  • Method for detecting defect in top cover of hydro turbine based on improved yolov8 model

    US20250131553A1