Material pile height detection method, device, equipment, medium and computer program product

Through deep learning technology, the material pile target detection model is trained and the material pile video data is detected in real time, which solves the problems of inefficiency, insufficient accuracy and safety hazards of traditional material pile height detection methods, and achieves efficient and accurate material pile height detection.

CN120070924APending Publication Date: 2025-05-30SHENHUA HUANGHUA PORT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133740.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional material pile height detection methods are inefficient, insufficient accuracy and safety hazards.

Method used

By obtaining multi-scale images of the material pile, performing similar deletion processing and material pile area annotation, the material pile object detection model is obtained, the material pile video data is detected in real time, the bounding box information is obtained, and the height exceeds the safety threshold is determined, and the corresponding instructions are issued.

Benefits of technology

It improves the efficiency and accuracy of the height detection of material piles, reduces labor costs, and reduces safety accidents caused by human negligence or environmental factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070924A_ABST
    Figure CN120070924A_ABST
Patent Text Reader

Abstract

The invention relates to a material pile height detection method and device, equipment, a storage medium and a computer program product, and the method comprises the steps: obtaining a multi-scale image set of a material pile, carrying out the similar deletion processing of the multi-scale image set, obtaining a pre-processed image set, carrying out the material pile region labeling processing of the pre-processed image set, and obtaining a material pile height detection result; the method comprises the steps of obtaining training data, training by utilizing the training data to obtain a material pile target detection model, obtaining material pile video data in real time, detecting the material pile video data in real time by utilizing the material pile target detection model to obtain material pile bounding box information, and obtaining the height of an upper frame according to the material pile bounding box information. Whether the height of the upper frame is larger than a preset safety height or not is judged, if the height of the upper frame is smaller than or equal to the safety height, the step of detecting the material pile video data in real time to obtain the material pile bounding box information is returned, and if the height of the upper frame is larger than the safety height, a material pile operation stopping instruction is sent out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning technology, and particularly to a method, device, equipment, storage medium and computer program product for detecting the height of a material pile. Background Art

[0002] In modern industrial production, material pile management is a crucial task, and its height detection directly affects the safety, efficiency and resource utilization rate of the production process. At present, traditional methods for detecting the height of a material pile mostly rely on manual regular inspections or simple mechanical measuring devices. These methods not only have a large labor intensity and low efficiency, but are also easily affected by human errors and environmental factors, resulting in inaccurate detection results and posing significant safety hazards. Especially in complex and changeable industrial environments, such as adverse conditions like dust, rain and fog, and at night, traditional detection means are even more inadequate.

[0003] With the rapid development of artificial intelligence and machine vision technologies, deep learning technology has shown great application potential in the field of industrial automation due to its powerful image recognition and processing capabilities. By applying deep learning to the detection of the height of a material pile, functions such as real-time detection, automatic analysis, and intelligent warning can be achieved, greatly improving the accuracy and efficiency of detection, reducing labor costs, and effectively reducing safety accidents caused by human negligence or environmental factors. Summary of the Invention

[0004] The present disclosure provides a method, device, equipment, storage medium and computer program product for detecting the height of a material pile to solve the problems of low efficiency, insufficient accuracy and safety hazards existing in traditional methods for detecting the height of a material pile.

[0005] In a first aspect, the present disclosure provides a method for detecting the height of a material pile, including:

[0006] Obtaining multi-scale images of the material pile to obtain a multi-scale image set, and performing similarity deletion processing on the multi-scale image set to obtain a preprocessed image set;

[0007] Performing material pile area annotation processing on the preprocessed image set to obtain training data, and training a material pile target detection model using the training data;

[0008] Real-time obtaining material pile video data, and performing real-time detection on the material pile video data using the material pile target detection model to obtain material pile bounding box information;

[0009] Obtaining the height of the upper border according to the material pile bounding box information;

[0010] Judging whether the height of the upper border is greater than a preset safety height;

[0011] If the height of the upper border is less than or equal to the safety height, return the step of obtaining the real-time stockpile video data and using the stockpile target detection model to perform real-time detection on the stockpile video data to obtain the stockpile bounding box information;

[0012] If the height of the upper border is greater than the safety height, issue a stop instruction for the stockpile operation.

[0013] In some embodiments, the processing of the multi-scale image set to perform similarity deletion to obtain the preprocessed image set includes:

[0014] Select images from the multi-scale image set according to a preset interval to obtain a selected image set;

[0015] Calculate the similarity between two adjacent images in the selected image set;

[0016] If the similarity between two adjacent images is greater than or equal to a preset similarity threshold, obtain the image information of the first image in the two adjacent images to obtain a deleted image information set;

[0017] Delete the corresponding images in the multi-scale image set according to the deleted image information set to obtain the preprocessed image set.

[0018] In some embodiments, the training of the stockpile target detection model using the training data includes:

[0019] Perform convolution processing on the preprocessed image set to obtain a convolution feature set;

[0020] Perform normalization processing on the convolution feature set to obtain a normalized feature set;

[0021] Use a preset activation function to introduce a non-linear factor to the normalized feature set to obtain activation features, where the activation function is represented by the following formula:

[0022] SiLU(x) = x · 1 / (1 + e -x )

[0023] where SiLU(x) is the activation function, x is the normalized feature set, and e is the base of the natural logarithm;

[0024] Use a pre-established YOLOv8 model to output training detection results according to the activation features;

[0025] Calculate the bounding box regression loss and classification loss between the training detection results and the standard annotation of the training data;

[0026] After adjusting the weight parameters and learning rate of the YOLOv8 model according to the bounding box regression loss and the classification loss, return the step of outputting the training detection result according to the activation features by using the pre-established YOLOv8 model;

[0027] When the preset number of iterations is reached, confirm that the model training is completed to obtain a stockpile target detection model.

[0028] In some embodiments, calculating the bounding box regression loss and the classification loss between the training detection result and the standard annotation of the training data includes:

[0029] Calculate the bounding box regression loss by using the following formula:

[0030]

[0031] where, L bbox is the bounding box regression loss, N is the number of samples, x i is the standard annotation corresponding to the i-th data in the training data, is the i-th detection result in the training detection result;

[0032] Calculate the classification loss by using the following formula:

[0033]

[0034] where, L cls is the classification loss, M is the number of classification types, c is the type number, y o,c is a binary function, if the sample o belongs to the type c, then output 1, otherwise output 0, p o represents the probability that the sample o belongs to the type c.

[0035] In some embodiments, using the stockpile target detection model to perform real-time detection on the stockpile video data to obtain stockpile bounding box information includes:

[0036] Use the stockpile target detection model to perform convolution processing and normalization processing on each frame of image in the stockpile video data to obtain stockpile image features;

[0037] Split the stockpile image features to obtain first stockpile image features and second stockpile image features;

[0038] Perform convolution splicing processing on the second stockpile image features to obtain convolution splicing features, and fuse the first stockpile image features and the convolution splicing features to obtain fusion features;

[0039] Extract pooling features of different scales from the fusion features to obtain pooling features of different scales;

[0040] Concatenate the pooling features of different scales in the channel dimension to obtain multi-scale features;

[0041] Use the stockpile target detection model to output stockpile bounding box data based on the multi-scale features.

[0042] In some embodiments, obtaining the upper border height according to the stockpile bounding box information includes:

[0043] Calculate the upper border height using the following formula:

[0044]

[0045] where y max is the upper border height, y c is the central point pixel height included in the stockpile bounding box information, and h is the bounding box height included in the stockpile bounding box information.

[0046] In a second aspect, the present disclosure provides a stockpile height detection device, including:

[0047] An image processing module, configured to obtain multi-scale images of a stockpile, obtain a multi-scale image set, perform similarity deletion processing on the multi-scale image set to obtain a preprocessed image set;

[0048] A model training module, configured to perform stockpile area annotation processing on the preprocessed image set to obtain training data, and use the training data to train a stockpile target detection model;

[0049] A height detection module, configured to obtain stockpile video data in real time, perform real-time detection on the stockpile video data using the stockpile target detection model to obtain stockpile bounding box information, and obtain the upper border height according to the stockpile bounding box information;

[0050] A height judgment module, configured to judge whether the upper border height is greater than a preset safety height. If the upper border height is less than or equal to the safety height, return to the step of obtaining the stockpile video data in real time and performing real-time detection on the stockpile video data using the stockpile target detection model to obtain stockpile bounding box information. If the upper border height is greater than the safety height, issue a stop stockpile operation instruction.

[0051] In a third aspect, the present disclosure provides a computer device, including a memory, a processor, and a computer program product stored on the memory, and the processor executes the computer program product to implement the steps of the method described in the above aspect.

[0052] Fourthly, the present disclosure provides a computer-readable storage medium, on which a computer program product is stored. When the computer program product is executed by a processor, the steps of the method described in the above aspects are implemented.

[0053] Fifthly, the present disclosure provides a computer program product, including a computer program product. When the computer program product is executed by a processor, the steps of the method described in the above aspects are implemented.

[0054] A method, device, equipment, storage medium and computer program product for detecting the height of a material pile provided by the present disclosure. By acquiring multi-scale images of the material pile to obtain a multi-scale image set, performing similarity deletion processing on the multi-scale image set to obtain a preprocessed image set, performing material pile area annotation processing on the preprocessed image set to obtain training data, training a material pile target detection model using the training data, acquiring material pile video data in real time, using the material pile target detection model to perform real-time detection on the material pile video data to obtain material pile bounding box information, obtaining the upper border height according to the material pile bounding box information, and determining whether the upper border height is greater than a preset safety height. If the upper border height is less than or equal to the safety height, then return to the step of acquiring the material pile video data in real time and using the material pile target detection model to perform real-time detection on the material pile video data to obtain material pile bounding box information. If the upper border height is greater than the safety height, then issue a stop instruction for the material pile operation, thereby solving the problems of low efficiency, insufficient accuracy and potential safety hazards of traditional material pile height detection methods, and improving the efficiency and accuracy of material pile height detection.

[0055] 1. Introduce non-linearity into the standardized feature set using a preset activation function to obtain activation features, and the technical effect is to avoid gradient disappearance;

[0056] 2. Perform convolutional splicing processing on the second material pile image features to obtain convolutional splicing features, and fuse the first material pile image features with the convolutional splicing features to obtain fusion features, and the technical effect is to enable the model to utilize multi-scale information;

[0057] 3. Extract pooling features of different scales of the fusion features to obtain pooling features of different scales, and splice the pooling features of different scales in the channel dimension to obtain multi-scale features, and the technical effect is to enrich the feature representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Hereinafter, the present disclosure will be described in more detail based on embodiments and with reference to the drawings:

[0059] Figure 1 It is a schematic flowchart of a method for detecting the height of a material pile provided by an embodiment of the present disclosure;

[0060] Figure 2 It is a functional module diagram of a stockpile height detection device provided by an embodiment of the present disclosure.

[0061] In the accompanying drawings, the same components are denoted by the same reference numerals, and the drawings are not drawn to actual scale. Detailed implementation manners

[0062] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand how the present disclosure uses technical means to solve technical problems and the implementation process of achieving corresponding technical effects and implement accordingly, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments of the present disclosure and each feature in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0063] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0064] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0065] Example 1

[0066] Figure 1 It is a flowchart of a stockpile height detection method provided by an embodiment of the present disclosure. As Figure 1 shown, a stockpile height detection method includes:

[0067] S1. Obtain multi-scale images of the stockpile to get a multi-scale image set, and perform similarity deletion processing on the multi-scale image set to obtain a preprocessed image set.

[0068] In an embodiment of the present invention, the obtaining of the multi-scale images of the stockpile means using a camera to obtain images of the stockpile under different heights, different distances, different positions, and different lighting conditions to ensure the diversity of image data.

[0069] In an embodiment of the present invention, the camera is installed on the gantry and has an adjustable camera bracket, and the camera bracket has a fine-tuning screw.

[0070] In an embodiment of the present invention, the performing of similarity deletion processing on the multi-scale image set to obtain a preprocessed image set includes:

[0071] Select images from the multi-scale image set according to a preset interval to obtain a selected image set;

[0072] Calculate the similarity between two adjacent images in the selected image set;

[0073] If the similarity between two adjacent images is greater than or equal to a preset similarity threshold, obtain the image information of the first image in the two adjacent images to obtain a deleted image information set;

[0074] Delete the corresponding images in the multi-scale image set according to the deleted image information set to obtain a preprocessed image set.

[0075] In an embodiment of the present invention, the calculating of the similarity between two adjacent images in the selected image set is to calculate the similarity between two images using the similarity calculation function in the OpenCV library. The similarity calculation function can be cv2.absdiff().

[0076] In an embodiment of the present invention, by obtaining multi-scale images of the stockpile to get a multi-scale image set, and performing similarity deletion processing on the multi-scale image set to obtain a preprocessed image set, the efficiency of subsequent annotation processing is improved.

[0077] S2. Perform stockpile area annotation processing on the preprocessed image set to obtain training data, and use the training data to train a stockpile target detection model.

[0078] In an embodiment of the present invention, the performing of stockpile area annotation processing on the preprocessed image set to obtain training data is to use an annotation tool to frame the stockpile area of the images in the preprocessed image set.

[0079] In the embodiments of the present invention, the stockpile target detection model is a YOLO series model, which can be a YOLOv8 model.

[0080] Specifically, the YOLOv8 model is a real-time target detection model.

[0081] In the embodiments of the present invention, the network structure of the stockpile target detection model includes three parts: a Backbone main network, a Neck neck network, and a Head head network. The Backbone main network uses a Cf2 module as the basic unit.

[0082] In the embodiments of the present invention, the stockpile target detection model includes a Conv Block module, and the ConvBlock module is the most basic module of the stockpile target detection model.

[0083] Furthermore, the Conv Block module includes a two-dimensional convolutional layer, a two-dimensional batch normalization layer, and an activation layer.

[0084] In the embodiments of the present invention, training the stockpile target detection model using the training data includes:

[0085] Performing convolutional processing on the preprocessed image set to obtain a convolutional feature set;

[0086] Performing normalization processing on the convolutional feature set to obtain a normalized feature set;

[0087] Introducing a non-linear factor for the normalized feature set using a preset activation function to obtain activation features, where the activation function is represented by the following formula:

[0088] SiLU(x) = x·1 / (1 + e -x )

[0089] where SiLU(x) is the activation function, x is the normalized feature set, and e is the base of the natural logarithm;

[0090] Using a pre-established YOLOv8 model to output training detection results based on the activation features;

[0091] Calculating the bounding box regression loss and classification loss between the training detection results and the standard annotations of the training data;

[0092] After adjusting the weight parameters and learning rate of the YOLOv8 model according to the bounding box regression loss and the classification loss, return to the step of using the pre-established YOLOv8 model to output training detection results based on the activation features;

[0093] After reaching the preset number of iterations, confirm that the model training is completed to obtain a stockpile target detection model.

[0094] In the embodiment of the present invention, the convolution process is to slide (or convolve) a convolution kernel on an image and calculate the dot product (sum of element products) of the convolution kernel and the local region of the image at each position. This process generates a new two-dimensional array, that is, convolution features.

[0095] Specifically, the bounding box regression loss is used to calculate the difference between the predicted bounding box and the ground truth bounding box. Mean Squared Error (MSE) is a commonly used loss function, which gives a higher penalty for larger errors, and this helps the model quickly correct large prediction errors.

[0096] Specifically, the classification loss is used to measure the difference between the class distribution predicted by the model and the ground truth label. Using the cross-entropy loss as the loss function, this function gives a higher penalty for incorrect predictions. The classification loss can help the model optimize its predictions in classification problems and make the predicted probability distribution as close as possible to the true label distribution.

[0097] In the embodiment of the present invention, calculating the bounding box regression loss and the classification loss between the training detection result and the standard annotation of the training data includes:

[0098] Calculate the bounding box regression loss using the following formula:

[0099]

[0100] where, L bbox is the bounding box regression loss, N is the number of samples, x i is the standard annotation corresponding to the i-th data in the training data, is the i-th detection result in the training detection result;

[0101] Calculate the classification loss using the following formula:

[0102]

[0103] where, L cls is the classification loss, M is the number of classification types, c is the type number, y o,c is a binary function, if the sample o belongs to the type c, then output 1, otherwise output 0, p o represents the probability that the sample o belongs to the type c.

[0104] In the embodiment of the present invention, by performing stockpile area annotation processing on the preprocessed image set to obtain training data, and training a stockpile target detection model using the training data, the efficiency of subsequent real-time detection of the stockpile is improved.

[0105] S3. Obtain the stockpile video data in real time, and use the stockpile target detection model to perform real-time detection on the stockpile video data to obtain the stockpile bounding box information.

[0106] In the embodiment of the present invention, the obtaining the stockpile video data in real time is to obtain the video data of the stockpile by using the camera.

[0107] In the embodiment of the present invention, the using the stockpile target detection model to perform real-time detection on the stockpile video data to obtain the stockpile bounding box information includes:

[0108] Use the stockpile target detection model to perform convolution processing and normalization processing on each frame of image in the stockpile video data to obtain the stockpile image features;

[0109] Split the stockpile image features to obtain the first stockpile image features and the second stockpile image features;

[0110] Perform convolution splicing processing on the second stockpile image features to obtain convolution splicing features, and fuse the first stockpile image features with the convolution splicing features to obtain fusion features;

[0111] Extract the pooling features of different scales of the fusion features to obtain the pooling features of different scales;

[0112] Splice the pooling features of different scales in the channel dimension to obtain multi-scale features;

[0113] Use the stockpile target detection model to output the stockpile bounding box data according to the multi-scale features.

[0114] Specifically, the performing convolution splicing processing on the second stockpile image features to obtain convolution splicing features is that the Bottleneck Block module adds a skip connection between the convolutional layers, directly connects the input to the output, and splices the output after the skip connection with the output of the convolutional layer to form the final output, which can increase the network depth while alleviating the problem of gradient disappearance.

[0115] Specifically, the Bottleneck Block module is a commonly used building block in deep convolutional neural networks, especially in ResNet (Residual Network). The design purpose of this module is to reduce the amount of computation and the number of parameters while increasing the network depth, thereby improving the efficiency and performance of the network.

[0116] Specifically, the pooling features of different scales of the fusion feature are extracted to obtain multi-scale features by performing pooling operations on the fusion feature using three MaxPool2d layers with different-sized convolutional kernels. The MaxPool2d layer is a commonly used pooling layer in convolutional neural networks. Its function is to extract the most important features from the feature map and reduce the spatial dimensions (height and width) of the feature map, thereby reducing the number of parameters and computational amount of subsequent layers, helping to prevent overfitting, and improving the generalization ability of the model.

[0117] In the embodiment of the present invention, by obtaining the stockpile video data in real time and using the stockpile target detection model to perform real-time detection on the stockpile video data, the stockpile bounding box information is obtained, which improves the efficiency and accuracy of stockpile height detection.

[0118] S4. Obtain the upper border height according to the stockpile bounding box information.

[0119] In the embodiment of the present invention, the upper border height refers to the height from the top of the stockpile to the bottom in the stockpile video data.

[0120] In the embodiment of the present invention, obtaining the upper border height according to the stockpile bounding box information includes:

[0121] Calculate the upper border height using the following formula:

[0122]

[0123] where y max is the upper border height, y c is the center point pixel height included in the stockpile bounding box information, and h is the bounding box height included in the stockpile bounding box information.

[0124] In the embodiment of the present invention, by obtaining the upper border height according to the stockpile bounding box information, the efficiency of subsequent determination of whether the upper border height is greater than the preset safety height is improved.

[0125] S5. Determine whether the upper border height is greater than the preset safety height.

[0126] In the embodiment of the present invention, the preset safety height can be half of the pixel height of the stockpile video data.

[0127] If the upper border height is less than or equal to the safety height, return to S3. Obtain the stockpile video data in real time and use the stockpile target detection model to perform real-time detection on the stockpile video data to obtain the stockpile bounding box information.

[0128] In an embodiment of the present invention, when the height of the upper border is less than or equal to the safety height, it indicates that the height of the material pile is in a safe state. Therefore, it is necessary to return to the step of real-time detection of the height of the material pile to achieve continuous monitoring of the height of the material pile.

[0129] If the height of the upper border is greater than the safety height, then execute S6, and issue an instruction to stop the material pile operation.

[0130] In an embodiment of the present invention, when the height of the upper border is greater than the safety height, it indicates that the height of the material pile is too high at this time, posing a safety hazard. Therefore, it is necessary to issue an instruction to stop the material pile operation to ensure production safety.

[0131] Example 2

[0132] Based on the above embodiments, this embodiment provides an application example.

[0133] S01, perform internal parameter calibration of the camera

[0134] In an embodiment of the present invention, a dedicated checkerboard is placed at different positions to obtain some images of the camera; then, based on these images, the calibration tool provided by OpenCV is used to obtain the internal parameters of the camera.

[0135] S02, camera installation and position calibration

[0136] In an embodiment of the present invention, the camera is installed on the gantry, and the height is the material pile protection height (such as 5 meters). An adjustable camera bracket needs to be used. The bracket usually has screws that can be finely adjusted, allowing adjustment of the horizontal and pitch angles of the camera. A horizontal line or bubble level mark is set at the installation position of the camera to help quickly check and adjust the level of the camera during installation.

[0137] A laser level is used for installation position calibration to ensure that the optical axis passing through the optical center of the camera is parallel to the horizontal plane; the camera is used to collect images, and image processing software is used to analyze the known horizontal and vertical lines in the images to determine the tilt angle of the camera and make adjustments to ensure that the horizontal line passing through the center point of the camera's field of view is parallel to the horizontal plane, and the vertical line passing through the center point of the camera's field of view is perpendicular to the horizontal plane.

[0138] S03, image acquisition and annotation

[0139] In the embodiments of the present invention, it is necessary to collect images of the stockpile at different heights during the stockpiling process, including images of the stockpile below the protection height and images of the stockpile above the protection height; it is also necessary to collect images of the stockpile at different positions and at different distances from the camera; images of the stockpile under different weather and lighting conditions should be collected to ensure the diversity of data. According to the above requirements, multiple videos are recorded using the installed camera, and an image dataset is made from the collected videos: an image is selected from the video at regular intervals, and the cv2.absdiff() function of the OpenCV library is used to compare the differences between the selected images. If the difference is less than the set threshold (i.e., the similarity between the two images is too high), the image is not retained.

[0140] Specifically, the dataset is labeled using a labeling tool, and the stockpile area in the image is framed using a target detection box.

[0141] S04, Selection and Training of Target Detection Model

[0142] S41, In the field of target detection, the YOLO series of models have always attracted much attention due to their excellent performance and flexibility. The most popular versions are YOLOv5 and YOLOv8. In terms of detection accuracy, YOLOv8 has higher accuracy than YOLOv5 and performs well in detecting small objects; in terms of detection speed, both YOLOv5 and YOLOv8 are suitable for real-time applications. The FPS of YOLOv5 on the CPU is higher than that of YOLOv8, but the FPS of YOLOv8 on the GPU is higher than that of YOLOv5. Therefore, YOLOv8 is selected as the model for stockpile detection.

[0143] S42. Specifically, the network structure of YOLOv8 mainly consists of three major parts, namely the Backbone main network, the Neck neck network, and the Head head network. The Backbone is mainly responsible for extracting features from the input image. This part uses the Cf2 module as the basic unit. The Cf2 module is a structure with few parameters and excellent feature extraction capabilities. The Backbone also uses depthwise separable convolution and dilated convolution to further enhance the feature extraction ability. The Neck is responsible for multi-scale feature fusion, fusing feature maps from different stages of the Backbone to enhance the feature representation ability. The Neck part contains three modules: The SPPF module uses pooling operations of different scales to splice feature maps of different scales together; the PAA module is responsible for intelligently allocating anchor boxes to optimize the selection of positive and negative samples; the PAN-FPN module aggregates features at different levels through two paths: bottom-up and top-down. The Head part contains a detection head and a classification head, which are responsible for the final object detection and classification tasks. The detection head is responsible for predicting the bounding box regression values and the confidence of the existence of the target for each anchor box; the classification head uses global average pooling to classify each feature map and outputs the probability distribution of each class.

[0144] In the embodiments of the present invention, the Conv Block is the most basic module in YOLOv8, including a Conv2d layer, a BatchNorm2d layer, and a SiLU activation function. The Conv2d layer performs a convolution operation on the input data to generate a feature map; the BatchNorm2d layer normalizes the features in each small batch of data, making the mean of each feature close to 0 and the variance close to 1 in the small batch, ensuring that the data passing through the network is neither too large nor too small, and guaranteeing the stable training of the model. The definition of the SiLU (Sigmoid Linear Unit) activation function is as follows:

[0145] SiLU(x) = x · σ(x)

[0146] where σ(x) is the Sigmoid function, defined as:

[0147] σ(x) = 1 / (1 + e -x )

[0148] The SiLU function has a smooth curve when approaching 0, which can effectively avoid the problem of gradient vanishing.

[0149] Specifically, the Bottleneck Block aims to reduce the computational complexity and the number of parameters while preserving the performance of the model. The specific implementation of the Bottleneck Block is achieved by adding a skip connection between convolutional layers, directly connecting the input to the output, and concatenating the output after the skip connection with the output of the convolutional layer to form the final output. This module can increase the network depth while alleviating the problem of vanishing gradients.

[0150] Specifically, the Cf2 module first processes the input feature map using the Conv Block, and then splits the generated intermediate feature map into two parts. One part is passed to the final Concat module, and the other part is passed to multiple Bottleneck Blocks for processing. The Concat module fuses the directly passed feature map and the processed feature map, enabling the model to comprehensively utilize multi-scale and multi-level information. In YOLOv8, the number of Bottleneck Blocks is defined by the depth_multiple parameter of the model, which means that the depth and computational complexity of the model can be flexibly adjusted according to requirements.

[0151] Specifically, the SPPF (Spatial Pyramid Pooling - Fast) module is used to capture multi-scale information, which is mainly implemented through three max pooling layers. The feature map is passed through three MaxPool2d layers, and each MaxPool2d layer performs pooling operations on the feature map using specific convolutional kernel sizes and strides. These different pooling operations can capture information at different scales. Finally, the output feature maps of the three MaxPool2d layers are concatenated in the channel dimension to fuse the multi-scale features into one feature map, enriching the feature representation.

[0152] In the embodiments of the present invention, the loss function of the model includes the bounding box regression loss and the classification loss. The bounding box regression loss is used to calculate the difference between the predicted bounding box and the ground truth bounding box. Mean Squared Error (MSE) is a commonly used loss function, which assigns a higher penalty for larger errors, helping the model quickly correct large prediction errors. Therefore, the bounding box regression loss calculates the sum of the squares of the differences between the predicted coordinates and the actual coordinates, and the formula is as follows:

[0153]

[0154] where x i represents the coordinates of the ground truth bounding box, represents the coordinates of the predicted bounding box. Using this loss function as the optimization objective, the model reduces the gap between the predicted box and the ground truth box during training.

[0155] In the embodiments of the present invention, the classification loss is used to measure the difference between the class distribution predicted by the model and the true label. The cross-entropy loss is used as the loss function, which gives a higher penalty for incorrect predictions. The classification loss can help the model optimize its predictions in classification problems, making the predicted probability distribution as close as possible to the true label distribution. The formula is:

[0156]

[0157] where y o,c is an indicator, which is 1 if the sample o belongs to class c, and 0 otherwise. p o is the probability that the model predicts that the sample o belongs to class c.

[0158] In the embodiments of the present invention, the stockpile detection is trained using the YOLOv8 model. The graphics card used is the RTX3090 with 24G video memory. The batch is set to 64, the input image size imgsz is set to 640, the learning rate is set to 0.01, and the final learning rate is 0.0001. The Adam optimizer is used to adjust the learning rate, and the official pre-trained weights are loaded for training.

[0159] S05, Comparison of the stockpile bounding box and the height of the protection line

[0160] In the embodiments of the present invention, the video stream capture function is implemented to obtain video data from the camera in real time. A horizontal line is drawn in the center of the video as the protection line. Assuming the video height is H pixels, the position of the protection line Apply the object detection model to each frame of the image to detect the position of the stockpile. The model outputs the predicted bounding box of the stockpile. The height of the upper border of the predicted bounding box is:

[0161]

[0162] where y c is the y value of the center point coordinate of the predicted bounding box, and h is the height of the predicted bounding box.

[0163] S06, Alarm function

[0164] In the embodiments of the present invention, when the predicted bounding box is higher than the protection line, that is, y max > y, an alarm is automatically issued. When there is an alarm in the portal machine, the stockpiling operation is stopped.

[0165] Example 3

[0166] The material pile height detection device 100 described in the present invention can be installed in an electronic device. According to the functions achieved, the material pile height detection device 100 can include an image processing module 101, a model training module 102, a height detection module 103, and a height judgment module 104. The modules described in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0167] In this embodiment, the functions of each module / unit are as follows:

[0168] The image processing module 101 is used to obtain multi-scale images of the material pile, obtain a multi-scale image set, perform similarity deletion processing on the multi-scale image set, and obtain a preprocessed image set;

[0169] The model training module 102 is used to perform material pile area annotation processing on the preprocessed image set to obtain training data, and use the training data to train a material pile target detection model;

[0170] The height detection module 103 is used to obtain material pile video data in real time, use the material pile target detection model to perform real-time detection on the material pile video data, obtain material pile bounding box information, and obtain the upper border height according to the material pile bounding box information;

[0171] The height judgment module 104 is used to judge whether the upper border height is greater than a preset safety height. If the upper border height is less than or equal to the safety height, the step of obtaining the material pile video data in real time and using the material pile target detection model to perform real-time detection on the material pile video data to obtain the material pile bounding box information is returned. If the upper border height is greater than the safety height, a stop material pile operation instruction is issued.

[0172] Example 4

[0173] Based on the above embodiments, this embodiment provides a computer device, including a memory, a processor, and a computer program product stored on the memory. The processor executes the computer program product to implement the steps of the method described in the above embodiments.

[0174] In some embodiments of this embodiment, a computer-readable storage medium is provided, on which a computer program product is stored. The computer program product is characterized in that when executed by a processor, it implements the steps of the method described in the above embodiments.

[0175] In some embodiments of the present embodiment, a computer program product is provided, including a computer program product, characterized in that when the computer program product is executed by a processor, the steps of the method described in the above embodiments are implemented.

[0176] The processor may include, but is not limited to, for example, one or more processors or microprocessors, etc. Each processor may be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute the method in the above embodiments.

[0177] The computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof. The computer-readable storage medium may include, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy disks, solid state drives, removable disks, Blu-ray discs, etc.).

[0178] The computer-readable storage medium may also store at least one computer-executable program, and the computer-executable program is, for example, a computer-readable instruction. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may include, for example, read-only memory (ROM), hard disks, flash memory, etc. For example, the non-transitory computer-readable storage medium may be connected to a computing device such as a computer. Then, when the computing device runs the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.

[0179] In addition, the computer device may further include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (such as a keyboard, a mouse, a speaker, etc.).

[0180] The processor may communicate with external devices via the I / O bus through a wired or wireless network.

[0181] In one embodiment, the at least one computer-executable instruction may also be compiled into or comprise a software product / computer program product, wherein when one or more computer-executable instructions are run by a processor, the various functions and / or method steps in the embodiments described in this technology are executed.

[0182] In the embodiments provided in this disclosure, it should be understood that the disclosed devices and methods may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0183] It should be noted that in this disclosure, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, the element limited by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0184] Although the disclosed embodiments are as above, the above content is only an embodiment adopted for the convenience of understanding this disclosure and is not used to limit this disclosure. Any person skilled in the art within the technical field to which this disclosure pertains may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed by this disclosure. However, the scope of patent protection of this disclosure shall still be subject to the scope defined by the appended claims.

Claims

1. A method for detecting the height of a pile of materials, characterized in that: include: Acquire a multi-scale image of the stockpile to obtain a multi-scale image set, and perform similarity removal processing on the multi-scale image set to obtain a pre-processed image set; Performing stockpile region labeling processing on the preprocessed image set to obtain training data, and using the training data to train a stockpile target detection model; Acquire the material pile video data in real time, and use the material pile target detection model to perform real-time detection on the material pile video data to obtain the material pile boundary box information; Obtaining the height of the upper frame according to the material pile boundary frame information; Determine whether the height of the upper frame is greater than a preset safety height; If the height of the upper frame is less than or equal to the safety height, returning to the step of acquiring the material pile video data in real time, using the material pile target detection model to perform real-time detection on the material pile video data to obtain the material pile boundary box information; If the height of the upper frame is greater than the safety height, an instruction to stop the stockpile operation is issued.

2. The method according to claim 1, characterized in that The performing similarity removal processing on the multi-scale image set to obtain a pre-processed image set includes: Selecting images from the multi-scale image set according to a preset interval to obtain a selected image set; Calculating the similarity between two adjacent images in the selected image set; If the similarity between two adjacent images is greater than or equal to a preset similarity threshold, the image information of the first image of the two adjacent images is obtained to obtain a deleted image information set; The corresponding images in the multi-scale image set are deleted according to the deleted image information set to obtain a preprocessed image set.

3. The method according to claim 1, characterized in that The method of using the training data to train a pile target detection model includes: Performing convolution processing on the preprocessed image set to obtain a convolution feature set; Performing standardization on the convolution feature set to obtain a standardized feature set; A nonlinear factor is introduced into the standardized feature set using a preset activation function to obtain an activation feature, wherein the activation function is expressed by the following formula: SiLU(x)=x·1 / (1+e -x ) Wherein, SiLU(x) is the activation function, x is the standardized feature set, and e is the base of the natural logarithm; Outputting a training detection result according to the activation feature using a pre-established YOLOv8 model; Calculating bounding box regression loss and classification loss of the training detection result and the standard annotation of the training data; After adjusting the weight parameters and the learning rate of the YOLOv8 model according to the bounding box regression loss and the classification loss, returning to the step of outputting the training detection result according to the activation feature using the pre-established YOLOv8 model; When the preset number of iterations is reached, the model training is confirmed to be completed and the stockpile target detection model is obtained.

4. The method according to claim 3, characterized in that The calculating the bounding box regression loss and the classification loss of the training detection result and the standard annotation of the training data includes: The bounding box regression loss is calculated using the following formula: Among them, L bbox is the bounding box regression loss, N is the number of samples, x i is the standard label corresponding to the i-th data in the training data, is the i-th test result in the training test results; The classification loss is calculated using the following formula: Among them, L cls is the classification loss, M is the number of classification types, c is the type number, y o,c is a binary function. If the sample o belongs to type c, it outputs 1, otherwise it outputs 0. i Represents the probability that sample o belongs to type c.

5. The method according to claim 4, characterized in that The method of using the pile target detection model to perform real-time detection on the pile video data to obtain the pile boundary box information includes: Using the material pile target detection model, convolution processing and standardization processing are performed on each frame of the material pile video data to obtain material pile image features; Splitting the material pile image feature to obtain a first material pile image feature and a second material pile image feature; Performing convolution stitching processing on the second material pile image feature to obtain a convolution stitching feature, and fusing the first material pile image feature with the convolution stitching feature to obtain a fusion feature; Extracting pooling features of different scales of the fusion features to obtain pooling features of different scales; The pooled features of different scales are spliced ​​in the channel dimension to obtain multi-scale features; The stockpile target detection model is used to output stockpile bounding box data according to the multi-scale features.

6. The method according to claim 1, characterized in that The obtaining the height of the upper frame according to the material pile boundary frame information includes: The height of the upper border is calculated using the following formula: Among them, y max is the height of the upper border, y c is the center point pixel height contained in the material pile boundary box information, and h is the frame height contained in the material pile boundary box information.

7. A material pile height detection device, characterized in that: include: An image processing module is used to obtain multi-scale images of the stockpile to obtain a multi-scale image set, and perform similarity removal processing on the multi-scale image set to obtain a pre-processed image set; A model training module is used to perform stockpile area labeling processing on the pre-processed image set to obtain training data, and use the training data to train a stockpile target detection model; A height detection module is used to obtain the material pile video data in real time, use the material pile target detection model to perform real-time detection on the material pile video data to obtain the material pile boundary box information, and obtain the upper frame height according to the material pile boundary box information; The height judgment module is used to judge whether the height of the upper frame is greater than the preset safety height. If the height of the upper frame is less than or equal to the safety height, the module returns to the step of acquiring the material pile video data in real time, and uses the material pile target detection model to perform real-time detection on the material pile video data to obtain the material pile boundary box information. If the height of the upper frame is greater than the safety height, a command to stop the material pile operation is issued.

8. A computer device comprising a memory, a processor and a computer program product stored in the memory, characterized in that: The processor executes the computer program product to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program product stored thereon, characterized in that: When the computer program product is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program product, characterized in that When the computer program product is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.