Deep learning-based ore material accurate classification method and equipment
By improving the YOLO-V8 seg model, the DepthSepConv module and the coordinate attention mechanism module were introduced, combined with the CARAFE module and the bidirectional self-attention converter, the problems of inaccurate specification judgment and low model calculation efficiency in material classification are solved, and the accurate classification and specification prediction of materials are realized, which is suitable for scenarios with resource limitations.
Patent Information
- Application Number
- CN202510048514.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to accurately judge the specifications of materials in material classification, resulting in insufficient accuracy in material classification. The existing YOLO-V8 seg model is inefficient in resource-constrained scenarios, and it is difficult to apply.
Improve the YOLO-V8 seg model, introduce the DepthSepConv module and the coordinate attention mechanism module, build a material instance segmentation model, and realize accurate classification and specification prediction of materials through multi-scale feature extraction and feature fusion, and introduce CARAFE module and bidirectional self-attention converter into the NECK network to improve computing efficiency and generalization capabilities.
It realizes fast and accurate classification and specification prediction of materials, reduces the number of parameters and calculations of the model, makes it suitable for mobile and embedded devices, and improves the computing efficiency and generalization capabilities of the model.
Smart Images

Figure CN119992173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computing model technology, and in particular to a method and device for accurately classifying ore materials based on deep learning. Background Art
[0002] During the transportation process, materials need to be classified into different categories, and the transportation costs are calculated based on the weight of the goods. Therefore, the correctness of the judgment of the category of the transported materials must be ensured. The material classification here is not only to determine the major categories of materials, such as weathered rock, rough stone, gravel, slag, sand, etc., but also to determine the specifications based on the particle size of the materials on the basis of the major categories, so as to finally determine the classification of materials. For example, gravel is formed by crushing pebbles, granite, basalt, rocks, etc. The transportation prices of gravels of different specifications vary greatly. Gravels below 4.5mm are classified as machine-made sand, and those above 4.5mm need to be subdivided into different specifications, such as gravel 1 (3*3-5*5cm), gravel 2 (1*1-3*3cm), soil and stone (5*5-30*30cm), rock (30*30-60*60cm), etc. Therefore, the final material classification corresponding to different specifications is different.
[0003] At present, the judgment of material category is mainly carried out through manual review. During the process of intelligent weighing on the vehicle scale, on-site personnel manually judge the material category, which consumes a lot of manpower and time costs.
[0004] The existing technology mainly uses semantic segmentation algorithms to classify material images at the pixel level and finally identify the category of the material. For example, a method, device, electronic device and storage medium for classification using a segmentation model are provided in the publication number CN111738310A. The material image is input into a single material model to output the semantic segmentation result and material attribution vector of each pixel of the material image. However, traditional material segmentation can only classify materials and obtain material boundaries, and cannot directly judge the specifications of the material based on the identified results. Therefore, it is impossible to classify materials more accurately.
[0005] Although the existing YOLO-V8 seg model can also achieve classification and target segmentation, it consumes a lot of memory and occupies a lot of resources, making it difficult to apply in resource-constrained scenarios such as embedded devices or mobile terminals. At the same time, the existing YOLO-V8seg model also has problems such as low computational efficiency and recognition and segmentation accuracy that cannot meet the requirements. Therefore, it needs to be improved to reduce resource overhead and improve recognition accuracy and computational efficiency. Summary of the invention
[0006] In order to solve the above problems, the purpose of the present invention is to provide a method for accurate classification of ore materials based on deep learning, which can realize material segmentation and identification while completing the prediction of the specification size of each material, ensuring fast and accurate distinction between material categories.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] Technical Solution 1
[0009] A method for accurate classification of ore materials based on deep learning, comprising the following steps:
[0010] Collect the top view of the vehicle that needs material classification and extract the image of the truck bed area from it;
[0011] Based on the YOLO-V8 seg training network, an improvement is made to build a material instance segmentation model: the material instance segmentation model includes a backbone feature extraction network, a NECK network, a HEAD network and Protonet The backbone feature extraction network introduces a DepthSepConv module and a coordinate attention mechanism module, performs downsampling in a separate convolution manner through multiple DepthSepConv modules, obtains a multi-scale feature extraction map, and inputs the smaller-scale feature map into the coordinate attention mechanism module to enhance the ability to focus on key features;
[0012] The truck bed area image is input into the backbone feature extraction network to extract feature maps of multiple scales. The NECK network performs feature fusion on the feature maps of different scales and transmits them to the HEAD network. At the same time, the feature map with the largest scale after feature fusion is transmitted to the HEAD network. Protonet network, the HEAD network outputs the category of the target material, the Protonet The network outputs a predicted native mask of the target material; the HEAD network also combines the predicted native mask to determine the boundary of the target material, and then performs instance segmentation to obtain a target material mask map.
[0013] Preferably, the backbone feature extraction network performs the following steps: the image of the truck bed area is convolved by several CBS modules and then successively passed through multiple DepthSepConv modules for feature extraction to obtain multiple feature maps with gradually smaller scales. For feature maps with larger scales, they are also output to the NECK network through the C2F module; for feature maps with smaller scales, they are also output to the NECK network by the C2F module after passing through the coordinate attention mechanism module; for feature maps with the smallest scale, they are further pooled through the SPPF module after passing through the coordinate attention mechanism module and the C2F module and then output to the NECK network.
[0014] Preferably, the NECK network introduces a CARAFE module and a bidirectional self-attention transformer, and the NECK network performs the following steps: the feature maps of different scales are gradually fused from the highest level to the lowest level in the direction of small to large scale, and each fusion is up-sampled by the CARAFE module before the fusion, and the lowest level fused feature map is transmitted to the Protonet The lowest level fusion feature map is fused with the up-sampling result before the previous level fusion and then input into the bidirectional self-attention transformer. After the global information is captured by the bidirectional self-attention transformer, it is output to the HEAD network and fused with the down-sampling result of the minimum scale feature map at the same time, and then output to the HEAD network.
[0015] Preferably, the HEAD network first processes each received fusion feature map through three branches to generate a prediction box feature map, a prediction category feature map and a Mask coefficient feature map, two of which are composed of two CBS convolution modules and one Conv2d convolution layer, and the other branch is a mask_coefficients branch; then the lowest scale prediction box feature map and category feature map are upsampled using the CARAFE module, the largest scale prediction box feature map and category feature map are convolutionally downsampled, and then concatenated into feature maps through Concat. Feature map, the spliced prediction box feature map and category feature map are processed by non-maximum suppression, the duplicate detection boxes are removed, and the prediction boxes and prediction categories with the highest confidence are retained. At the same time, the Mask coefficient feature maps of different scales are fused, and the mask coefficient with the highest confidence is selected. The native mask output by the Protonet network and the mask coefficient with the highest confidence are linearly combined to generate the mask of the current target. Then, the prediction mask is cropped according to the prediction box to remove the background part irrelevant to the target area, and only the mask information in the box is retained to obtain the mask map of the target material.
[0016] Preferably, the top view of the vehicle is input into a trained key point detection model to obtain the vertex pixel coordinates of the four corners of the truck bed, and the length and width of the truck bed in the pixels are calculated; the actual length and width of the truck bed surface in the record are extracted, and combined with the length and width of the truck bed in the image, the length and width scale factors are calculated respectively; the length and width of the target material mask image are combined with the scale factor to deduce the actual length and width of the material, thereby determining the specifications of the material; according to the material category output by the instance segmentation model, combined with the specifications of the material, an accurate material classification is finally obtained.
[0017] Based on the same inventive concept, the present invention also provides a precise classification device for ore materials based on deep learning.
[0018] Technical Solution 2
[0019] A device for accurately classifying ore materials based on deep learning, comprising a processor and a memory storing an executable program, wherein the processor runs the executable program to execute the steps of technical solution one.
[0020] The present invention has the following beneficial effects:
[0021] 1. The present invention provides a method for accurate classification of ore materials based on deep learning. The target detection model of YOLO-V8 is used to quickly locate the position of the truck bucket area, obtain the image of the truck bucket area, and then combine the improved YOLO-V8 instance segmentation model. By introducing the attention mechanism and high-dimensional features, the complex and diverse material contours can be better extracted and learned, and the targets can be more finely identified and distinguished, so that the mask segmented by the model can more accurately express our material contours. At the same time, the improvement of the existing YOLO-V8 seg training network can reduce the number of parameters and the amount of calculation of the model, making it lightweight, and improve the computational efficiency and generalization ability of the model, making the model more suitable for resource-constrained scenarios such as mobile terminals and embedded devices.
[0022] 2. The present invention also adds key point detection of the four corners of the truck bed, locates the positions of the four corners of the truck bed in pixels, and thus infers the actual size of the material based on the actual length and width of the truck bed and the precise target material mask of the instance segmentation model. The accurate classification of the material is determined by the size, which can better adapt to more material scene requirements.
[0023] 3. The present invention provides a method for accurate classification of ore materials based on deep learning, which uses a variety of deep neural networks to infer and identify the types and size specifications of materials, directly avoiding manual intervention and improving the accuracy and reliability of vehicle weighing data. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of the present invention;
[0025] Figure 2 The improved instance segmentation model of the present invention;
[0026] Figure 3 It is a network structure diagram of the DepthSepConv module of the present invention;
[0027] Figure 4 A network structure diagram of the coordinate attention mechanism module of the present invention;
[0028] Figure 5 It is a network structure diagram of the C2F module of the present invention;
[0029] Figure 6 It is a network structure diagram of the SPPF module of the present invention;
[0030] Figure 7 A network structure diagram of a bidirectional self-attention transformer of the present invention;
[0031] Figure 8 This is the Protonet network structure diagram of the present invention. DETAILED DESCRIPTION
[0032] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0033] See also Figure 1 , a method for accurate classification of ore materials based on deep learning, comprising the following steps:
[0034] Step 1: Collect the top view of the vehicle that needs material classification, and extract the image of the truck bed area and the pixel coordinates of the four corners of the truck bed. At the same time, obtain the length and width of the truck bed from the vehicle registration information.
[0035] The process of collecting the vehicle top view is as follows: when a vehicle carrying materials passes through the weighbridge, the angle of the industrial camera installed above the weighbridge is adjusted so that the camera can collect a complete top view of the vehicle and keep it horizontally aligned with the weighbridge.
[0036] The extraction of the truck bucket area image can be obtained by training the YOLO-V8 target detection model to obtain a truck bucket detection weight model. The truck bucket detection weight model identifies the truck bucket position and outputs the position of the truck bucket prediction box, that is, the upper left corner pixel coordinates and the lower right corner pixel coordinates of the truck bucket prediction box. The pixel coordinates of the four corner vertices of the truck bucket (not the truck bucket prediction box coordinates, but the actual vertex position coordinates of the truck bucket) can be obtained by training the YOLO-V8 Pose key point detection model to obtain a truck bucket key point detection model. The truck bucket key point detection model outputs the vertex pixel coordinates of the four corners of the truck bucket. The specific process is as follows:
[0037] S11, collecting the top view of the vehicle, performing image preprocessing such as denoising, and then using it as a sample data set for vehicle bucket detection, and dividing it into a training set, a test set, and a validation set;
[0038] S12, the truck bucket detection sample data set is used to train the YOLO-V8 target detection model and the YOLO-V8 Pose key point detection model, and the specific steps are as follows: 1) Use labelme to mark the truck bucket position of the vehicle image in the data set. Add the key point ID annotation while marking the rectangular box. 2) After the annotation is completed, set a certain number of training times and parameters, and train the YOLO-V8 target detection model and the YOLO-V8 Pose key point model for the labeled images to extract the feature information of the truck bucket, and continuously adjust the optimization parameters to finally obtain the trained YOLO-V8 truck bucket detection weight model and the YOLO-V8 truck bucket key point detection model. 3) Import the images in the verification set into the YOLO-V8 truck bucket detection weight model and the YOLO-V8 key point detection model for testing. The test output result is the pixel coordinates of the positions of the truck bucket rectangular prediction box pixels (left1, top1), (right1, bottom1) and the four corner vertices of the truck bucket in the top view of the vehicle taken by the industrial camera.
[0039] S13, input the top view of the vehicle into the trained YOLO-V8 bucket detection weight model, obtain the positions of the pixel points of the bucket rectangular prediction box (left1, top1), (right1, bottom1), crop the bucket area image, and use it for the instance segmentation model in step 2. Input the top view of the vehicle into the bucket key point detection model to obtain the pixel point coordinates of the four corner vertices of the bucket, which are used in step 3 to calculate the scale factor.
[0040] Step 2: Improve the YOLO-V8 seg training network and build a material instance segmentation model.
[0041] The truck bed area image is input into a trained material instance segmentation model, the category of the target material and the native mask of the material are output, and then instance segmentation is performed to obtain a mask map of the target material.
[0042] See also Figure 2 , the construction process of the material instance segmentation model is:
[0043] The material instance segmentation model includes a backbone feature extraction network, a NECK network, a HEAD network and a Protonet network.
[0044] See also Figures 3 to 6, the backbone feature extraction network is composed of a DepthSepConv module (DepthSepConv module, depth separable convolution module), a CBS module, a C2f module (Cross-StageFeature, cross-stage feature fusion module), an SPPF module and a coordinate attention mechanism module (i.e., a CA module, Coordinate Attention coordinate attention mechanism), wherein the CBS module represents a convolution layer, a batch normalization BN and an activation function SILU, which are used to perform convolution feature extraction on the input image or feature map. By adjusting the step size of the convolution kernel, the size of the feature map can be controlled to achieve downsampling or keep the size unchanged. In this embodiment, 4 DepthSepConv modules and 2 CBS modules are used, and downsampling is performed in a separate convolution manner through multiple DepthSepConv modules to obtain a multi-scale feature extraction map, and the feature map with a smaller scale is input into the coordinate attention mechanism module to enhance the ability to focus on key features. For example, the input 640*640 image can be changed to 20*20, and the number of input channels can be changed from 3 to 512. Specifically, the image of the truck bed area is convolved by two CBS modules and then sequentially passed through four DepthSepConv modules for feature extraction to obtain multiple feature maps with gradually smaller scales. Figure 2 , for feature maps with larger scales, such as feature Figure 1 (1,128,160,160), Features Figure 2 (1,256,80,80), and is also output to the NECK network through the C2F module. For feature maps with smaller scales, such as feature Figure 3 (1,512,40,40), and then output to the NECK network by the C2F module after passing through the coordinate attention mechanism module. For the smallest feature map, such as feature Figure 4
[0045] (1,512,20,20), after passing through the coordinate attention mechanism module and the C2F module, it is further processed by the SPPF module and then output to the NECK network. In this embodiment, the DepthSepConv module (depth separable convolution module) is introduced, which is a lightweight convolution module. It uses a separated convolution method to replace the traditional convolution module to lightweight the entire network, improve the computational efficiency and generalization ability of the model, reduce the number of parameters and the amount of calculation of the model, and make the model more suitable for resource-constrained scenarios such as mobile terminals and embedded devices. At the same time, adding a coordinate attention mechanism module to the backbone feature extraction network is an attention mechanism that combines channel attention and position information. By weighting the different importance of channel features, it enhances the ability to focus on key features to highlight the representation of objects of interest. After introducing the coordinate attention mechanism module, the model can more accurately locate the edge position of the material and identify the specific features of the material that need to be paid attention to, without increasing the computational overhead, and significantly improving the performance of the model. The SPPF module (spatial pyramid pooling layer) can generate its own features. Figure 4 (1,512,20,20) After passing through multiple small-sized maximum pooling kernels, multiple features are fused and passed to the next part of the Neck network. In the SPPF module, the performance of the SPPF module is optimized by combining the global average pooling layer and the global maximum pooling layer to improve the accuracy and robustness of the model. In the present invention, only one SPPF module can be used, otherwise it will bring some negative effects: first, it will increase the amount of calculation and memory consumption, because each additional layer will require more calculations and storage. Secondly, too many SPPF layers may cause the model to overfit the training data, causing it to perform poorly on new data. Furthermore, if too many features are fused, redundant information may be generated, but the performance will not be improved much. Finally, the model will become more complicated and more difficult to understand and explain. Therefore, in combination with the application scenario of the present invention, only one SPPF module is selected to process the feature map with the smallest scale (i.e., the lowest resolution).
[0046] The NECK network performs feature fusion on feature maps of different scales and transmits them to the HEAD network. At the same time, the feature map with the largest scale (i.e., the highest resolution) after feature fusion is transmitted to the Protonet network, and the HEAD network outputs the category of the target material. In the present invention, the NECK network introduces a CARAFE module and a bidirectional self-attention transformer (i.e., a BiFormer module). The NECK network performs the following steps: the feature maps of different scales are fused step by step from the highest level to the lowest level in the direction of scale from small to large, and each time before fusion, the CARAFE module is used for upsampling, and the lowest level of fused feature maps are transmitted to the Protonet network and the HEAD network. At the same time, the lowest level of fused feature maps are also fused with the upsampling results before the previous level fusion and then input into the bidirectional self-attention transformer. After capturing global information through the bidirectional self-attention transformer, it is output to the HEAD network and simultaneously fused with the downsampling results of the minimum scale feature map, and then output to the HEAD network. In the present invention, an upsampling operator CARAFE module is introduced, which adopts a lightweight and efficient upsampling method, which upsamples the feature map through content-aware feature reorganization to extract richer semantic feature information. At each position, the CARAFE module uses the underlying content information to predict the reorganization kernel and performs feature reorganization in a predefined neighborhood area. In addition, such a design enables the CARAFE module to adaptively use the optimized reorganization kernel at different positions. Existing public upsampling methods include nearest neighbor interpolation and bilinear interpolation, which perform upsampling by the spatial position of pixels, but fail to utilize the semantic information of feature maps. Due to the small receptive field, these methods are usually unable to capture the rich semantic information required for the task of predicting the aggregation of large materials. The upsampling module of the existing YOLO-V8 seg model adopts a deconvolution upsampling method, which performs a convolution operation on the input feature map and uses zero padding and step size adjustment to expand the size of the output feature map. However, since the same convolution kernel is applied to the entire image, this method limits the responsiveness to local changes, and also affects the computational efficiency due to the large number of parameters. The present invention adopts content-aware function reorganization to replace the original Upsample upsampling module, which improves the regression accuracy and convergence speed of the network and can also extract richer feature information. Figure 7The bidirectional self-attention transformer uses a self-attention mechanism to capture global information, allowing the information at each position of the material image to directly interact with the information at all other positions. At the same time, the weights between elements are dynamically calculated, so that the model can focus on the features of the input that are most relevant to the material. The bidirectional self-attention mechanism processes the flow of information simultaneously in the encoding and decoding stages, thereby improving the ability to understand material features in context and enhancing the model's attention and weight allocation to input material features. By automatically learning the relevance and importance in the data, the model can process different inputs more specifically.
[0047] In the present invention, the NECK network also adds the fusion of the lowest layer of large-scale feature maps. Since the large-scale feature maps are located in the earlier stage of the network and have higher resolution, they can capture the details in the image more finely, thereby improving the model's ability to distinguish small-sized target materials from the background, and improving the model's segmentation performance for small-sized target materials.
[0048] The four features output by the backbone feature extraction network are Figure 1 (1,128,160,160), Features Figure 2 (1,256,80,80), features Figure 3 (1,512,40,40) and features Figure 4 (1,512,20,20) as an example, after NECK network fusion, the fused features are output Figure 5 (1,512,20,20), Features Figure 6 (1,512,40,40) and features Figure 7 (1,256,80,80). Figure 5 ,feature Figure 6 And features Figure 7 are all input into the HEAD network, and at the same time, the largest feature Figure 7 (1,256,80,80) is input to the Protonet network, and the Protonet network outputs the network predicted native mask feature map (1,32,160,160).
[0049] The HEAD network first processes each received fusion feature map through three branches to generate a prediction box feature map, a prediction category feature map, and a Mask coefficient feature map. Two of the branches consist of two CBS convolution modules and one Conv2d convolution layer, and the other branch is
[0050] mask_coefficients branch (i.e., mask coefficient branch); then use the CARAFE module to upsample the lowest-scale prediction box feature map and category feature map, convolution downsample the largest-scale prediction box feature map and category feature map, and then concatenate them into feature maps through Concat. After concatenating the prediction box feature map and category feature map, they are processed by non-maximum suppression to remove duplicate detection frames and retain the prediction frames and prediction categories with the highest confidence. At the same time, the Mask coefficient feature maps of different scales are fused, the mask coefficient with the highest confidence is selected, and the native mask output by the Protonet network and the mask coefficient with the highest confidence are linearly combined to generate the mask of the current target, that is, to determine the boundary of the target material, and then crop the prediction mask according to the prediction box to remove the background part that is not related to the target area, and only retain the mask information in the box. Preferably, the cropped prediction mask is subjected to threshold processing and binarized to simplify the pixel values in the mask into two levels: usually black or white (0 or 1), where white (or 1) represents the target area and black (or 0) represents the background. The threshold is selected as 0.5, and a binary prediction mask is finally obtained. The binarized mask image performs edge detection and contour smoothing on the contour to optimize the appearance and accuracy of the contour. The predicted contour is a material contour line surrounded by multiple pixels. Opencv is used to draw a transparent category color on the original image to highlight the segmented area and category. After segmentation, the target material mask image is obtained.
[0051] The truck bed area image obtained in the training process in step 1 is used as the material segmentation dataset to train the improved YOLO-V8 seg training network and obtain the material instance segmentation model. The training process is as follows: All images in the material segmentation dataset are converted to 640*640 size to fix the image format for training, and the contour lines of the stones in the processed images are annotated with labelme. After the annotation is completed, each image will generate a label in Json format. By reading each Json file, all the coordinates and category data in the file are transferred to a new txt file according to the structure of (category, coordinate x1, coordinate y1, coordinate x2, coordinate y2...) to realize the conversion of Json to the txt label format required for YOLO training. Each image corresponds to a label in json format and txt format, and finally divided into training set, test set, and validation set in the manner of 8:1:1. Prepare the txt files required for model training, train.txt, test.txt and val.txt, which are the index files required for model training set, test set and validation set respectively; put the processed material segmentation data set into the improved YOLO-V8 seg training network for training to obtain the model weight; use the exponential decay ExponentialLR method to reduce the learning rate, use the RandomErasing and CutMix methods for data enhancement, cut out a part of the area, randomly fill in the remaining pixel values of other data in the data set, and occlude or overlap the image information in the annotated original image to simulate occlusion or overlap for training. The model parameters are lr=1e-4, Batch_size=8, and Epoch=300.
[0052] Import the validation set image samples of the material segmentation dataset into the YOLO-V8 seg training network for test analysis. In terms of the loss function, CIoU (Center Intersection over Union) loss is used as the positioning loss for the target detection part. CIOU (A,B)=α·d Chamfer (A, B) + β·IoU(A, B), where dchamfe(A, B) represents the difference in distance between the AB shapes, IoU(A, B) is the intersection-over-union ratio, and α and β are weighting coefficients. For the semantic segmentation part, Dice loss is introduced to enhance the ability to handle category imbalance, solve problems such as category imbalance and insufficient positioning accuracy, and improve training efficiency and the model's ability to handle complex scenes. Among them, p i represents the predicted probability of the i-th pixel, t i represents the true label (0 or 1) of the i-th pixel. ∈ is a small positive number used to prevent the denominator from being zero.
[0053] Step 3. Calculate the length and width of the bucket in the pixel points according to the vertex pixel coordinates of the four corners of the bucket output by the key point detection model; extract the actual length and width of the bucket surface in the record, and calculate the length and width proportional factors respectively by combining the length and width of the bucket in the image; combine the length and width of the target material mask image with the proportional factor to deduce the actual length and width of the material, thereby determining the specifications of the material; finally obtain accurate material classification according to the material category output by the instance segmentation model and the material specifications.
[0054] The specific process is as follows:
[0055] Step 31: The coordinates of the four corner pixel points of the truck bed output by the key point detection model are: A(x1,y1), B(x2,y2), C(x3,y3), and D(x4,y4). The straight-line distances between A and B, and between B and C are calculated as the length and width of the truck bed pixel points.
[0056] Step 32: Extract the actual length and width of the truck bed surface in the record, combine the length and width of the pixels displayed in the image, and use the following formula to calculate the length and width parameter ratio factor between the distance displayed in the image and the actual distance: U l =CL s / CL t , U w =CW s / CW t , where CL s Represents the actual length of the truck bed, CL t Represents the length of the truck body shown in the image, CW s Represents the actual width of the truck bed, CW t Represents the width of the truck bed displayed in the image, U l , U w They are the scale factors of the length and width of the image display distance and the real distance, respectively, indicating the conversion relationship from the image length and width to the actual length and width.
[0057] Step 32: According to formula L s =U l *L t , W s =U w *W t , infer the actual length and width of the target material, where L s Represents the actual length of the material, L t Represents the length of the material displayed in the image, W s Represents the actual material width, W t Represents the width of the material displayed in the image.
[0058] Step 33: Based on the inferred actual length and width information of the material, the material category identified by the instance segmentation model is judged to finally obtain the material category loaded by the vehicle. For materials that do not require specifications to accurately determine the material category, the material category identified by the instance segmentation model is directly used as the final recognition result, such as sand, coal, etc.
[0059] The present invention provides a method for accurately classifying ore materials based on deep learning. The target detection model of YOLO-V8 is used to quickly locate the position of the bucket area, obtain the image of the bucket area, and then combine with the improved YOLO-V8 instance segmentation model to more finely identify and distinguish targets, better solve the segmentation of overlapping or occluded targets, and obtain accurate material categories. The present invention also adds key point detection of the four corners of the bucket to locate the positions of the four corners of the bucket in pixels, thereby inferring the actual size of the material based on the actual length and width of the bucket and the accurate target material mask of the instance segmentation model, and accurately classifying the material by judging the size, which can better adapt to more material scene requirements.
[0060] The present invention discloses a method for accurately classifying ore materials based on deep learning, which improves the existing YOLO-V8 training network to make it lightweight, improves the computational efficiency and generalization ability of the model, reduces the number of parameters and the amount of computation of the model, and makes the model more suitable for resource-constrained scenarios such as mobile terminals and embedded devices.
[0061] Embodiment 2
[0062] A device for accurately classifying ore materials based on deep learning comprises a processor and a memory storing an executable program, wherein the processor runs the executable program to execute the method steps in the first embodiment.
[0063] Those skilled in the art can understand the specific implementation of the device, which will not be described in detail here. Any system that executes the method described in the first embodiment of the present invention belongs to the scope of protection of the present invention.
[0064] The above description is only a specific implementation mode of the present invention, and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for accurate classification of ore materials based on deep learning, characterized by: The steps include: Collect the top view of the vehicle that needs material classification and extract the image of the truck bed area from it; Based on the YOLO-V8 seg training network, an improvement is made to build a material instance segmentation model: the material instance segmentation model includes a backbone feature extraction network, a NECK network, a HEAD network and Protonet The backbone feature extraction network introduces a DepthSepConv module and a coordinate attention mechanism module, performs downsampling in a separate convolution manner through multiple DepthSepConv modules, obtains a multi-scale feature extraction map, and inputs the smaller-scale feature map into the coordinate attention mechanism module to enhance the ability to focus on key features; The truck bed area image is input into the backbone feature extraction network to extract feature maps of multiple scales. The NECK network performs feature fusion on the feature maps of different scales and transmits them to the HEAD network. At the same time, the feature map with the largest scale after feature fusion is transmitted to the HEAD network. Protonet network, the HEAD network outputs the category of the target material, the Protonet The network outputs a predicted native mask of the target material; the HEAD network also combines the predicted native mask to determine the boundary of the target material, and then performs instance segmentation to obtain a target material mask map.
2. The method for accurate classification of ore materials based on deep learning according to claim 1, characterized in that: The backbone feature extraction network performs the following steps: the image of the truck bed area is convolved by several CBS modules and then sequentially passed through multiple DepthSepConv modules for feature extraction to obtain multiple feature maps with gradually smaller scales. For feature maps with larger scales, they are also output to the NECK network through the C2F module. For feature maps with smaller scales, they are also output to the NECK network by the C2F module after passing through the coordinate attention mechanism module. For feature maps with the smallest scale, they are further pooled through the SPPF module after passing through the coordinate attention mechanism module and the C2F module and then output to the NECK network.
3. The method for accurate classification of ore materials based on deep learning according to claim 2 is characterized in that: The NECK network introduces the CARAFE module and the bidirectional self-attention transformer. The NECK network performs the following steps: the feature maps of different scales are fused step by step from the highest level to the lowest level in the direction of small to large scale, and each fusion is up-sampled by the CARAFE module before the fusion, and the lowest level fused feature map is transmitted to the Protonet The lowest level fusion feature map is fused with the up-sampling result before the previous level fusion and then input into the bidirectional self-attention transformer. After the global information is captured by the bidirectional self-attention transformer, it is output to the HEAD network and fused with the down-sampling result of the minimum scale feature map at the same time, and then output to the HEAD network.
4. The method for accurate classification of ore materials based on deep learning according to claim 3 is characterized in that: The HEAD network first processes each received fusion feature map through three branches to generate a prediction box feature map, a prediction category feature map and a Mask coefficient feature map, two of which are composed of two CBS convolution modules and one Conv2d convolution layer, and the other branch is a mask_coefficients branch; then the lowest scale prediction box feature map and category feature map are up-sampled using the CARAFE module, and the largest scale prediction box feature map and category feature map are convolutionally down-sampled, and then spliced into feature maps through Concat, and the spliced prediction box feature map and category feature map are subjected to non-maximum suppression processing to remove duplicate detection frames, and retain the prediction frame and prediction category with the highest confidence. At the same time, the Mask coefficient feature maps of different scales are fused, and the mask coefficient with the highest confidence is selected. The native mask output by the Protonet network and the mask coefficient with the highest confidence are linearly combined to generate a mask of the current target, and then the prediction mask is cropped according to the prediction box to remove the background part irrelevant to the target area, and only the mask information in the box is retained to obtain the target material mask map.
5. The method for accurate classification of ore materials based on deep learning according to claim 4 is characterized in that: Input the top view of the vehicle into the trained key point detection model, obtain the vertex pixel coordinates of the four corners of the truck bucket, and calculate the length and width of the truck bucket in the pixel points; extract the actual length and width of the truck bucket surface in the record, and calculate the length and width scale factors respectively in combination with the length and width of the truck bucket in the image; Combine the length and width of the target material mask image with the scale factor to calculate the actual length and width of the material, thereby determining the specifications of the material; Based on the material category output by the instance segmentation model and combined with the material specifications, accurate material classification is finally obtained.
6. A device for accurate classification of ore materials based on deep learning, characterized in that: The method comprises a processor and a memory storing an executable program, wherein the processor runs the executable program to execute any one of the steps in claims 1 to 5.
Citation Information
Patent Citations
Material classification method and device, electronic equipment and storage medium
CN111738310A
Cited By
Surface ore segmentation and particle size analysis method based on YOLO and prompt learning
CN121564077A