Ore crushing energy consumption formula, energy consumption prediction method, device, equipment and medium
By constructing the data set and using the DeepSet and KAN network models to learn the relationship between ore particle size and crushing energy consumption, combining the target identification network model to detect and track ore particle size, the problem of large-scale changes in crushing energy consumption caused by the uncertainty of ore particle size distribution is solved, and high-efficiency and energy-saving ore crushing is achieved.
Patent Information
- Application Number
- CN202510565234.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-19
AI Technical Summary
In the prior art, due to the uncertainty of explosive explosion energy, the ore particle size distribution is large, which in turn leads to a large change in the crushing energy consumption of the ore during crushing, making it difficult to control the energy consumption of the crusher.
By constructing the data set, the relationship between ore particle size distribution and crushing energy consumption is learned using the DeepSet network and KAN network model, the ore crushing energy consumption prediction formula is obtained, and the ore particle size distribution is detected and tracked in combination with the target recognition network model, and the blasting scheme is optimized to reduce crushing energy consumption.
It achieves blasting ore with the smallest possible crushing energy consumption, achieving the purpose of efficient and energy saving, accurately obtaining the relationship between ore crushing energy consumption and particle size distribution, and optimizing the blasting plan to reduce energy consumption.
Smart Images

Figure CN120509295A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and mineral processing technology, and in particular to an ore crushing energy consumption formula and an energy consumption prediction method, device, equipment and medium. Background Art
[0002] When crushing ore, the crusher's energy consumption is primarily related to the ore's particle size distribution. Controlling ore particle size can ultimately control energy consumption in the beneficiation process. However, in existing technologies, the uncertainty of explosive blast energy leads to a wide range of ore particle size distribution, which in turn results in significant variations in crushing energy consumption. To minimize crusher energy consumption, it is necessary to understand the relationship between ore crushing energy consumption and ore particle size distribution. Summary of the Invention
[0003] The main purpose of this application is to provide an ore crushing energy consumption formula and an energy consumption prediction method, device, equipment and medium, aiming to solve the technical problem of how to obtain the relationship between ore crushing energy consumption and ore particle size distribution.
[0004] In a first aspect, the present application provides a method for obtaining a prediction formula for ore crushing energy consumption, the method comprising:
[0005] Constructing a dataset based on the ore particle size distribution and corresponding ore crushing energy consumption of multiple groups of ore piles;
[0006] The data set is input into the DeepSet network to obtain a first eigenvector about the ore particle size distribution and a second eigenvector about the corresponding ore crushing energy consumption, and multiple groups of first eigenvectors and corresponding second eigenvectors are input into the KAN network to obtain an ore crushing energy consumption prediction formula; alternatively, the data set is directly input into the KAN network to obtain an ore crushing energy consumption prediction formula; the ore crushing energy consumption prediction formula is used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile.
[0007] The present application also provides a method for predicting ore crushing energy consumption, which includes:
[0008] Obtaining an ore crushing energy consumption prediction formula according to any of the above methods for obtaining an ore crushing energy consumption prediction formula;
[0009] The obtained ore particle size distribution of the target ore pile is substituted into the ore crushing energy consumption prediction formula to obtain the ore crushing predicted energy consumption of the target ore pile.
[0010] The present application also provides a device for obtaining a prediction formula for energy consumption of ore crushing, the device for obtaining a prediction formula for energy consumption of ore crushing comprising:
[0011] A data set construction module, for constructing a data set based on the ore particle size distribution and corresponding ore crushing energy consumption of multiple groups of ore piles;
[0012] A formula prediction module is used to input a data set into a DeepSet network to obtain a first eigenvector about the ore particle size distribution and a second eigenvector about the corresponding ore crushing energy consumption, and input multiple groups of first eigenvectors and corresponding second eigenvectors into a KAN network to obtain an ore crushing energy consumption prediction formula; or, directly input the data set into the KAN network to obtain an ore crushing energy consumption prediction formula; the ore crushing energy consumption prediction formula is used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile.
[0013] The fourth aspect of the present application provides a computer device, comprising: a memory and at least one processor, wherein instructions are stored in the memory; at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned method for obtaining the ore crushing energy consumption prediction formula, or so that the computer device executes the above-mentioned method for predicting ore crushing energy consumption.
[0014] The fifth aspect of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the above-mentioned method for obtaining the ore crushing energy consumption prediction formula, or, when the instructions are executed by the processor, implements the above-mentioned method for predicting ore crushing energy consumption.
[0015] This application obtains the ore particle size distribution and ore crushing energy consumption of multiple ore piles, and uses the KNA network or a combination of the DeepSet model and the KAN network model to learn and mine the relationship between ore crushing energy consumption and ore particle size distribution. This can accurately obtain an ore crushing energy consumption prediction formula that can be used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile. Based on the relationship between crushing energy consumption and ore particle size distribution represented by the ore crushing energy consumption prediction formula, the ore blasting plan can be modified accordingly to blast the ore when the possibility of crushing energy depletion is small, thereby achieving the goal of high efficiency and energy saving. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of a first embodiment of a method for obtaining a prediction formula for ore crushing energy consumption in an embodiment of the present application;
[0017] Figure 2 Schematic diagram of DeepSet network + Kan network in one embodiment of the present application;
[0018] Figure 3 This is a schematic diagram of the effect of ore target detection in one embodiment of the present application;
[0019] Figure 4This is a schematic diagram of ore target detection in a designated area in another embodiment of the present application;
[0020] Figure 5 This is a schematic diagram of the structure of the I3SS feature fusion module in one embodiment of the present application;
[0021] Figure 6 This is a schematic diagram of the network architecture of the target recognition network model in one embodiment of the present application;
[0022] Figure 7 This is a structural table of the backbone network in one embodiment of the present application;
[0023] Figure 8 Schematic diagram of the bounding box of the ore identified in one embodiment of the present application;
[0024] Figure 9 This is a schematic diagram of a camera calibration plate image in one embodiment of the present application;
[0025] Figure 10 A schematic diagram of the maximum inscribed ellipse of the identification frame and a schematic diagram of the equivalent sphere diameter in one embodiment of the present application;
[0026] Figure 11 This is a functional module diagram of an embodiment of a device for obtaining an ore crushing energy consumption prediction formula in an embodiment of the present application;
[0027] Figure 12 This is a schematic diagram of an embodiment of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] The environment in the ore crushing section is complex, and crushing is accompanied by a large amount of dust. Traditional image recognition has difficulty in clearly and accurately identifying ore in such a complex environment, which leads to excessive errors in the output ore particle size distribution, affecting subsequent energy consumption predictions.
[0030] When crushing ore, the crusher's energy consumption is primarily related to the ore's particle size distribution. Controlling ore particle size can ultimately control energy consumption in the beneficiation process. However, in existing technologies, the uncertainty of explosive blast energy leads to a wide range of ore particle size distribution, which in turn causes significant variability in crushing energy consumption. To minimize crusher energy consumption, it is necessary to understand the relationship between crushing energy consumption and ore particle size distribution.
[0031] By finding the relationship between crushing energy consumption and ore particle size distribution, the ore blasting plan can be changed accordingly to blast the ore when the crushing energy depletion is less likely.
[0032] refer to Figure 1 In one embodiment of the present application, a method for obtaining a prediction formula for energy consumption of ore crushing is provided. The method for obtaining a prediction formula for energy consumption of ore crushing includes:
[0033] S100: Constructing a data set based on the ore particle size distribution and corresponding ore crushing energy consumption of multiple groups of ore piles.
[0034] Specifically, visual data of the ore before crushing the ore in the ore pile and the ore crushing energy consumption generated by crushing the ore are obtained.
[0035] The ore in the ore pile may be blasted ore obtained in advance by explosive blasting.
[0036] More specifically, an ore pile contains a pile of ore to be crushed. Once the ore pile enters the ore crushing section, it is crushed using a crusher. Before the ore is crushed, the ore in the pile is photographed to generate visual data (pre-crushing visual data). This visual data can be either an image or a video of the ore.
[0037] Ore from the ore pile is conveyed by a conveyor belt into a crushing port for crushing by a crusher. Visual data of the ore entering the crushing port can be captured by a camera. The camera can be an explosion-proof high-definition camera.
[0038] The crusher consumes electricity to crush the ore, so it is also necessary to obtain energy consumption data of the energy used to crush the ore.
[0039] In one specific embodiment, the power consumption data of the crusher can be collected as energy consumption data. For example, the power consumption data of the ore in the ore pile can be collected from before to after crushing. The difference in power consumption data before and after crushing can accurately indicate energy consumption.
[0040] The ore particle size distribution of the ore included in the ore pile is obtained based on the ore visual data.
[0041] More specifically, the ore in the ore pile is successively fed into the crushing port by, for example, a conveyor belt and then crushed by a crusher. The shape and other data of the ore entering the crushing port before being crushed can be identified from the ore visual data (ore video or ore image). For example, the bounding box of the ore before crushing can be identified from the ore video or ore image using the YOLO series of target detection models, and the particle size and volume of the ore before crushing can be calculated based on the bounding box. This method can be used to dynamically track, count, and calculate ore data of the crushed ore in the ore pile before crushing, where the ore data includes the particle size and volume of the ore. The ore particle size distribution can be calculated based on the ore data of the crushed ore in the ore pile and the number of crushed ore.
[0042] For ore particles that are nearly spherical, particle size typically refers to the diameter of the ore particle. For irregular particles, particle size generally refers to the maximum linear distance between the ore particles. Ore particle size distribution includes multiple different ore distribution intervals and the maximum particle size within each ore distribution interval.
[0043] S200: Input the data set into the DeepSet network to obtain a first eigenvector about the ore particle size distribution and a second eigenvector about the corresponding ore crushing energy consumption, and input multiple groups of the first eigenvectors and the corresponding second eigenvectors into the KAN network to obtain an ore crushing energy consumption prediction formula; or, directly input the data set into the KAN network to obtain an ore crushing energy consumption prediction formula; the ore crushing energy consumption prediction formula is used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile.
[0044] Specifically, the ore particle size distribution and the ore crushing energy consumption are input into the ore crushing energy consumption prediction model to obtain the target relationship between energy consumption and ore particle size distribution.
[0045] Specifically, during ore crushing, the energy consumption of the crusher is usually related to the ore particle size distribution. This embodiment can obtain multiple sets of ore particle size distribution and energy consumption data of different ore piles, and input multiple sets of ore particle size distribution and energy consumption data into the trained ore crushing energy consumption prediction model. The ore crushing energy consumption prediction model is used to learn the relationship between ore particle size distribution and energy consumption, and then predict the ore crushing energy consumption prediction formula between ore particle size distribution and energy consumption.
[0046] More specifically, the ore crushing energy consumption prediction model includes a KAN network model, or the ore crushing energy consumption prediction model includes a DeepSet network model and a KAN network model.
[0047] The DeepSet network is a model for processing vector sets. It is permutation-invariant and suitable for non-vector samples such as graphs and point clouds. It learns a model through a combination of embedding and Dense layers, performing linear and nonlinear operations on the elements of a set. Permutation equivariant functions, a more sophisticated model, assign an independent output function to each element at each position, better capturing relationships between elements.
[0048] A sample is typically viewed as a vector. The sample label is then fed to the machine to learn the model. There are often samples that are not vectors, such as graphs, point clouds (matrices), continuous graphs (sets of 2D vectors), and text (sequences of vectors).
[0049] The object processed by the DeepSet network is a set S of vectors, and the output is a real number (or vector).
[0050] This embodiment solves the problem of energy consumption prediction in mine crushing by using DeepSet as the core modeling. Since the input data is essentially an unordered set, that is, the order of arrangement of the particle size does not affect the final energy consumption, the model needs to be insensitive to the input order. However, traditional neural networks (such as MLP or CNN) usually require a fixed input order and dimension, and it is difficult to directly process set data. The quantitative analysis functions and common solutions adopted in the prior art include padding the particle size set to a fixed length input, but this method may lead to the introduction of invalid information; or calculating statistical features (such as mean, variance, median, etc.) as input features, but it is easy to lose the complex relationship within the set. Therefore, this embodiment selects DeepSet as the modeling method. DeepSet adopts a neural network architecture based on permutation invariance. It first performs a nonlinear transformation on each element, and then uses aggregation functions such as summation to extract global features, thereby ensuring that the input order does not affect the output. At the same time, DeepSet has strong flexibility, does not require the input set size to be fixed, and can automatically learn richer feature representations, making it more suitable for the task of this application than traditional methods.
[0051] Specifically, the core idea of DeepSet is to map the input set x to a fixed-length representation through a permutation-invariant method and use it for downstream regression tasks. Its mathematical form is shown in the following formula 1:
[0052]
[0053] Among them, Φ(.) is the feature mapping function, which is parameterized by θ1.
[0054] g(.) is an aggregation function that ensures that the input order does not affect the output. It can typically take the average, maximum, or sum. Preferably, this embodiment uses max pooling. ρ(.) is a subsequent prediction function (such as an MLP) used to generate the final energy consumption regression value, parameterized by θ2.
[0055] In this example, DeepSet is applied to the regression problem of mine particle size data. Specifically, each sample represents a set of ore particle sizes, and the network output corresponds to the energy e required to crush these ores i The introduction of the DeepSet structure enables the model to automatically learn the relationship between particle size and energy consumption while maintaining the invariance of the set data.
[0056] To optimize the DeepSet model and ensure that it can accurately predict the energy consumption of ore crushing, this example uses the mean squared error (MSE) as the loss function. The mean squared error is the most common optimization objective in regression tasks and is defined as shown in the following formula 2:
[0057]
[0058] Formula 2
[0059] Where N is the total number of training samples; e i represents the actual ore crushing energy consumption of the i-th sample; is the energy consumption value predicted by the model; θ represents the trainable parameters of the model (including and ρ).
[0060] The goal of the mean squared error loss function is to minimize the predicted value and the true value e i The sum of squared errors between the two sets of samples improves the model’s fitting ability. In addition, since MSE imposes a greater penalty when the error is large, it can effectively guide the model to reduce the impact of high-error samples and improve the overall prediction accuracy.
[0061] During the training process, the Adam optimizer can be used for gradient descent optimization to accelerate convergence and avoid local optimality problems. The optimization process can be expressed as shown in Formula 3:
[0062]
[0063] Formula 3
[0064] Where η is the learning rate, represents the gradient of the loss function.
[0065] Specifically, compared with traditional modeling methods, DeepSet is a black box model as a whole, which is not interpretable. It is also difficult to find out which links play a more important role in the processing of ore particles on the processing line.
[0066] DeepSet provides an effective and adaptive statistical extraction method for ore particle size identification data. However, the model learned by DeepSet still lacks interpretability. To solve this problem, this embodiment further combines the KAN network to ensure that the model input has sufficient information to prevent the loss of model prediction accuracy due to the omission of model input information. It also takes into account that the model as a whole can output a visible formula, making the output structure more interpretable.
[0067] Among them, the KAN network is Kolmogorov-Arnold Networks. The main theoretical basis of KAN is the Kolmogorov-Arnold theorem, which is called the Kolmogorov-Arnold representation theorem in Chinese.
[0068] Figure 2 This is a schematic diagram of the DeepSet network + Kan network in an embodiment of the present application; Figure 2 There are n groups of ore piles (ore piles) on the ore particle processing line, namely ore pile 1, ore pile 2...ore pile n. The ore particle size distribution and the corresponding ore crushing energy consumption of each group of ore piles are obtained.
[0069] A dataset is constructed using the ore particle size distribution and corresponding ore crushing energy consumption of n groups of ore piles. The ore particle size distribution and corresponding ore crushing energy consumption in the dataset are converted into vector data tensor-k1, tensor-k2, tensor-k3, ..., tensor-kn, respectively. These vector data are uniformly standardized to obtain the standardized vector data tensor-k1, tensor-k2, tensor-k3, ..., tensor-kn. The standardized vector data tensor-k1, tensor-k2, tensor-k3, ..., tensor-kn are input into the DeepSet model to train the DeepSet model, and the intermediate vectors tensor-t1, tensor-t2, tensor-t3, ..., tensor-tn are obtained as the output of the trained DeepSet model. Tensor-t1, tensor-t2, tensor-t3, ..., tensor-tn are input into the KAN network model, which attempts to find an explicit functional relationship between the ore particle size distribution and the corresponding ore crushing energy consumption. The KAN network model retains the optimal branches through pruning and constructs a visible and interpretable prediction formula for ore crushing energy consumption.
[0070] In this process, the ore particle size distribution and the corresponding ore crushing energy consumption in the data set are converted into a matrix X and output to the DeepSet network (DeepSetModel). Let h be The DeepSet model normalizes vectors of inconsistent dimensions into a fixed length, that is, extracts the data statistics of the set. The intermediate vectors Tensor-Ti output by the DeepSet model serve as input to the KAN network. The KAN network then performs pruning and filtering of basic data symbols to produce a visual formula after training. The DeepSet model is a pre-trained model. Each set of intermediate vectors Tensor-Ti includes a first eigenvector of the ore particle size distribution and a second eigenvector of the corresponding ore crushing energy consumption.
[0071] This embodiment combines the advantages of the DeepSet and KAN models, achieving good performance while also providing a certain degree of interpretability.
[0072] This application adopts a combination of DeepSet+Kan network to convert the ore particle size distribution data into vectors and input them into DeepSetModel. The DeepSet model uniformly standardizes the vectors with inconsistent dimensions into a fixed length, and then outputs the result tensor-ki to the DeepSet model for alignment training. After the model training is completed, the intermediate vector Tensor-Ti output by the trained model is input into the KAN network. Due to the KAN network, pruning and screening of basic data symbols are performed, and after training, a visual formula for energy consumption and ore particle size distribution is obtained. Alternatively, in another embodiment, the data set consisting of the ore particle size distribution and the corresponding ore crushing energy consumption can be directly input into the KAN network to obtain the corresponding ore crushing energy consumption prediction formula.
[0073] Before the data set is input into the KAN network, it can be first input into the gradation statistical model to obtain gradation data, such as a low-dimensional vector (such as a 7-dimensional vector), and then the low-dimensional vector is input into the KAN network.
[0074] The gradation statistical model is an empirically designed statistical model. For example, data such as D5 and D20 are obtained. For example, D5 represents the ore particle size corresponding to 5% of the total mass, and D20 represents the ore particle size corresponding to 20% of the total mass. The specific data depends on the actual configuration and is not limited by this application.
[0075] It should be noted that inputting a dataset into DeepSet yields high-dimensional features, which is different from inputting a dataset into a graded statistical model.
[0076] This embodiment obtains the ore particle size distribution and ore crushing energy consumption from multiple ore piles, and uses a KNA network or a combination of a DeepSet model and a KAN model to learn and mine the relationship between ore crushing energy consumption and ore particle size distribution. This accurately obtains an ore crushing energy consumption prediction formula that can be used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile. Based on the relationship between crushing energy consumption and ore particle size distribution expressed in the ore crushing energy consumption prediction formula, the ore blasting plan can be modified accordingly to blast the ore when the likelihood of crushing energy depletion is low, thereby achieving efficient energy conservation.
[0077] In one embodiment, before constructing a data set using the ore particle size distributions and corresponding ore crushing energy consumptions of multiple groups of ore piles in step S100, the method for obtaining the ore crushing energy consumption prediction formula further includes:
[0078] Using the visual data of the ore pile as input, the target recognition network model is used to detect the ore target on the ore pile to obtain the ore target detection result. The target recognition network model is built based on the YOLO model.
[0079] Track each ore in the ore target detection results and calculate the ore particle size to obtain the ore particle size distribution of the ore pile;
[0080] Get the ore crushing energy consumed to complete ore crushing in the ore pile.
[0081] Specifically, the electric meter readings before and after ore crushing and the time of electric meter reading are obtained; the ore crushing energy consumption consumed to complete the ore crushing can be obtained based on the electric meter readings before and after the ore crushing; the original ore visual data is obtained, and the original ore visual data is the video before ore crushing; the video before ore crushing is time-synchronized and aligned according to the time of electric meter reading and then cropped to obtain ore visual data; the ore visual data is input into the trained target recognition network model for ore target detection to obtain ore target detection results, wherein the ore visual data is the cropped video before ore crushing, or the ore visual data is the image frame intercepted from the cropped video before ore crushing; each ore is tracked and ore data is calculated according to the ore target detection results, wherein the ore data includes the number of ores, the ore particle size of each ore and the ore volume; the ore particle size distribution is obtained according to the ore data.
[0082] More specifically, the crusher is powered by electricity and supplied by a power supply. Video data or images of the electricity meter of the power supply can be collected. For example, video data or images of the electricity meter can be collected using a miniature smart camera.
[0083] The meter reading time and meter reading can be extracted from the meter video data using an optical character recognition algorithm. Alternatively, the meter reading can be identified from a meter image using an optical character recognition algorithm. Furthermore, the timestamp of the meter image can be obtained, which is the meter reading time.
[0084] A high-quality dataset can be formed by extracting meter time and readings through optical character recognition algorithms, cropping the video with a sliding window, and synchronizing the meter readings with the ore video time.
[0085] By calculating the difference in the electric meter readings before and after the ore is crushed, we can get the electric energy consumed by the ore crusher that crushes the ore pile, that is, the ore crushing energy consumption.
[0086] Visually capturing meter data and pre-crushed ore data isn't always completely synchronized. To ensure synchronization and consistency between the two captured data, the embodiment crops the pre-crushed ore video based on the times before and after the meter readings. The cropped video includes video data from before and after ore crushing, that is, between the times before and after the meter readings.
[0087] In addition, cropping the video can reduce the amount of data processing and improve the efficiency of subsequent steps such as ore target detection.
[0088] The video before ore crushing includes the video of the ore from the ore pile entering the crushing port one by one. The ore entering the crushing port will be crushed by the crusher.
[0089] The two meter reading times include the meter reading time before the ore in the ore pile starts to be crushed and the meter reading time after the ore in the ore pile is successively crushed by the crusher and the crusher stops crushing.
[0090] The cropped video of the ore before crushing is obtained, and the cropped video of the ore before crushing can be input into the trained target recognition network model to perform ore target detection and ore target detection results.
[0091] Alternatively, image frames can be captured from the cropped video before ore crushing, and the image frames can be input into the trained target recognition network model for ore target detection to obtain the ore target detection results.
[0092] The ore target detection result includes, for example, a predicted bounding box of the identified ore passing through the crushing opening. The ore target detection result may also include, for example, the category and confidence level of the identified ore passing through the crushing opening.
[0093] In one specific embodiment, a target recognition network model can be trained to identify ores in specific areas within a video or image. For example, it can identify ores entering a crushing opening or passing through a predetermined straight line. This can reduce the target detection range when there are a large number of ores, improve target detection accuracy, reduce repeated detection, and reduce computational overhead.
[0094] Figure 3 This is a schematic diagram of the effect of ore target detection in one embodiment of the present application; Figure 4 This is a schematic diagram of ore target detection in a designated area in another embodiment of the present application; Figure 3 ,The trained ore recognition module can be used to detect and identify the ores in the video or picture, and the identified ores are selected using the predicted bounding box.
[0095] The more ore that enters the crusher, the greater the energy consumption of the crusher. Ore that does not enter the crushing port may not be crushed by the crusher, and thus will not consume electricity. If the uncrushed ore is used as the detection target, it will cause an error between energy consumption and ore particle size distribution. Based on this, refer to Figure 4, you can set up target detection for the ore near the No. 2 jaw break (crushing port), that is, detect the ore passing through the No. 2 jaw break (crushing port). This can prevent the ore that has not entered the crushing port from being detected, ensuring that the identified ore is all crushed ore, thereby ensuring the consistency of ore particle size distribution and energy consumption.
[0096] Alternatively, the ore target detection result includes the predicted bounding boxes of all ores that can be identified by the target recognition network model, and when tracking the ore and calculating the ore data, only the ore passing through the No. 2 jaw rupture (crushing opening) is tracked and the ore data is calculated.
[0097] In one specific embodiment, the BoTSORT (Bytetrack with Occlusion and Scale-aware Re-identification) algorithm is used to track the motion trajectory of ore particles in real time. BoTSORT effectively addresses tracking interruptions caused by occlusion or target motion, enabling ore flow statistics in videos or images. The BoTSORT algorithm tracks the ore identified by the target recognition network model in real time and performs unique counts to determine the number of ore particles entering the crusher.
[0098] In addition, based on the predicted bounding box of the ore, the ore particle size and ore volume of the identified ore target can also be calculated.
[0099] According to the ore quantity, ore particle size and ore volume, the ore particle size distribution of the ore pile can be obtained.
[0100] In one specific embodiment, the ores are sorted by volume, for example, in ascending or descending order. The maximum particle size of the ores ranked between 0-p1% by volume is taken; the maximum particle size of the ores ranked between 0-p2% by volume is taken; the maximum particle size of the ores ranked between 0-p3% by volume is taken; and the maximum particle size of the ores ranked between 0-pm% by volume is taken. In this way, the maximum particle size of multiple different ore distribution intervals can be obtained, thereby determining the ore particle size distribution of the ore pile. Here, p1, p2, p3, ..., pm increase in order.
[0101] For example, the maximum particle size of each of the seven ore distribution intervals of 0-5%, 0-20%, 0-50%, 0-63.2%, 0-75%, 0-80%, and 0-90% is taken.
[0102] This embodiment can perform target detection on the ore in the ore visual data through the target recognition network model, and can obtain accurate ore particle size distribution by tracking the ore detected in the target detection results and calculating the ore data, and can also accurately obtain the ore crushing energy consumption of each ore pile.
[0103] In one embodiment, the object recognition network model includes a backbone network, a neck network, and an object detection head;
[0104] The backbone network is used to extract features from ore visual data and obtain feature maps of different scales;
[0105] The neck network is used to perform multi-scale feature fusion on the feature map to obtain a fused feature map;
[0106] The target detection head is used to detect and locate ore targets based on the fused feature map to obtain ore target detection results.
[0107] Specifically, the target recognition network model of this embodiment can be built based on the YOLO series of target detection models, for example, based on one of the YOLO V5 network model and the YOLO V8 network model. YOLO (you only look once) is a target detection and image segmentation model. YOLOv8 supports a full range of visual AI tasks, including classification, detection, segmentation, pose estimation, and tracking. The YOLO model can run on hardware platforms ranging from CPUs to GPUs.
[0108] The backbone network, also known as the feature extraction backbone, consists of multiple sequentially organized convolutional layers that extract relevant features from the input image. This is used for image feature extraction and produces feature maps at different scales. The feature maps extracted by the backbone network serve as the base feature maps.
[0109] The neck network is used to fuse the feature maps of different stages of the backbone network to enhance the feature representation capability.
[0110] The target detection head is used for decoding, including bbox bounding box decoding and class loss classification decoding, that is, to achieve target detection, including predicting bounding boxes, categories and confidence levels.
[0111] The fused feature map is input into the detection head network in the target recognition network model to perform bounding box regression reasoning and target classification reasoning, and output the bounding box correction value and category probability value of the predicted box.
[0112] This embodiment builds a target recognition network model based on the YOLO model architecture, which can accurately identify ore targets in visual data.
[0113] In one embodiment, the target recognition network model includes a backbone network,
[0114] The backbone network includes a feature extraction backbone, or the backbone network includes a dust removal network layer and a feature extraction backbone connected in sequence;
[0115] The dust removal network layer is used to perform dust removal processing on the ore visual data to obtain dust-free visual data;
[0116] The feature extraction backbone is used to extract features from the cropped visual data or the dust-removed visual data to obtain feature maps of different scales.
[0117] Specifically, the target recognition network model is built based on the YOLO network model, for example, the target detection algorithm based on YOLOv8, and is designed and improved accordingly based on the characteristics and difficulties of blasting ore detection in the mine pile.
[0118] To address the problem of large amounts of dust blocking the field of view that often occurs during ore transportation, a multi-level feature fusion dehazing network or dust removal network (FFA-Net) that fuses channel attention and pixel attention is introduced.
[0119] The dust removal network layer is also called the defogging network layer, which can include the FFA-Net module.
[0120] FFA-Net (Feature Fusion Attention Network for single image) feature fusion attention network.
[0121] The key concept of FFA-Net is to directly restore dust-free videos or images using a Feature Fusion Attention Network. Through its innovative design of a feature attention mechanism and a feature fusion attention architecture, FFA-Net effectively improves the model's performance in removing dust and haze from videos or images. By cleverly combining channel and pixel attention, along with local residual learning, the network can more accurately handle haze or dust in different regions, achieving significant improvements in detail preservation and color fidelity.
[0122] The FFA-Net architecture consists of three key components: (1) a novel feature attention (FA) module that combines channel attention and pixel attention mechanisms, taking into account that different channel features contain completely different weighted information and that dust is unevenly distributed across different image pixels. FA treats different features and pixels unequally, which provides additional flexibility for processing different types of information. This expands the representation capabilities of CNNs. (2) a basic block structure consisting of local residual learning and feature attention. Local residual learning allows less important information such as light dust areas or low frequencies to be bypassed through multiple local residual connections, allowing the main network architecture to focus on more effective information. (3) an attention-based feature fusion (FFA) structure at different levels. Feature weights are adaptively learned through the feature attention (FA) module, giving more weight to important features. This structure can also retain shallow information and pass it to deep layers.
[0123] This embodiment builds an efficient feature extraction backbone by adding a dust removal network layer (FFA-Net) to the backbone network.
[0124] The model is optimized for scenes with large amounts of smoke and dust to improve the accuracy of ore detection. In ore particle size detection, a large amount of suspended particles such as smoke, smog and mist will be generated due to ore crushing. These interference factors can easily cause problems such as image or video color distortion, blurring and contrast reduction, which significantly affect the quality of the image or video. This degradation in image or video quality increases the difficulty of the target detection task and may even cause a large number of gravel target detection failures. In order to solve this problem, this embodiment introduces the FAA-Net dust removal network layer into the model structure to pre-process the input damaged image or video, thereby restoring a clearer image or video to assist in the implementation of subsequent detection tasks.
[0125] The FFA-Net module, or image dust removal feature fusion network, consists of multiple modules. Its input is a blurred image, which is first processed by a shallow feature extraction module and then passed to N groups of architectures containing multiple skip connections. Within these architectures, the output features of each group are fused through a feature attention module (FA). The fused features are then passed to a reconstruction module and a global residual learning structure to generate a clear dust-free image output (haze-free image output). Each group of architectures consists of B basic blocks and a local residual learning module. The basic blocks combine skip connections with a feature attention mechanism. The local residual learning within the basic blocks bypasses secondary information such as haze or low-frequency areas through multiple local residual connections, allowing the main network to focus on more critical feature information. The FeatureAttention module includes CA (ChannelAttention) and a pixel attention mechanism. This attention mechanism provides greater flexibility in processing different types of information.
[0126] The embedded CA attention mechanism focuses on target coordinate information related to channel information and direction, weaving coordinate dimension information into the channel attention mechanism to capture information within a wider area. This attention mechanism adopts a split strategy, concretizing it into a parallel process focused on one-dimensional feature encoding to reduce the position information loss caused by 2D global pooling. The CA module uses two pooling kernels of different sizes, (1, W) and (H, 1), to encode the spatial coordinates of each channel in the horizontal and vertical directions, respectively. This enables the module to efficiently aggregate features in two spatial dimensions, ensuring that the generated feature map retains key position information. The calculation formula is shown in Equation 4, followed by a convolution transformation using Equation 5, and the output is shown in Equation 6.
[0127]
[0128] in, and (indexed by height h and width w) is the output value of the cth channel. f is the intermediate feature, F h and F w is the corresponding convolution operation. δ is used to introduce nonlinear elements. [.,.] performs feature concatenation in the spatial dimension. The sigmoid activation function σ performs nonlinear transformation on the intermediate features, mapping them to the [0,1] interval, thereby generating the weight g h and g w Finally, these weights are combined with the original feature map to produce the output y of the CA module c .
[0129] Given the uneven distribution of gravel dust across image pixels, the pixel attention mechanism (PA) focuses on extracting specific informative features, such as pixels in areas with high dust content and high-frequency image details. PA processes the input of the channel attention (CA) by passing it through two convolutional layers, which use ReLU and Sigmoid activation functions, respectively, ultimately transforming the data shape from C×H×W to 1×H×W.
[0130] PA=σ(Conv(δ(Conv(y c ))))
[0131] Formula 7
[0132] Finally, y c and PA are element-wise multiplied, where F a is the output of the FA module.
[0133]
[0134] The captured ore images are preprocessed using the FFA-Net dust removal network, which makes the high-frequency components more prominent and highlights the details in the original image. The ore objects in the images are clearer, with more distinct outlines and stronger contrast, improving the recognition accuracy of the target detection algorithm.
[0135] The clearer feature map obtained after dust removal of the original image through the FFA-Net network is input into the feature extraction backbone. From low to high layers, the receptive field of the ore image is gradually expanded to fully capture the ore features of all scales.
[0136] This embodiment introduces the FFA-Net dust removal network into the backbone network, constructs an efficient dust removal backbone network that integrates feature attention and a composite scalable backbone, enables the model to efficiently extract features of ore images, and enhances the model's robustness to dust interference during ore crushing.
[0137] In one embodiment, the feature extraction backbone included in the backbone network includes a Conv convolution layer, a shifted flip bottleneck convolution layer, a fast convolution layer, and an SPPF layer connected in sequence;
[0138] Among them, the Conv convolution layer includes a Conv module, and the input of the Conv module is visual data input;
[0139] The mobile flip bottleneck convolution layer includes at least two mobile flip bottleneck convolution blocks connected in sequence, and the fast convolution layer includes at least two fast convolution blocks; or, the sum of the number of convolution blocks of the mobile flip bottleneck convolution block contained in the mobile flip bottleneck convolution layer and the fast convolution block contained in the fast convolution layer is not less than 4; wherein each mobile flip bottleneck convolution block includes an MBConv module; each fast convolution block includes multiple stacked FasterNet network modules; the SPPF layer includes an SPPF module.
[0140] Specifically, the target recognition network model is built based on the YOLO network model, for example, the object detection algorithm based on YOLOv8, and is designed and improved to address the characteristics and difficulties of blasted ore detection in ore piles. To address the problem of large amounts of dust obstructing the field of view that often occurs during ore transportation, a multi-level feature fusion dehazing network, or FFA-Net, is introduced that combines channel attention and pixel attention. Furthermore, a scalable ore feature extraction backbone is constructed using the MBConv module based on depthwise separable convolution.
[0141] The target recognition network model of this embodiment is an improved target detection model based on the YOLO architecture. For example, the target recognition network model of this embodiment can be obtained by improving YOLO V8.
[0142] The feature extraction backbone included in the backbone network is a composite scalable high-level multi-level feature extraction backbone.
[0143] A Conv convolutional layer includes a Conv module. For example, it contains a Conv module. A Conv module, also known as a convolutional module, includes a convolutional layer, a batch normalization layer, and an activation function (i.e., a Conv2d layer, a BatchNorm2d layer, and a SiLu activation function), implementing the convolution-->normalization-->activation function. Convolution: Uses a 2D convolution operation (nn.Conv2d) to extract local features.
[0144] Normalization: Use BatchNormalization (nn.BatchNorm2d) to accelerate training and improve stability. Activation: Use the nonlinear activation function SiLU to introduce nonlinear capabilities. The visual data input to the Conv convolution layer can be cropped visual data or dust-removed visual data. The dust-removed visual data is the visual data obtained by dust removal from the cropped visual data. The Conv convolution layer performs a preliminary convolution calculation on the visual data input to obtain a first intermediate feature map, which is then transmitted to the moving flip bottleneck convolution layer.
[0145] The moving flip bottleneck convolution layer may include 2 moving flip bottleneck convolution blocks, and the fast convolution layer may include a convolution block of 2 fast convolution blocks; or, the moving flip bottleneck convolution layer may include 1 moving flip bottleneck convolution block, and the fast convolution layer may include a convolution block of 3 fast convolution blocks; or, the moving flip bottleneck convolution layer may include 2 or 3 moving flip bottleneck convolution blocks, and the fast convolution layer may include a convolution block of 3 or 4 fast convolution blocks, etc.
[0146] The moving flip bottleneck convolution layer is an MBConv convolution layer (or, MBConv layer), and the moving flip bottleneck convolution layer includes at least two moving flip bottleneck convolution blocks connected in sequence, wherein each moving flip bottleneck convolution block includes an MBConv module.
[0147] The moving flip bottleneck convolution block is the moving flip bottleneck subconvolution layer or MBConv Block.
[0148] Each moving flip bottleneck convolution block contains at least one MBConv module.
[0149] The MBConv module stands for mobile inverted bottleneck convolution. Its channel expansion and compression rates are adjustable, making it an inverted linear bottleneck layer with depthwise separable convolution. Based on this depthwise separable convolution, the MBConv module enables the construction of a scalable feature extraction backbone.
[0150] FasterNet is a new neural network backbone model designed to improve computational efficiency by reducing redundant computations and memory accesses. FasterNet's core innovation lies in the introduction of partial convolution (PConv), a convolution operator that convolves only a subset of the channels in the input feature map while leaving the rest unchanged, thereby reducing computational complexity.
[0151] In a specific embodiment, the moving flip bottleneck convolution layer includes three moving flip bottleneck convolution blocks connected in sequence. The first moving flip bottleneck convolution block includes an MBConv module. The input of the first moving flip bottleneck convolution block is the first intermediate feature map output by the Conv convolution layer, and the output of the first moving flip bottleneck convolution block is the second intermediate feature map.
[0152] The second moving flip bottleneck convolution block and the third moving flip bottleneck convolution block each include at least two stacked MBConv modules, for example, include 3 stacked MBConv modules.
[0153] The input of the second moving flip bottleneck convolution block is the second intermediate feature map output by the first moving flip bottleneck convolution block, and the output is the third intermediate feature map.
[0154] The input of the third moving flip bottleneck convolution block is the third intermediate feature map output by the second moving flip bottleneck convolution block, and the output is the fourth intermediate feature map.
[0155] In a specific embodiment, the MBConv module may be an MBConv 3×3 layer.
[0156] The fast convolution layer includes at least two fast convolution blocks, and the fast convolution block is a fast sub-convolution layer.
[0157] Each fast convolution block consists of multiple stacked FasterNet network modules.
[0158] In a specific embodiment, the fast convolution layer includes 4 fast convolution blocks (FasterNet Block).
[0159] The input of the first fast convolution block is the output of the moving flip bottleneck convolution layer, for example, the fourth intermediate feature map output by the third moving flip bottleneck convolution block, and its output is the fifth intermediate feature map.
[0160] The input of the second fast convolution block is the output of the first fast convolution block, for example, the input is the fifth intermediate feature map, and its output is the sixth intermediate feature map.
[0161] The input of the third fast convolution block is the output of the second fast convolution block, for example, the input is the sixth intermediate feature map, and its output is the seventh intermediate feature map.
[0162] The input of the fourth fast convolution block is the output of the third fast convolution block, for example, the input is the seventh intermediate feature map, and its output is the eighth intermediate feature map.
[0163] In a specific embodiment, the first fast convolution block and the second fast convolution block each include 4 stacked FasterNet network modules; the third fast convolution block includes 5 stacked FasterNet network modules; and the fourth fast convolution block includes 2 stacked FasterNet network modules.
[0164] FasterNet network module is FasterNet Block.
[0165] The SPPF layer includes the SPPF module, or Spatial Pyramid Pooling with Fixed-size feature maps. This module performs pooling operations at different scales, concatenating feature maps of different scales to improve detection of objects of varying sizes. The SPPF module acquires more comprehensive spatial information through weighted fusion of global and local features.
[0166] In a specific embodiment, the input of the SPPF module is the output of the fast convolution layer, for example, the eighth intermediate feature map output by the fourth fast convolution block, and the output of the SPPF module is the ninth intermediate feature map.
[0167] In a specific embodiment, the MBConv module includes: a 1x1 ordinary convolution (dimensionality increase, including BN and Swish activation), a Depthwise Conv convolution (including BN and Swish activation), an SE module, a 1x1 ordinary convolution (dimensionality reduction, including BN and linear activation, linear activation y=x) and an add operation.
[0168] The first 1x1 convolution layer is used to increase the dimension and increase the number of channels of the input feature map to expand the feature dimension. Depthwise Convolution: Use depthwise separable convolution to perform independent convolution operations on each channel, which greatly reduces the amount of computation. SE module: Consists of a global average pooling and two fully connected layers. The features are recalibrated through the SE-Net module to enhance important features and suppress unimportant features. The SE-Net module includes Squeeze and Excitation operations. The former compresses the feature map to the channel dimension, and the latter recalibrates the features by learning the weights of each channel. 1x1 dimensionality reduction convolution: Finally, 1x1 convolution is used to reduce the number of channels of the feature map to reduce the amount of computation and the number of parameters. Add operation: Add the input feature map to the output feature map to form a residual connection, which helps the gradient backpropagation and improves the training stability.
[0169] refer to Figure 6 , the MBConv module includes expansion convolution, depthwise convolution and projection convolution connected in sequence.
[0170] Alternatively, the MBConv module includes expansion convolution, depthwise convolution, and projection convolution connected in sequence, and the input and output ends use residual connections.
[0171] Expansion convolution is primarily used in deep learning models to increase the dimensionality and number of channels of feature maps, enabling better feature extraction in subsequent depthwise separable convolutions. By increasing the number of channels, the model can learn more feature representations, thereby improving its expressiveness and generalization capabilities.
[0172] Depthwise Convolution is a technique used in convolutional neural networks, which is mainly used to reduce the amount of calculation and the number of parameters while maintaining the ability of feature extraction.
[0173] Depthwise convolution is a lightweight convolution operation that reduces computational complexity by splitting standard convolution into two parts: depthwise convolution and pointwise convolution. Depthwise separable convolution is widely used in mobile applications, effectively reducing computational complexity and parameter count, while improving model inference speed and operational efficiency.
[0174] In contrast to Expansion Convolution, Projection Convolution outputs a much smaller number of channels than the input, thereby limiting the size of the model. Using Projection Convolution can ensure that the number of channels does not increase too much, and can even keep the number of output channels equal to the number of input channels.
[0175] The core idea of residual connections is to introduce a "shortcut" or "skip connection," allowing the input signal to bypass some layers and be added to their outputs. This way, the network no longer needs to learn a complete function mapping input to output, but instead learns a residual function, which is the difference between the input and the desired output. This structure simplifies the optimization process and makes training more stable. Residual connections are usually implemented as short-circuit connections, where the input is directly added to the layer's output.
[0176] In a specific embodiment, the MBConv module, i.e., the MBConv basic block, first performs a 1×1 convolution operation to expand the number of channels of the input ore feature map, then uses an efficient depth-wise separable convolution operation to extract features such as ore edges and textures, and finally performs another 1×1 convolution operation to reduce the number of channels of the ore feature map for calculation by subsequent modules.
[0177] During this process, the channel expansion rate and compression rate of each MBConv module can be adjusted, so the entire network can be compound scaled. Then, through grid search, the network structure with the highest efficiency and the best mineral image feature extraction performance is found, thereby constructing an efficient feature extraction backbone network.
[0178] If the dust removal network (FFA-Net) is added to the efficient feature extraction backbone network, a dust removal and efficient feature extraction backbone network can be constructed.
[0179] The overall expression of the backbone network is shown in Formula 9 below:
[0180]
[0181] Among them, H i ×W i is the input ore feature map resolution, C i is the number of channels of the input ore characteristic map, X is the operation F i Input information, E i is the dimension magnification of 1×1 convolution, F i is the operation corresponding to network layer i, L i The number of network layers to operate on.
[0182] The Conv convolution layer in this embodiment is used to further process the feature map. The MBConv convolution layer is a lightweight convolution module. FasterNetBlock is a faster network block used to accelerate calculations.
[0183] This embodiment improves the backbone network based on the YOLO model architecture and constructs a composite and scalable ore feature extraction backbone based on the MBConv module of deep separable convolution, which can efficiently and accurately extract feature maps from ore visual data.
[0184] In one embodiment, the target recognition network model further includes a neck network, which includes a feature pyramid network (FPN) and a path aggregation network (PAN, PANet).
[0185] The feature pyramid network (FPN) includes at least two first feature fusion modules;
[0186] The path aggregation network (PAN) includes at least two second feature fusion modules,
[0187] Each first feature fusion module includes an upsampling module, a Concat module and an I3SS feature fusion module connected in sequence;
[0188] Each second feature fusion module includes a Conv module, a Concat module and an I3SS feature fusion module connected in sequence;
[0189] The I3SS feature fusion module is constructed by combining the convolutional model and the sequence model. The I3SS feature fusion module includes the I3 module, SS2D module and MLP module connected in sequence;
[0190] Among them, the I3 module includes the SPLIT module, Conv module, LN module and BN module.
[0191] Feature Pyramid Network (FPN) is used to start from deep features, upsample layer by layer, obtain upsampled features, fuse the upsampled features of each layer with the corresponding low-level features, obtain intermediate features or part of the final features, and construct a top-down feature pyramid through FPN to achieve preliminary fusion of multi-scale features.
[0192] In one embodiment, the target recognition network model further comprises a neck network;
[0193] The neck network includes a first feature fusion network and a second feature fusion network;
[0194] The first feature fusion network includes at least two first feature fusion modules;
[0195] The second feature fusion network includes at least two second feature fusion modules;
[0196] Each first feature fusion module includes an upsampling module, a Concat module and an I3SS feature fusion module connected in sequence;
[0197] Each second feature fusion module includes a Conv module, a Concat module and an I3SS feature fusion module connected in sequence;
[0198] The I3SS feature fusion module is constructed by combining the convolutional model and the sequence model. The I3SS feature fusion module includes the I3 module, SS2D module and MLP module connected in sequence;
[0199] Among them, the I3 module includes the SPLIT module, Conv module, LN module and BN module.
[0200] Specifically, the original neck network of the original YOLO network, such as YOLOv8, fuses three feature maps derived from the three lower-resolution feature layers of 80×80, 40×40, and 20×20 in the high-level backbone. It then uses upsampling and convolutional downsampling to form a bidirectional path to fully integrate the different feature information in the high- and low-resolution layers. This structure is suitable for general target detection, but blasting ore targets exhibit drastic scale variations. Many small-sized ore features are lost in the low- and medium-resolution feature maps in the high-level layers. Furthermore, because the bidirectional path is too long, the blasting ore features must cross at least two lateral connections from the backbone to the detection head. This process loses much feature information of fine ore, resulting in a decrease in ore detection rate.
[0201] Based on this, the target recognition network model of this embodiment is built based on the YOLO network model, for example, the target detection algorithm based on YOLOv8, and is designed and improved accordingly to address the characteristics and difficulties of blasted ore detection in ore piles. To address the dense accumulation of blasted ore and the dramatic changes in ore particle size, a Vmamba model combining convolution and a sequence-based model is introduced, and a neck feature fusion network (neck network) using the BiFPN cross-path bidirectional feature fusion architecture is adopted.
[0202] This embodiment improves the original neck network based on the YOLO architecture to obtain a neck network using the BiFPN cross-path bidirectional feature fusion architecture. The neck network using the BiFPN cross-path bidirectional feature fusion architecture includes a first feature fusion network and a second feature fusion network.
[0203] The first feature fusion network is an architecture designed to enhance the performance of object detection and image segmentation. It utilizes the outputs from different convolutional layers to perform feature fusion, allowing for better detection of objects of different scales.
[0204] The goal of the first feature fusion network is to detect objects at multiple scales using not only the output of the final convolutional layer, but also the output of several intermediate layers. This multi-scale detection capability is the key to the effectiveness of the first feature fusion network.
[0205] Bidirectional Feature Pyramid Network (BiFPN) establishes bidirectional connections between top-down and bottom-up paths, allowing for more efficient flow, exchange, and fusion of information between features of different scales. This allows for easy and rapid fusion of multi-scale features and optimizes cross-scale connections. This design improves the efficiency and effectiveness of feature fusion by strengthening the bidirectional flow of features, thereby improving object detection performance.
[0206] Compared with the traditional unidirectional feature pyramid network, BiFPN can more efficiently fuse features between different levels without significantly increasing the computational cost.
[0207] The first feature fusion network is used to start from the deep features, upsample layer by layer, obtain upsampled features, fuse the upsampled features of each layer with the corresponding low-level features, obtain intermediate features or part of the final features, and construct a top-down feature fusion network through the first feature fusion network to achieve preliminary fusion of multi-scale features.
[0208] The first feature fusion network allows the model to understand images at multiple scales.
[0209] The second feature fusion network is used to start from the bottom-level features, transfer features layer by layer, and fuse the feature maps of each layer with the high-level features through a low-to-up path.
[0210] Based on the first feature fusion network, the bottom-up path of the second feature fusion network is introduced to pass the underlying features to the higher layers layer by layer, further enriching the multi-scale features. Through horizontal connections, features of different scales are fused to ensure that the features of each layer contain rich contextual information.
[0211] By fusing bidirectional paths, the feature map contains richer contextual and semantic information, which enhances the model's ability to detect objects of different scales.
[0212] The second feature fusion network is an improvement to the first feature fusion network architecture. The second feature fusion network strengthens the connections between different feature scales and introduces additional mechanisms to better aggregate information. The second feature fusion network ensures that all information (from the top and bottom) is thoroughly aggregated.
[0213] The first feature fusion network of this embodiment includes at least two first feature fusion modules; each first feature fusion module includes an upsampling module (Upsample), a Concat module and an I3SS feature fusion module connected in sequence. The input of the first first feature fusion module includes the output of the SPPF module of the backbone network and the intermediate feature maps output by other layers of the backbone network. The inputs of the other first feature fusion modules all include the output of the previous first feature fusion module and the intermediate feature maps output by other layers of the backbone network. Among them, each first feature fusion module includes an upsampling module (Upsample), a Concat module and an I3SS feature fusion module.
[0214] The input of the upsample module (Upsample) of the first feature fusion module is the output of the SPPF module of the backbone network.
[0215] The input of the upsampling modules (Upsample) of other first feature fusion modules is the output of the previous first feature fusion module.
[0216] The input of the Concat module of each first feature fusion module includes the output of the upsampling module (Upsample) in the same group and the intermediate feature maps output by other layers of the backbone network. The Concat module is a splicing module used for concat operation.
[0217] The output of the Concat module of each first feature fusion module is transmitted to the I3SS feature fusion module in the same group.
[0218] The output of the I3SS feature fusion module of the last first feature fusion module is transmitted as input to the object detection head and the second feature fusion network.
[0219] The output of the I3SS feature fusion module of the other first feature fusion modules is transmitted as input to the upsampling module (Upsample) of the next first feature fusion module, and the second feature fusion network.
[0220] In a specific embodiment, the first feature fusion network includes three first feature fusion modules.
[0221] The input of the first first feature fusion module includes the output of the SPPF module of the backbone network and the intermediate feature maps output by other layers or modules of the backbone network. More specifically, for example, the input of the first first feature fusion module includes the output of the SPPF module of the backbone network and the output of the second-to-last fast convolution block.
[0222] The input of the second first feature fusion module includes the output of the first first feature fusion module and the intermediate feature maps output by other layers or modules of the backbone network. More specifically, for example, the input of the second first feature fusion module includes the output of the first first feature fusion module and the output of the first fast convolution block.
[0223] The input of the third first feature fusion module includes the output of the second first feature fusion module and the intermediate feature maps output by other layers or modules of the backbone network. More specifically, for example, the input of the third first feature fusion module includes the output of the second first feature fusion module and the output of the last shifted flip bottleneck convolution block.
[0224] The second feature fusion network of this embodiment includes at least two second feature fusion modules.
[0225] Each second feature fusion module includes a Conv module, a Concat module and an I3SS feature fusion module connected in sequence.
[0226] The input of the first second feature fusion module includes the output of the last first feature fusion module in the first feature fusion network, the outputs of the first feature fusion modules of other parts in the first feature fusion network, and the outputs of other layers or modules in the backbone network. More specifically, for example, the input of the first second feature fusion module includes the output of the last first feature fusion module in the first feature fusion network, the output of the second-to-last first feature fusion module, and the output of the first fast convolution block.
[0227] The input of each middle second feature fusion module includes the output of the previous second feature fusion module, the output of the first feature fusion modules of other parts in the first feature fusion network, and the output of some layers or modules in the backbone network.
[0228] The input of the last second feature fusion module includes the output of the previous second feature fusion module and the output of the SPPF module of the backbone network.
[0229] The input of the Conv module of the first second feature fusion module includes the output of the last first feature fusion module in the first feature fusion network.
[0230] The input of the Concat module of the first second feature fusion module includes the output of the Conv module in the same group, the output of the first feature fusion module of other parts in the first feature fusion network, and the output of other layers or modules in the backbone network.
[0231] The input of the Conv module of the middle second feature fusion module includes the output of the previous second feature fusion module.
[0232] The input of the Concat module of the middle second feature fusion module includes the output of the Conv module in the same group, the output of the first feature fusion module of other parts in the first feature fusion network, and the output of other layers or modules in the backbone network.
[0233] The input of the Conv module of the last second feature fusion module includes the output of the previous second feature fusion module.
[0234] The input of the Concat module of the last second feature fusion module includes the output of the Conv module in the same group and the output of the SPPF module of the backbone network.
[0235] The input of the I3SS feature fusion module of each second feature fusion module is the output of the Concat module in the same group.
[0236] The final output of each second feature fusion module is the output of the same group of I3SS feature fusion modules.
[0237] In a specific embodiment, the second feature fusion network includes three second feature fusion modules.
[0238] In addition, the modules in the neck network are mainly used to extract and fuse features of different resolutions of ores of different particle sizes obtained from the backbone network. The feature extraction and fusion performance of the modules affects the subsequent ore positioning accuracy and the accuracy of distinguishing ore from the background. The original YOLO model, such as the C2f module used in YOLOv8, is suitable for conventional target feature extraction, but its performance is insufficient in blasting ore detection tasks under various environmental interferences. For this reason, this embodiment introduces the I3SS module that combines the convolutional model with the Vmamba model based on the sequence model to replace the original C2f module to fuse multi-scale ore features. It can improve the detection accuracy and positioning accuracy, so that the detector has better detection performance for ore targets, and at the same time provide a good foundation for accurate tracking and statistics of ore targets in the video.
[0239] This embodiment improves the neck network by replacing the CSPLayer_2Conv module in the original YOLOv8 neck network with the I3SS feature fusion module. The I3SS feature fusion module combines a convolutional model with a sequence model, which can be a Vmamba model or a Vmamba sequence module.
[0240] Figure 5 This is a schematic diagram of the structure of the I3SS feature fusion module in one embodiment of the present application; Figure 5 ,The I3SS feature fusion module includes the I3 module, SS2D module and MLP module (i.e., I3Block, SS2D Block, MLPBlock) connected in sequence.
[0241] The input of the I3 module is the output of the Concat module in the I3SS feature fusion module. The MLP module is the final output of the same group of I3SS feature fusion modules.
[0242] The I3 module includes the Conv module, BN module, LN module and SPLIT module. Figure 5The input of the SPLIT module is the output of the Concat module in the first feature fusion module, or the output of the Concat module in the second feature fusion module. The output of the SPLIT module is given to the Conv5×5 module, the Conv3×3 module and the LN module respectively. The output of the Conv5×5 module is given to the first BN module, and the output of the Conv3×3 module is given to the second BN module. The output of the first BN module, the output of the second BN module and the output of the LN module are fused to obtain the output of the I3 module.
[0243] The output x of the I3 module is fed to the SS2D module in the same group. The output Y of the SS2D module in the same group is taken as the tensor product with the output x of the I3 module to obtain the output Y'. Y' is fed to the MLP module in the same group to obtain the output Z of the MLP module. The output Z is then taken as the tensor product with the output Y' to obtain the final output of the I3SS feature fusion module.
[0244] The MLP Block consists of a DWConv module, an LN module, and an FFN module connected in sequence. The input of the DWConv module is Y', and the output of the FFN module is Z.
[0245] The SS2D Block is a module in the VMamba network architecture.
[0246] The core idea of VMamba is to segment the input image into patches and gradually extract hierarchical features through multiple downsampling stages and the VSS module. The SS2D module can scan information in different directions to effectively extract contextual features while maintaining linear computational complexity.
[0247] VMamba is a visual backbone network based on a state-space model (SSM) with linear complexity. The core of this architecture is the Visual State Space (VSS) module and the 2D Selective Scan (SS2D) module. By traversing four scanning paths, the SS2D module (2D Selective Scan) enables the acquisition of contextual information from different directions while reducing computational cost.
[0248] VMamba, a novel SSM-based vision network, handles visual representation learning tasks with linear time complexity. It proposes 2D Selective Sweep (SS2D), extending the 1D array scan to 2D plane traversal. VMamba demonstrates excellent performance in image classification, object detection, and semantic segmentation tasks, and demonstrates advantages in input scale expansion.
[0249] The data forward process in SS2D consists of three steps: Cross-scan: SS2D first expands the incoming patches into sequences along four different traversal paths. Each patch sequence is processed in parallel using S6 blocks. Cross-merge: The resulting sequences are rearranged and merged to form the output graph.
[0250] By using complementary 1D traversal paths, SS2D allows each pixel in the image to integrate information from all other pixels in different directions. This integration is beneficial for building a global receptive field in 2D space.
[0251] More specifically, the blasted ore feature tensor extracted from the backbone network is input into the I3SS module and firstly performs convolution operation using the I3 module. The I3 module uses two sizes of convolution kernels to process the ore feature tensor respectively and splice it with the original feature. The calculation formula of the I3 module is shown in the following formula 10:
[0252] I3=Bn(conv5(I)+conv3(I)+Ln(I))
[0253] Formula 10
[0254] The convolution operation of the I3 module enables the model to further extract and fuse local features of ore images at different scales. The ore feature tensor, after convolution and concatenation by the I3 module, is input into the SS2D module for further calculations. The core of the SS2D module is the state-space model Mamba (S6) for processing sequence data. The S6 state transition formula is shown in Equations 11 and 12 below:
[0255] h i =Ah i-1 +Bx i
[0256] Formula 11
[0257] y i =Ch i +Dx i
[0258] Formula 12
[0259] x is the pixel sequence generated by the Scan Expand module, h is the current state information, A is the state transfer matrix, which represents the dynamic characteristics of the state change. B is the input control matrix, which represents how the input x affects the state. y is the output feature sequence, C is the state control matrix, which represents how the state h controls the output. D is the direct transfer matrix, which represents the direct impact of input x on output y. Since the S6 module processes sequence data, and the ore image and extracted feature tensor data have three dimensions, SS2D transforms the ore feature tensor in the H×W dimension along the following lines: Figure 5 The image is expanded in the four directions shown and input into the S6 module. This module fully learns the long-range dependencies between each ore feature pixel and other pixels, enabling more accurate positioning of the ore target frame. The I3 module, combined with the SS2D module, effectively integrates local features such as texture edges of individual ores while separating the features of different ores, facilitating the output of ore targets. Finally, the MLP module performs a nonlinear transformation on the output ore feature tensor to enhance the model's expressiveness.
[0260] By introducing a new type of combined convolutional model and Vmamba sequence model in the neck network and fusing multi-scale features across layers into the feature fusion network, the model can effectively adapt to the drastic scale changes of blasted ores, capture fine and large-sized ores, avoid ore omission, and improve the accuracy of particle size statistics.
[0261] This embodiment introduces a neck cross-path bidirectional feature fusion architecture, which introduces the feature map containing rich detailed information of fine ore in the trunk (for example, a 160×160 resolution feature map) into the neck, and at the same time adds cross-path connections from the trunk directly to the outermost path, reducing the loss of fine ore and ore edge and texture detail information. Combined with the I3SS feature fusion module, it can effectively improve the model's detection and positioning capabilities for small-sized ores.
[0262] In one embodiment, the feature extraction backbone includes a Conv convolution layer, a shifted-flip bottleneck convolution layer, a fast convolution layer, and an SPPF layer connected in sequence;
[0263] The moving flip bottleneck convolution layer includes three moving flip bottleneck convolution blocks connected in sequence, where each moving flip bottleneck convolution block includes an MBConv module;
[0264] The fast convolution layer includes 4 fast convolution blocks, each of which includes multiple stacked FasterNet network modules;
[0265] The first feature fusion network includes three first feature fusion modules, and the second feature fusion network includes three second feature fusion modules;
[0266] The input of the first feature fusion module includes the output of the SPPF layer and the output of the third fast convolution block;
[0267] The input of the second first feature fusion module includes the output of the first first feature fusion module and the output of the first fast convolution block;
[0268] The input of the third first feature fusion module includes the output of the second first feature fusion module and the output of the third moving flip bottleneck convolution block;
[0269] The input of the first second feature fusion module includes the output of the third first feature fusion module, the output of the second first feature fusion module and the output of the first fast convolution block;
[0270] The input of the second second feature fusion module includes the output of the first second feature fusion module, the output of the first first feature fusion module and the output of the third fast convolution block;
[0271] The input of the third second feature fusion module includes the output of the second second feature fusion module and the output of the SPPF layer.
[0272] Specifically, the Conv convolution layer includes a Conv module, and the input of the Conv module is visual data input; the SPPF layer includes an SPPF module.
[0273] Figure 6 This is a schematic diagram of the network architecture of the target recognition network model in one embodiment of the present application; Figure 6 The feature extraction backbone and neck network of the backbone network (Backbone) in .
[0274] The feature extraction backbone of the backbone network (Backbone) includes three moving flip bottleneck convolution blocks, where each moving flip bottleneck convolution block includes an MBConv module. More specifically, the first moving flip bottleneck convolution block is an MBConv3×3 module.
[0275] The input of the second mobile flip bottleneck convolution block is a feature map with a resolution of 320×320. The second mobile flip bottleneck convolution block contains 3 MBConv3×3 modules (Block×3).
[0276] The input of the third mobile flip bottleneck convolution block is a feature map with a resolution of 160×160. The third mobile flip bottleneck convolution block contains 3 MBConv3×3 modules (Block×3).
[0277] The input of the first fast convolution block is a feature map with a resolution of 80×80. The first fast convolution block contains 4 FasterNet network modules (FasterNetBlock), namely (Block×4).
[0278] The input of the second fast convolution block is a feature map with a resolution of 40×40. The second fast convolution block contains 4 FasterNet network modules (FasterNetBlock), i.e. (Block×4).
[0279] The input of the third fast convolution block is a feature map with a resolution of 40×40. The third fast convolution block contains 5 FasterNet network modules (FasterNetBlock), i.e. (Block×5).
[0280] The input of the fourth fast convolution block is a feature map with a resolution of 20×20. The fourth fast convolution block contains 2 FasterNet network modules (FasterNetBlock), namely (Block×2).
[0281] The input of the SPPF module is the output of the fourth fast convolution block.
[0282] The output of the SPPF layer is the output of the SPPF module.
[0283] refer to Figure 6 , the Conv module is a Conv3×3 module.
[0284] In the backbone network, the Conv module, the first moving flip bottleneck convolution block, the second moving flip bottleneck convolution block, the third moving flip bottleneck convolution block, the first fast convolution block, the second fast convolution block, the third fast convolution block, the fourth fast convolution block, and the SPPF module are connected from top to bottom.
[0285] If the dust removal network (FFA-Net) is added to the backbone network, the dust removal network (FFA-Net) is located before the Conv module.
[0286] In the neck network, each first feature fusion module includes an Upsample module, a Concat module, and an I3SS module connected in sequence. Each second feature fusion module includes a Conv module, a Concat module, and an I3SS module connected in sequence.
[0287] The input of the first feature fusion module includes the output of the SPPF layer and the output of the third fast convolution block;
[0288] The input of the second first feature fusion module includes the output of the first first feature fusion module and the output of the first fast convolution block;
[0289] The input of the third first feature fusion module includes the output of the second first feature fusion module and the output of the third moving flip bottleneck convolution block;
[0290] The input of the first second feature fusion module includes the output of the third first feature fusion module, the output of the second first feature fusion module and the output of the first fast convolution block;
[0291] The input of the second second feature fusion module includes the output of the first second feature fusion module, the output of the first first feature fusion module and the output of the third fast convolution block;
[0292] The input of the third second feature fusion module includes the output of the second second feature fusion module and the output of the SPPF layer.
[0293] This embodiment introduces a neck cross-path bidirectional feature fusion architecture, which introduces the feature map containing rich detailed information of fine ore in the trunk (for example, a 160×160 resolution feature map) into the neck, and at the same time adds cross-path connections from the trunk directly to the outermost path, reducing the loss of fine ore and ore edge and texture detail information. Combined with the I3SS feature fusion module, it can effectively improve the model's detection and positioning capabilities for small-sized ores.
[0294] In one embodiment, the target recognition network model further includes a target detection head, which is a dynamic head.
[0295] The target detection head is used to enhance the features of the fused feature map in the feature level dimension, channel dimension and ore spatial position dimension through the scale-aware attention module, the task-aware attention module and the spatial-aware attention module, and perform ore target detection and positioning based on the enhanced fused feature map to obtain the ore target detection result.
[0296] Specifically, the target recognition network model is built based on the YOLO network model, for example, the target detection algorithm based on YOLOv8, and is designed and improved to address the characteristics and difficulties of detecting blasted ore in ore piles. To address the problem of unclear edges of ore wetted for dust suppression, conventional detection and segmentation methods tend to merge with surrounding ore, thus affecting particle size calculation. The multi-dimensional feature fusion capabilities of the Dynamic Head are introduced to optimize the target detection head. This further optimizes feature fusion, making the feature tensor obtained from ore images suitable for ore detection and location tasks.
[0297] Dynamic Head is Dynamic Head.
[0298] The dynamic head applies scale, task, and spatial perception attention π to the different feature layer dimensions, channel dimensions, and ore spatial position dimensions of the blasting ore feature tensor processed by the feature fusion network of the neck network. L , π C and π S ,like Figure 6 As shown in the Dynamic Head section. This makes the edge and texture feature information of ores of various particle sizes in the feature tensor clearer, allowing the model to pay more attention to the location of the ore, and making the ore feature tensor input to the detection head more suitable for ore target positioning, distinguishing ore targets from background, and other ores. This enables the detection head to have better detection capabilities for ores of various sizes in scenes with densely packed ores and unclear edges after wetting. The dynamic head applies the attention calculation formula to the blasted ore feature tensor F as shown in Equation 13:
[0299] W(F)=π C (π S (π L (F)·F)·F)·F
[0300] Formula 13
[0301] Among them, scale-aware attention is only deployed in the feature level dimension, and by learning the relative importance of different semantic levels, it enhances the features of the appropriate level based on the scale of the object.
[0302] Spatial-aware Attention: Deployed in the spatial dimension (height × width), it first makes attention learning sparse through deformable convolution, and then aggregates cross-level features at the same spatial position to focus on discriminative regions that are consistent in spatial position and feature level.
[0303] Task-aware Attention: Deployed on the channel, it dynamically switches feature channels to support different tasks (such as classification, box regression, center / keypoint learning, etc.) based on the different convolution kernel responses of the object.
[0304] The scale-aware attention module is used to enhance the features of different feature levels of the fused feature map;
[0305] The spatial perception attention module is used to enhance cross-layer feature aggregation at the same spatial position in the fusion feature map;
[0306] Dynamically switch channels of the fused feature map through the task-aware attention module.
[0307] The feature fusion network of the neck network can be a feature fusion network constructed using a cross-path bidirectional feature fusion architecture.
[0308] The object detection head is primarily responsible for the final regression prediction, using the feature maps fused by the neck network to detect the location and category of the object. The output predictions include information such as the category of each object and its corresponding bounding box coordinates.
[0309] refer to Figure 6 The output of the target detection head is Object Classification and Box Regression, for example, predicting the position and size of the bounding box of the target object.
[0310] refer to Figure 6 In a specific embodiment, the input of the object detection head includes the outputs of all second feature fusion modules and the output of the last first feature fusion module.
[0311] The target detection head of this embodiment introduces the multi-dimensional feature fusion capability of Dynamic Head, optimizes the target detection head, further optimizes feature fusion, further improves the applicability of the feature tensor extracted from the ore image for ore detection and positioning tasks, and improves the model's ability to distinguish ore.
[0312] In a specific embodiment, the target recognition network model is constructed based on the YOLO network model, for example, based on the target detection algorithm of YOLOv8, and corresponding designs and improvements are made to the characteristics and difficulties of blasting ore detection in the ore pile. In order to solve the problem of a large amount of dust blocking the field of view that often occurs during ore transportation, a multi-level feature fusion defogging network or dust removal network FFA-Net that fuses channel attention and pixel attention is introduced, and a composite scalable ore feature extraction backbone is constructed using the MBConv module based on deep separable convolution. In order to solve the problem of dense accumulation of blasting ore and drastic changes in ore particle size, a neck feature fusion network that combines convolution with a Vmamba model based on a sequence model and adopts a BiFPN cross-path bidirectional feature fusion architecture is introduced. In order to solve the problem that the edges of ore wetted for dust suppression are unclear and conventional detection and segmentation methods are easily integrated with surrounding ore, thereby affecting the particle size calculation, the multi-dimensional feature fusion capability of Dynamic Head is introduced to optimize the target detection head and further optimize the feature fusion so that the feature tensor obtained from the ore image is suitable for the detection and positioning tasks of the ore. The network architecture of the target recognition network model is as follows: Figure 6 As shown:
[0313] The backbone network (Backbone) includes a dust removal network (FFA-Net), Conv3×3, MBConv3×3, 3 MBConv3×3 (320×320), 3 MBConv3×3 (160×160), 4 FasterNet Block (80×80), 4 FasterNet Block (40×40), 5 FasterNetBlock (40×40), 2 FasterNetBlock (20×20) and SPPF modules connected in sequence.
[0314] The neck network (Neck) includes an Upsample module, a Concat module, an I3SS feature fusion module, an Upsample module, a Concat module, an I3SS feature fusion module, an Upsample module, a Concat module, an I3SS feature fusion module, a Conv module, a Concat module, an I3SS feature fusion module, a Conv module, a Concat module, an I3SS feature fusion module, a Conv module, a Concat module and an I3SS feature fusion module, which are connected in sequence.
[0315] FFA-Net: This is a feature fusion network for extracting multi-scale features in images.
[0316] Conv3×3: 3x3 convolutional layer for further processing feature maps.
[0317] MBConv3×3: Mobile InvertedBottleneck Convolution, 3×3 mobile inverted bottleneck convolution, which is a lightweight convolution module commonly used in models on mobile devices.
[0318] Block x 3: indicates that it contains three identical modules or blocks, for example, it contains 3 MBConv3×3 or 3 FasterNet Block.
[0319] FasterNetBlock: Faster network block for accelerated computation.
[0320] SPPF:Spatial Pyramid Pooling with Fixed-size feature maps, spatial pyramid pooling, fixed-size feature maps.
[0321] Figure 7 This is a structural table of the backbone network in one embodiment of the present application; Figure 7 , the FasterNetBlock in the backbone network can also be replaced by the MBConv module. Figure 7 The network layer, input resolution, operation, dimension magnification, number of output channels and number of network layers of each module in the backbone network are illustrated.
[0322] In one embodiment, the ore particle size distribution includes n ore distribution intervals and a maximum ore particle size in each ore distribution interval;
[0323] If n=7, the energy consumption prediction formula for ore crushing is as shown in the following formula 14:
[0324] E=0.0079D1+0.0013D2-0.003D3+0.0009X c +0.0089D5-0.0069D6+0.0114D7+exp(0.114D1-0.0263D2-0.05D3-0.0359X c -0.0404D5-0.0131D6-0.04825D7)+30.5838
[0325] Formula 14
[0326] Where E is the energy consumption of ore crushing, D1 is the maximum ore particle size of the first ore distribution interval in the ore particle size distribution, D2 is the maximum ore particle size of the second ore distribution interval in the ore particle size distribution, D3 is the maximum ore particle size of the third ore distribution interval in the ore particle size distribution, and X c D5 is the maximum ore particle size in the fifth ore distribution interval in the ore particle size distribution, D6 is the maximum ore particle size in the sixth ore distribution interval in the ore particle size distribution, and D7 is the maximum ore particle size in the seventh ore distribution interval in the ore particle size distribution.
[0327] Specifically, through the KAN network, the formula between the n-dimensional ore particle size distribution statistics and energy consumption can be obtained.
[0328] The gradation statistical model is used to extract gradation data or statistics, which are, for example, low-dimensional vectors (e.g., 7-dimensional vectors). The gradation data is input into the KAN network to obtain a 7-dimensional formula as shown in Formula 14.
[0329] Formula 14 can also be obtained using DeepSet and KAN. It should be noted that there may be differences between the formulas obtained using DeepSet + KAN and using KAN alone. Furthermore, formulas for other dimensions, such as 8 and 9, can also be obtained using DeepSet + KAN and using KAN alone, and this application does not limit this.
[0330] In a specific embodiment, for example, the identified ore pile ores are sorted by ore volume, and the maximum particle size of each of the seven ore distribution intervals in the sorting results, namely 0-5% (top 5%), 0-20% (top 20%), 0-50% (top 50%), 0-63.2% (top 63.2%), 0-75% (top 75%), 0-80% (top 80%), and 0-90% (top 90%), is taken.
[0331] The energy consumption prediction formula for ore crushing is:
[0332] E=0.0079D5+0.0013D 20 -0.003D 50 +0.0009X c +0.0089D 75 -0.0069D 80 +0.0114D 90 +exp(0.114D5-0.0263D 20 -0.05D 50 -0.0359X c -0.0404D 75 -0.0131D 80 -0.04825D 90 )+30.5838
[0333] Among them, D5 is the maximum particle size in the 0-5% ore distribution range, D 20 The maximum particle size of the 0-20% ore distribution range, D 50 The maximum particle size of the 0-50% ore distribution range, X c The maximum particle size of the ore distribution range of 0-63.2%, D 75 The maximum particle size of the 0-75% ore distribution range, D 80 The maximum particle size of the 0-80% ore distribution range, D 90 The maximum particle size in the 0-90% ore distribution range.
[0334] When n=7, the ore crushing energy consumption prediction formula trained from the KAN network is a function of the 7-dimensional statistics of the ore particle size distribution. After expanding the formula, it can be found that the data of the D5 and D20 dimensions have a larger coefficient weight. Therefore, it is believed that the data of this dimension plays a vital role in the prediction of crushing energy consumption. Combined with actual blasting and crushing operations, it is found that the larger the D5 and D20 data, the greater the energy consumption, indicating that if the proportion of ore with smaller particle size is high, the crushing energy efficiency is relatively low.
[0335] In one embodiment, multi-source data fusion, in addition to using ore image data or ore video data, can also combine other types of sensor data (such as acoustics, vibration, temperature, etc.) to more comprehensively assess ore characteristics and equipment status. This helps improve the accuracy and robustness of the prediction model.
[0336] Adaptive learning mechanism, introducing online learning or reinforcement learning algorithms, enables the system to automatically update model parameters based on the latest data to adapt to changes in ore composition or the impact of equipment aging.
[0337] In one embodiment, the object detection result includes a bounding box of the identified ore;
[0338] The ore particle size is obtained by the following steps:
[0339] Estimate the equivalent area of the identified ore based on the width and height of the bounding box;
[0340] The equivalent diameter is calculated as the ore particle size based on the equivalent circle particle size estimation method according to the equivalent area.
[0341] Specifically, accurate estimation of ore particle size is a key step in ore particle size analysis. This example introduces an equivalent circular particle size estimation method, combined with camera intrinsic calibration technology, to improve the accuracy and consistency of particle size measurements. This method first uses a target recognition network model to obtain a bounding box for the ore particles. The area of the ore is then estimated by calculating the width and height of the bounding box. Pixel area is converted to physical dimensions through actual conversion, and finally, the equivalent circular particle size estimation method is used to derive the ore's equivalent diameter.
[0342] Figure 8 Schematic diagram of the bounding box of the ore identified in one embodiment of the present application; Figure 9 This is a schematic diagram of a camera calibration plate image in one embodiment of the present application; Figure 10 This is a schematic diagram of the maximum inscribed ellipse of the identification frame and a schematic diagram of the equivalent sphere diameter in one embodiment of the present application.
[0343] (1) Camera intrinsic calibration
[0344] In order to ensure the accuracy of particle size estimation, this embodiment first performs a strict internal calibration of the camera. This process is based on Zhang's calibration method. It can accurately determine the focal length f, principal point coordinates (u0, v0) and pixel size d of the camera. x d y , thus constructing the camera's intrinsic parameter matrix K as shown in the following formula 15:
[0345]
[0346] Among them, f x and fy They represent the horizontal and vertical focal lengths respectively, in pixels; u0 and v0 are the pixel coordinates of the camera's principal point. Figure 9 An image of a calibration plate used for camera intrinsic calibration is shown. The known geometric features on the calibration plate are used to calculate the camera's intrinsic parameter matrix, which is then applied to pixel size correction in subsequent images to improve the accuracy of ore particle size measurement.
[0347] (2) Equivalent circular particle size estimation method
[0348] In the particle size estimation method, the target recognition network model is used to detect the ore particles. First, the detection frame of each ore particle is obtained through the target recognition network model. Each detection frame represents the target area of the ore particle in the image. Based on the width W and height H of the detection frame, the maximum inscribed ellipse of the identification frame is used as the equivalent area S of the ore, and then the equivalent circle method is used for estimation. The core of the equivalent circle method is to regard the two-dimensional plane image of the ore particle as a projection of an approximate circle, and calculate the equivalent diameter based on the projection area. Specifically, the equivalent area S eq , equivalent diameter D eq and equivalent volume V eq As shown in the following formula 16-18:
[0349]
[0350] The present application also provides a method for predicting ore crushing energy consumption, which includes:
[0351] Obtaining an ore crushing energy consumption prediction formula according to any of the above methods for obtaining an ore crushing energy consumption prediction formula;
[0352] The obtained ore particle size distribution of the target ore pile is substituted into the ore crushing energy consumption prediction formula to obtain the ore crushing predicted energy consumption of the target ore pile.
[0353] Specifically, the method for obtaining the ore crushing energy consumption prediction formula of this embodiment is described above and will not be repeated here.
[0354] refer to Figure 11 In one embodiment of the present application, a device for obtaining an ore crushing energy consumption prediction formula is further provided. The device for obtaining an ore crushing energy consumption prediction formula includes:
[0355] A data set construction module 100 is used to construct a data set based on the ore particle size distribution and corresponding ore crushing energy consumption of multiple groups of ore piles;
[0356] The formula prediction module 200 is used to input the data set into the DeepSet network to obtain a first eigenvector about the ore particle size distribution and a second eigenvector about the corresponding ore crushing energy consumption, and input multiple groups of first eigenvectors and corresponding second eigenvectors into the KAN network to obtain an ore crushing energy consumption prediction formula; or, directly input the data set into the KAN network to obtain an ore crushing energy consumption prediction formula; the ore crushing energy consumption prediction formula is used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile.
[0357] In one embodiment, the device for obtaining the ore crushing energy consumption prediction formula further includes:
[0358] The target detection module is used to detect ore targets in the ore pile using the ore visual data as input through the target recognition network model, and obtain the ore target detection results. The target recognition network model is built based on the YOLO model;
[0359] The particle size distribution acquisition module is used to track each ore in the ore target detection result and calculate the ore particle size to obtain the ore particle size distribution of the ore pile;
[0360] The energy consumption acquisition module is used to obtain the ore crushing energy consumption consumed by the ore pile to complete ore crushing.
[0361] In one embodiment, the target recognition network model includes a backbone network,
[0362] The backbone network includes a feature extraction backbone, or the backbone network includes a dust removal network layer and a feature extraction backbone connected in sequence;
[0363] The dust removal network layer is used to perform dust removal processing on the ore visual data to obtain dust-free visual data;
[0364] The feature extraction backbone is used to extract features from ore visual data or dust removal visual data to obtain feature maps of different scales.
[0365] In one embodiment, the feature extraction backbone includes a Conv convolution layer, a shifted-flip bottleneck convolution layer, a fast convolution layer, and an SPPF layer connected in sequence;
[0366] The Conv convolution layer includes a Conv module, and the input of the Conv module is visual data input, wherein the visual data input is ore visual data or dust removal visual data;
[0367] The moving flip bottleneck convolution layer includes at least two moving flip bottleneck convolution blocks connected in sequence, wherein each moving flip bottleneck convolution block includes an MBConv module;
[0368] The fast convolution layer includes at least two fast convolution blocks, wherein each fast convolution block includes multiple stacked FasterNet network modules;
[0369] The SPPF layer includes the SPPF module.
[0370] In one embodiment, the target recognition network model includes a neck network,
[0371] The neck network includes the first feature fusion network and the second feature fusion network.
[0372] The first feature fusion network includes at least two first feature fusion modules;
[0373] The second feature fusion network includes at least two second feature fusion modules.
[0374] Each first feature fusion module includes an upsampling module, a Concat module and an I3SS feature fusion module connected in sequence;
[0375] Each second feature fusion module includes a Conv module, a Concat module and an I3SS feature fusion module connected in sequence;
[0376] The I3SS feature fusion module is constructed by combining the convolutional model and the sequence model. The I3SS feature fusion module includes the I3 module, SS2D module and MLP module connected in sequence;
[0377] Among them, the I3 module includes the SPLIT module, Conv module, LN module and BN module.
[0378] In one embodiment, the feature extraction backbone includes a Conv convolution layer, a shifted-flip bottleneck convolution layer, a fast convolution layer, and an SPPF layer connected in sequence;
[0379] The moving flip bottleneck convolution layer includes three moving flip bottleneck convolution blocks connected in sequence, where each moving flip bottleneck convolution block includes an MBConv module;
[0380] The fast convolution layer includes 4 fast convolution blocks, each of which includes multiple stacked FasterNet network modules;
[0381] The first feature fusion network includes three first feature fusion modules, and the second feature fusion network includes three second feature fusion modules;
[0382] The input of the first feature fusion module includes the output of the SPPF layer and the output of the third fast convolution block;
[0383] The input of the second first feature fusion module includes the output of the first first feature fusion module and the output of the first fast convolution block;
[0384] The input of the third first feature fusion module includes the output of the second first feature fusion module and the output of the third moving flip bottleneck convolution block;
[0385] The input of the first second feature fusion module includes the output of the third first feature fusion module, the output of the second first feature fusion module and the output of the first fast convolution block;
[0386] The input of the second second feature fusion module includes the output of the first second feature fusion module, the output of the first first feature fusion module and the output of the third fast convolution block;
[0387] The input of the third second feature fusion module includes the output of the second second feature fusion module and the output of the SPPF layer.
[0388] In one embodiment, the target recognition network model further includes a target detection head, which is a dynamic head.
[0389] The target detection head is used to enhance the features of the fused feature map in the feature level dimension, channel dimension and ore spatial position dimension through the scale-aware attention module, the task-aware attention module and the spatial-aware attention module, and perform ore target detection and positioning based on the enhanced fused feature map to obtain the ore target detection result.
[0390] In one embodiment, the ore particle size distribution includes n ore distribution intervals and a maximum ore particle size in each ore distribution interval;
[0391] If n=7, the energy consumption prediction formula for ore crushing is:
[0392] E=0.0079D1+0.0013D2-0.003D3+0.0009X c +0.0089D5-0.0069D6+0.0114D7+exp(0.114D1-0.0263D2-0.05D3-0.0359X c -0.0404D5-0.0131D6-0.04825D7)+30.5838
[0393] Where E is the energy consumption of ore crushing, D1 is the maximum ore particle size of the first ore distribution interval in the ore particle size distribution, D2 is the maximum ore particle size of the second ore distribution interval in the ore particle size distribution, D3 is the maximum ore particle size of the third ore distribution interval in the ore particle size distribution, and X cD5 is the maximum ore particle size in the fifth ore distribution interval in the ore particle size distribution, D6 is the maximum ore particle size in the sixth ore distribution interval in the ore particle size distribution, and D7 is the maximum ore particle size in the seventh ore distribution interval in the ore particle size distribution.
[0394] In one embodiment, the device for obtaining the ore crushing energy consumption prediction formula further includes:
[0395] The visual data acquisition module is used to obtain the original visual data of the ore before the ore is crushed. The original visual data of the ore is the video before the ore is crushed;
[0396] The electric meter data acquisition module is used to obtain the electric meter readings before and after the ore is crushed, as well as the time when the electric meter is read;
[0397] The synchronization alignment module is used to perform time synchronization alignment and cropping on the video before ore crushing according to the meter reading time to obtain ore visual data;
[0398] The energy consumption acquisition module is specifically used to obtain the ore crushing energy consumption consumed to complete the ore crushing according to the electric meter readings before and after the ore crushing.
[0399] In one embodiment, the present application further provides a device for predicting ore crushing energy consumption, the device comprising: a device for obtaining any one of the above-mentioned ore crushing energy consumption prediction formulas, for obtaining the ore crushing energy consumption prediction formula;
[0400] The calculation module is used to substitute the obtained ore particle size distribution of the target ore pile into the ore crushing energy consumption prediction formula to obtain the ore crushing predicted energy consumption of the target ore pile.
[0401] This application proposes a solution for predicting ore crushing energy consumption, based on a dust removal recognition network. It innovatively combines advanced image processing techniques and machine learning algorithms to efficiently predict energy consumption during the crushing process and determine the particle size that minimizes energy consumption. By incorporating a specially designed dust removal recognition network, this solution enables more accurate analysis and evaluation of ore images. This technological advancement provides an effective solution to the low recognition rate of traditional image recognition methods in complex dusty environments, significantly improving the accuracy and reliability of ore image analysis.
[0402] The prediction model uses a combination of DeepSet+Kan network to convert the ore particle size distribution data output from the recognition model into a vector. After training, a clear formula for energy consumption and ore particle size distribution is obtained.
[0403] For ore target tracking, the BoTSORT algorithm (Bytetrack with Occlusion and Scale-aware Re-identification) is used to track the trajectory of ore particles in real time. BoTSORT effectively solves tracking interruptions caused by occlusion or target motion.
[0404] Figure 12 7 is a structural diagram of a computer device provided in an embodiment of the present application. The computer device 700 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 (for example, one or more mass storage devices) storing application programs 733 or data 732. Among them, the memory 720 and the storage medium 730 can be temporary storage or permanent storage. The program stored in the storage medium 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 700. Furthermore, the processor 710 can be configured to communicate with the storage medium 730 to execute a series of instruction operations in the storage medium 730 on the computer device 700.
[0405] The computer device 700 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input and output interfaces 760, and / or one or more operating systems 731, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 12 The illustrated computer device structure does not limit the computer device and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0406] The present application also provides a computer device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the method for obtaining the ore crushing energy consumption prediction formula in the above-mentioned embodiments.
[0407] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to execute the steps of a method for obtaining a formula for predicting energy consumption for ore crushing.
[0408] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0409] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0410] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for obtaining a prediction formula for ore crushing energy consumption, characterized in that: The method for obtaining the ore crushing energy consumption prediction formula includes: Constructing a dataset based on the ore particle size distribution and corresponding ore crushing energy consumption of multiple groups of ore piles; The data set is input into the DeepSet network to obtain a first eigenvector about the ore particle size distribution and a second eigenvector about the corresponding ore crushing energy consumption, and multiple groups of the first eigenvectors and the corresponding second eigenvectors are input into the KAN network to obtain an ore crushing energy consumption prediction formula; or, the data set is directly input into the KAN network to obtain an ore crushing energy consumption prediction formula; the ore crushing energy consumption prediction formula is used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile.
2. The method for obtaining the ore crushing energy consumption prediction formula according to claim 1, characterized in that: Before constructing a data set based on the ore particle size distributions and corresponding ore crushing energy consumptions of multiple groups of ore piles, the method further includes: Using the visual data of the ore pile as input, performing ore target detection on the ore pile through a target recognition network model to obtain an ore target detection result, wherein the target recognition network model is constructed based on the YOLO model; Tracking each ore in the ore target detection result and calculating the ore particle size to obtain the ore particle size distribution of the ore pile; Obtain the ore crushing energy consumption consumed by the ore pile to complete ore crushing.
3. The method for obtaining the ore crushing energy consumption prediction formula according to claim 2, characterized in that: The target recognition network model includes a backbone network, The backbone network includes a feature extraction backbone, or the backbone network includes a dust removal network layer and a feature extraction backbone connected in sequence; The dust removal network layer is used to perform dust removal processing on the ore visual data to obtain dust-removed visual data; The feature extraction backbone is used to perform feature extraction on the ore visual data or the dust removal visual data to obtain feature maps of different scales.
4. The method for obtaining the ore crushing energy consumption prediction formula according to claim 3, characterized in that: The feature extraction backbone includes a Conv convolution layer, a moving flip bottleneck convolution layer, a fast convolution layer and an SPPF layer connected in sequence; The Conv convolution layer includes a Conv module, and the input of the Conv module is visual data input, wherein the visual data input is ore visual data or dust removal visual data; The moving flip bottleneck convolution layer includes at least two moving flip bottleneck convolution blocks connected in sequence, and the fast convolution layer includes at least two fast convolution blocks; or, the sum of the number of convolution blocks of the moving flip bottleneck convolution block included in the moving flip bottleneck convolution layer and the fast convolution block included in the fast convolution layer is not less than 4; wherein each moving flip bottleneck convolution block includes an MBConv module; and each fast convolution block includes multiple stacked FasterNet network modules; The SPPF layer includes an SPPF module.
5. The method for obtaining the ore crushing energy consumption prediction formula according to claim 3 or 4, characterized in that: The target recognition network model includes a neck network, The neck network includes a first feature fusion network and a second feature fusion network, The first feature fusion network includes at least two first feature fusion modules; The second feature fusion network includes at least two second feature fusion modules, Each first feature fusion module includes an upsampling module, a Concat module and an I3SS feature fusion module connected in sequence; Each second feature fusion module includes a Conv module, a Concat module and an I3SS feature fusion module connected in sequence; The I3SS feature fusion module is constructed by combining the convolution model and the sequence model. The I3SS feature fusion module includes an I3 module, an SS2D module and an MLP module connected in sequence; The I3 module includes a SPLIT module, a Conv module, an LN module and a BN module.
6. The method for obtaining the ore crushing energy consumption prediction formula according to claim 5, characterized in that: The feature extraction backbone includes a Conv convolution layer, a moving flip bottleneck convolution layer, a fast convolution layer and an SPPF layer connected in sequence; The moving flip bottleneck convolution layer includes three moving flip bottleneck convolution blocks connected in sequence, wherein each moving flip bottleneck convolution block includes an MBConv module; The fast convolution layer includes 4 fast convolution blocks, wherein each fast convolution block includes multiple stacked FasterNet network modules; The first feature fusion network includes three first feature fusion modules, and the second feature fusion network includes three second feature fusion modules; The input of the first feature fusion module includes the output of the SPPF layer and the output of the third fast convolution block; The input of the second first feature fusion module includes the output of the first first feature fusion module and the output of the first fast convolution block; The input of the third first feature fusion module includes the output of the second first feature fusion module and the output of the third moving flip bottleneck convolution block; The input of the first second feature fusion module includes the output of the third first feature fusion module, the output of the second first feature fusion module and the output of the first fast convolution block; The input of the second second feature fusion module includes the output of the first second feature fusion module, the output of the first first feature fusion module and the output of the third fast convolution block; The input of the third second feature fusion module includes the output of the second second feature fusion module and the output of the SPPF layer.
7. The method for obtaining the ore crushing energy consumption prediction formula according to claim 5, characterized in that: The target recognition network model further includes a target detection head, which is a dynamic head. The target detection head is used to perform feature enhancement on the fused feature map in the feature level dimension, channel dimension and ore spatial position dimension through a scale-aware attention module, a task-aware attention module and a space-aware attention module, and perform ore target detection and positioning based on the enhanced fused feature map to obtain an ore target detection result.
8. The method for obtaining the ore crushing energy consumption prediction formula according to any one of claims 1-4, 6-7, characterized in that: The ore particle size distribution includes n ore distribution intervals and the maximum ore particle size of each ore distribution interval; If n=7, the ore crushing energy consumption prediction formula is: E=0.0079D1+0.0013D2-0.003D3+0.0009X c +0.0089D5-0.0069D6+0.0114D7+exp(0.114D1-0.0263D2-0.05D3-0.0359X c -0.0404D5-0.0131D6-0.04825D7)+30.5838 Where E is the energy consumption of ore crushing, D1 is the maximum ore particle size of the first ore distribution interval in the ore particle size distribution, D2 is the maximum ore particle size of the second ore distribution interval in the ore particle size distribution, D3 is the maximum ore particle size of the third ore distribution interval in the ore particle size distribution, and X c D5 is the maximum ore particle size in the fifth ore distribution interval in the ore particle size distribution, D6 is the maximum ore particle size in the sixth ore distribution interval in the ore particle size distribution, and D7 is the maximum ore particle size in the seventh ore distribution interval in the ore particle size distribution.
9. The method for obtaining the ore crushing energy consumption prediction formula according to claim 2, characterized in that: Before performing ore target detection on the ore pile by using the target recognition network model to obtain the ore target detection result, the method further includes: Obtaining original visual data of the ore before ore crushing in the ore pile, wherein the original visual data of the ore is a video before ore crushing; Obtain the meter readings before and after the ore is crushed, as well as the meter reading times; Performing time synchronization alignment on the video before ore crushing according to the meter reading time and then cropping it to obtain ore visual data; The obtaining of the ore crushing energy consumption consumed by the ore pile to complete ore crushing includes: The ore crushing energy consumption consumed to complete the ore crushing is obtained based on the electric meter readings before and after the ore crushing.
10. The method for obtaining the ore crushing energy consumption prediction formula according to claim 2, characterized in that: The ore target detection result includes a bounding box of the identified ore; The ore particle size is obtained by the following steps: estimating the equivalent area of the identified ore based on the width and height of the bounding box; The equivalent diameter is calculated as the ore particle size based on the equivalent area using the equivalent circle particle size estimation method.
11. A method for predicting energy consumption of ore crushing, characterized in that: The method for predicting the energy consumption of ore crushing includes: Obtaining an ore crushing energy consumption prediction formula according to the method for obtaining an ore crushing energy consumption prediction formula according to any one of claims 1 to 10; The obtained ore particle size distribution of the target ore pile is substituted into the ore crushing energy consumption prediction formula to obtain the ore crushing predicted energy consumption of the target ore pile.
12. A device for obtaining a prediction formula for ore crushing energy consumption, characterized in that: The device for obtaining the ore crushing energy consumption prediction formula includes: A data set construction module, for constructing a data set based on the ore particle size distribution and corresponding ore crushing energy consumption of multiple groups of ore piles; A formula prediction module is used to input the data set into the DeepSet network to obtain a first eigenvector about the ore particle size distribution and a second eigenvector about the corresponding ore crushing energy consumption, and input multiple groups of the first eigenvectors and the corresponding second eigenvectors into the KAN network to obtain an ore crushing energy consumption prediction formula; or, directly input the data set into the KAN network to obtain an ore crushing energy consumption prediction formula; the ore crushing energy consumption prediction formula is used to predict ore crushing energy consumption based on the ore particle size distribution of the ore pile.
13. A computer device, characterized in that: The computer device includes: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the computer device executes the method for obtaining the ore crushing energy consumption prediction formula as described in any one of claims 1 to 10, or so that the computer device executes the method for predicting ore crushing energy consumption as described in claim 11.
14. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instruction is executed by the processor, the method for obtaining the ore crushing energy consumption prediction formula as described in any one of claims 1 to 10 is implemented, or when the instruction is executed by the processor, the method for predicting the ore crushing energy consumption as described in claim 11 is implemented.