Tailing ash value prediction system and method based on bionic neural network
By constructing a bionic neural network with a multi-layer relational attention mechanism and simulating the adjustment of brain neuron connections, the accuracy problem of traditional neural networks in predicting the water-ash content of tailings coal slime is solved, and high-precision and stable ash content prediction is achieved, which is suitable for on-site application in coal preparation plants.
Patent Information
- Application Number
- CN202511120650.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-12
AI Technical Summary
The existing machine learning linear regression method cannot accurately predict the ash content of tailings slime water. The traditional neural network training method is single and fixed and cannot capture the nonlinear relationship of the ash content of tailings slime water.
A multi-layer relational attention mechanism based on meta-learning theory is adopted to construct a bionic neural network through multi-dimensional feature extraction and dynamic weighting, simulate the connection strength and relationship adjustment of brain neurons, establish a multi-layer attention mechanism and a dynamic learning rate adjustment mechanism, and achieve accurate prediction of gray value.
It achieves accurate prediction of the water-ash content of tailings slime, improves prediction accuracy and training stability, adapts to various scenarios, is safe and easy to operate, and is suitable for on-site deployment in coal preparation plants.
Smart Images

Figure CN120635677A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a tailings ash value prediction system and method based on a bionic neural network, belonging to the technical field of coal mine ash value prediction. Background Art
[0002] After the clean coal is recovered from the floating coal slurry in the flotation machine, the remaining part is tailings. Ash is the incombustible product of the coal slurry, and the residue consists of gangue (useless minerals in the ore) and mineral debris. It usually contains a small amount of valuable components that are not fully recovered, such as fine-grained coal and coal-containing components mixed with gangue. Flotation reagents can be used to separate these valuable components for recovery. If the ash value can be predicted in advance and the raw materials of the flotation reagent can be accurately proportioned, the recovery rate will be higher. Therefore, it is particularly important to predict the ash value of the tailings in advance.
[0003] In recent years, with the deepening of research, camera photography has been used to capture the characteristics of coal slime water's light reflection and sampled images, such as energy and grayscale values. By analyzing the relationship between these characteristics, conditions, and ash content, and based on a large amount of data, using the characteristics as conditions and the actual measured ash content as the target value, a machine learning linear regression method is used to determine the target ash content based on the characteristic conditions. However, the machine learning linear regression method can only capture the linear relationship between the conditions and the target, and the neural network training method is single and fixed, making it unable to accurately predict the ash content of tailings coal slime water. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a tailings ash value prediction system and method based on a bionic neural network. In view of the shortcomings of traditional training methods, the meta-learning theory is used to construct a multi-layer relational attention mechanism, breaking through the traditional single and fixed neural network training method, and realizing accurate prediction of the water ash value of tailings coal slime.
[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions: In a first aspect, the present invention provides a method for predicting tailings ash content based on a bionic neural network, comprising: Acquire tailings slime water collection images; Perform multi-dimensional feature extraction on tailings slime water collection images; Dynamically weighted based on the extracted multi-dimensional features through a multi-layer relational attention mechanism; The dynamically weighted features are input into the pre-trained neural network to predict the gray value; Output the ash value prediction result.
[0006] Furthermore, the multidimensional features include grayscale features, color features, texture features and edge features, wherein: The grayscale feature includes grayscale mean; The color features include R / G / B channel mean, R / G / B channel standard deviation, and R / G / B channel variance; The edge feature is a Sobel edge feature, including edge mean, variance, and energy; The texture features are GLCM texture features, including contrast, homogeneity, correlation, energy, and entropy.
[0007] Furthermore, the multi-layer relational attention mechanism includes: a feature-target relational attention module, a feature-feature relational attention module, a feature-feature relational attention module, and a cross-modal multi-interaction relational attention module, wherein: The feature-target relationship attention module is: Where: is the feature-target attention weight, is the Sigmoid function, is the first linear layer, is the second linear layer, is the activation function, is the input feature vector.
[0008] Furthermore, the feature-feature relationship attention module is: Where: is the relationship weight, is the Sigmoid function, is the input feature vector, is the projection matrix, is the relationship matrix, is the number of heads of multi-head attention.
[0009] Furthermore, the mathematical model of the feature-feature relationship attention module is: Where: is the new relationship matrix after the interaction between heads, is the normalized exponential function, is the multi-head relationship weight matrix, is element-wise multiplication, is a learnable bilinear matrix, is the Sigmoid function, is the gate weight, is the gating weight matrix used to Mapping to gating weights Dimensions, is the mean value function, is the final result of the relational attention module between feature-feature relations.
[0010] Furthermore, the mathematical model of the cross-modal multi-interaction attention module is: Where: is the projection vector of the feature-target relationship, is the mapping function that projects the feature-target attention weights into the shared interaction space, is the feature-target attention weight, is the matrix used to linearly transform the feature-target attention weights, is the projection vector of relation-relation weight, is the multi-head relationship weight matrix, is the mapping function that projects the relationship-relationship weights into the shared interaction space, is the matrix used to linearly transform the multi-head relationship weight matrix, is the adjustment factor calculated through cross-modal interaction, is the cross-modal interaction relationship weight of the final output, is a linear rectification function, Interactive core.
[0011] Furthermore, the neural network includes a neural network model, a similarity calculation model, a weight adjustment model and a learning rate adjustment model, wherein: The neural network model is: Where: is the predicted output value of the neural network for the target ash value, is the activation function is the input feature vector, is the final layer; The similarity calculation model is: Where: For the current sample The comprehensive similarity score with historical samples, is the current sample vector feature, is the vector feature of the historical sample, is the feature-target relationship weight vector of the current sample, For historical samples The feature-target relationship weight vector, is the feature-feature relationship weight matrix of the current sample, For historical samples The feature-feature relationship weight matrix, is the cosine similarity function, For historical memory The weighting coefficient of For current samples and historical memory The comprehensive similarity of For historical memory Confidence in predictions; is the total number of samples stored in the historical memory library, m is the historical sample sum index, is the comprehensive similarity between the current sample and the historical sample m, is the prediction confidence of historical sample m; By judging the difference between the predicted value and the verification target value, the weight of the sampled image feature is adjusted. The weight adjustment model is: Where: is the set of all learnable parameters of the neural network, To express the parameter Minimize the value of the subsequent expression for the optimization goal, is the loss function, is the function mapping, For the The feature vector of the training samples, represents the prediction function of the neural network model, For the The true gray value of the training samples, is the regularization coefficient; is the calculation function of the attention mechanism, is the weight matrix of the attention mechanism, is the query vector, is the key vector, is a value vector, is the normalized exponential function, is the transpose of the key vector matrix, is the dimension of the key vector, is the weight element; The learning rate adjustment model is: Where: The step length is The learning rate, is the initial learning rate, is the target association influence coefficient, is the average attention weight of feature importance, is the influence coefficient of the relationship between conditions, is the variance of the feature relationship.
[0012] In a second aspect, the present invention provides a tailings ash value prediction system based on a bionic neural network, comprising: Image acquisition module: used to acquire tailings slime water collection images; Feature extraction module: used to extract multi-dimensional features from tailings slime water images; Dynamic weighting module: used to dynamically weight the extracted multi-dimensional features through a multi-layer relational attention mechanism; Gray value prediction module: used to input the dynamically weighted features into the pre-trained neural network to predict the gray value; Ash value output module: used to output ash value prediction results.
[0013] In a third aspect, the present invention provides a device for predicting tailings ash content based on a bionic neural network, comprising a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of any of the above methods.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above methods when executed by a processor.
[0015] Compared with the prior art, the present invention has the following beneficial effects: First, the present invention addresses the shortcomings of traditional training methods and uses meta-learning theory to construct a multi-layer relational attention mechanism. First, multi-dimensional features are extracted from tailings slime water images. Based on the extracted multi-dimensional features, the multi-layer relational attention mechanism is used to dynamically weight them. The dynamically weighted features are then input into a pre-trained neural network for ash content prediction. This method breaks through the traditional single, fixed neural network training method and achieves accurate prediction of tailings slime water ash content. 2. The present invention constructs a multi-connection attention mechanism based on the basic neuronal connection relationship, and uses error feedback and synaptic plasticity functions between neurons to adjust the connection strength between them through the feedback of electrical signals. The biological principle of the connection method is used to establish a regulatory relationship backpropagation based on the neural network loss function and a dynamic adjustment mechanism of neuron weights; guided by attention, a multi-layer connection is established between neurons, groups, and domains to strengthen the information transmission path and speed to establish a feature-target, feature-feature and other multi-layer attention relationship mechanism; guided by the brain adjusting the amount of neurotransmitter release and switching neurons according to the degree of task responsibility, a dynamic learning rate adjustment mechanism is established to improve the learning efficiency of the neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings: Figure 1 A schematic flow chart of a method for predicting tailings ash content based on a bionic neural network provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0017] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other.
[0018] The following detailed description is an exemplary description and is intended to provide further detailed description of the present invention. Unless otherwise indicated, all technical terms used in the present invention have the same meaning as those generally understood by those skilled in the art to which the present invention belongs. The terms used in the present invention are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.
[0019] Example 1: This embodiment proposes a method for predicting the ash content of tailings based on a bionic neural network. To address the shortcomings of traditional training methods, the method utilizes meta-learning theory, imitates the learning process of primate brains learning instructions and the way of observing things, and takes the "minimum consciousness system" as the starting point to simulate the layered perception mechanism of the visual system. It imitates the connection strength, mode, and relationship of brain neurons for dynamic adjustment. That is, no pre-programming is required during the training process. Goal-oriented behavior can be achieved only by feedback-adjusting the connection strength, mode, and selection of neurons, i.e., neural network training nodes. This breaks through the traditional single and fixed training method. This solution is based on MAE difference feedback, multi-layer relationships and attention mechanisms, and heuristic reasoning to establish predicted ash values. By adding feature-feature relationships, relationship-relationships (i.e., feature-feature relationship relationships), relationship-relationships, and feature-target relationship cross-modal multiple interactive relationships within neural network elements on the basis of feature-target relationships, and adding attention mechanisms to the above relationships; using the gap between the predicted results and the true ash value, feedback is adjusted to the neural network and the parameters of the above multiple relationships and neural networks are adjusted, multiple iterations are performed, and the training results and parameters of each iteration are stored in the neural network. At the next prediction, the corresponding empirical relationship is called out according to the predicted body features to participate in the prediction to increase the prediction accuracy. The present invention is based on establishing multiple neural network parameter relationships and combining multi-layer relationship attention mechanisms to adjust the various parameters of the neural network, and can memorize each training result and its characteristics, thereby achieving accurate prediction of the tailings coal slime water ash value. Specifically including the following contents: 1. Bionic neural network gray value prediction method based on MAE feedback, multi-layer relational attention mechanism and heuristic reasoning Establish the relationship between multiple parameters of the neural network and use the multi-head attention mechanism. After each training, use the MAE prediction method to evaluate the prediction results, feed back to the neural network to adjust the above relationship and related parameters, and store each training result for quick reference during the next prediction. After multiple iterations, the best result is saved. The specific operation scheme includes the following: 1. Preprocess and extract parameters of the camera-sampled images: Use OpenCV to read the images and convert them into a processable format. The texture feature (GLCM) and edge feature (Sobel) operators are both image features. They are converted into tensors usable in the PyTorch architecture using Tensor, allowing the GPU to participate in the calculation and reducing CPU overhead.
[0020] 2. Use the model to predict the gray value based on the sampled image and the extracted parameters, and then use the neural network to train it. After obtaining the training feature results, load them into the prediction end to predict the gray value.
[0021] 2. Implementation of the association between the bionic neural network and the minimal consciousness system based on 18-dimensional features 1. Bio-perceptual mapping with 18-dimensional feature processing Given the relationship between grayscale value, grayscale mean, and other light characteristics, 18-dimensional features are extracted. This covers all visual dimensions related to grayscale value prediction from a physical perspective: 1) Basic visual layer (grayscale / color features)—10 dimensions in total: grayscale mean, R, G, and B channel means, R, G, and B channel standard deviations, and R, G, and B channel variances. Grayscale mean directly reflects the overall grayscale value, which is proportional to brightness. RGB channel statistics identify mineral type trends and reflective interference.
[0022] 2) Intermediate feature layer (texture / edge features): Sobel edge feature (3D, edge mean, variance, energy): It captures edge information by calculating image gradients. Its statistical characteristics directly reflect the physical properties of particles in coal slime water, quantify particle sharpness and contour clarity, and the ash value is proportional to the edge strength and variance. The extreme ash value is positively correlated with the edge energy.
[0023] GLCM (Gray Level Co-occurrence Matrix) texture features (5 dimensions, contrast, homogeneity, correlation, energy, entropy): Texture is quantified by statistically analyzing the spatial relationship of pixel pairs. Its core parameters directly reflect the physical state of coal slime water and analyze the particle distribution state. Contrast and correlation are proportional to the ash value, homogeneity is inversely proportional to the ash value, and ash extreme values are proportional to energy.
[0024] 3. Four Enhanced Connection Computing Modules of Neural Nodes 1. Feature-target relationship attention module k-target (k-target refers to the correlation between features (key) and targets (target), "k" is the abbreviation of "key", which is directly derived from the Query-Key-Value (QKV) model, the core of the attention mechanism, where k refers to Key (feature) and Target refers to target) neurons selectively focus on key stimuli.
[0025] Mathematical model: (Synaptic weight generation) In the formula, the input feature vector First pass through the first linear layer Compress the dimension and then pass The activation function introduces nonlinearity and then passes through a second linear layer Restore the original dimension and then use the Sigmoid function Generate attention weights , and apply the attention feature to the original feature, and attention weights Multiplying by levels , and finally apply the attention weights to the original features.
[0026] The FeatureAttention module uses linear transformation and activation function to generate attention weights on the feature dimension through two layers of linear transformation to calculate the relationship between 18 features (such as RGB channel mean and standard deviation) and the target gray value, and dynamically adjust the weight of each feature. For example: 1) The red channel has a high mean value: This may correspond to a specific area in the image (such as impurities in coal). The model will enhance the influence of this feature on the ash value prediction through attention weighting.
[0027] 2) Large color variance: This may indicate complex image texture, and the model will adjust the importance of this feature based on historical feedback.
[0028] 2. Feature-feature relationship attention (kk, the internal correlation between feature key and other feature keys) Neuron cluster interaction simulates the collaborative processing of neuron clusters in the cerebral cortex, such as the association weights between visual processing neurons and motor neurons.
[0029] Mathematical model: (Multi-cluster parallel processing) In the formula, input By different projection matrices Mapping to multiple subspaces, learning different relationship representations, each projected feature and the learnable relationship matrix Multiply and then pass through the Sigmoid function Generate relationship weights , and finally the weighted features of each head are averaged and merged, allowing the model to capture feature relationships from different subspaces. Represents the number of multi-head attention heads. Multi-head attention is a core mechanism of the Transformer. This attention calculation projects the features of the input head (i.e., the prediction body) into multiple subspaces. This allows each input feature, such as RGB mean, standard deviation, and contrast, to learn different attention patterns in the mapping network. Attention is then calculated in independent subspaces, and the multiple results are finally concatenated. This effectively captures complex feature combinations. The ConditionalRelationAttention module uses a multi-head mechanism to model complex relationships between features. For example: Dependencies between texture features: Regions with high contrast may also have low homogeneity. These two features together reflect the particle distribution characteristics of coal. The multi-head attention mechanism can learn these multi-dimensional dependencies in parallel.
[0030] 3. Relation-relation (i.e., feature-feature relationship) attention kk-kk, where kk represents the internal correlation between a feature (key) and other features (keys), and kk-kk represents the relationship between these correlations. Interaction and coordination of high-level neural circuits.
[0031] Mathematical model: Where: is the new relationship matrix (intermediate variable) after the interaction between heads, is the normalized exponential function, is the multi-head relationship weight matrix, It is element-by-element multiplication, which is used to multiply the weights and matrices by position. is a learnable bilinear matrix (used to model the interaction between different attention heads), is the Sigmoid function, is the gate weight, is the gating weight matrix used to Mapping to gating weights Dimensions, is the mean value function, It is an enhanced relationship matrix that integrates multi-level feature relationship information and is the final result of the feature-feature relationship attention module.
[0032] Input multi-head relationship weight matrix and bilinear matrix Multiply and calculate the interaction between different attention heads to get a new relationship matrix and through Normalization. By gating weights Weighted fusion of multi-head relationship weight matrix and inter-head interaction relationship , to achieve dynamic adjustment. The RelationRelationAttention module handles the relationship between different attention heads (representing different feature clusters). For example: Interaction between color feature cluster and texture feature cluster: The first attention head may focus on color features (RGB mean, variance), and the second head may focus on texture features (contrast, correlation). This module learns the interaction weights between these two feature clusters through a bilinear matrix.
[0033] 4. Cross-modal interaction (k-target—kk—kk): (k-target—kk—kk, representing the relationship between feature-target relationship attention k-target and relationship-relationship kk—kk) The integration and decision-making of multi-source information simulates the integration of multimodal information such as vision and touch by the prefrontal cortex.
[0034] Mathematical model: Where: is the projection vector of the feature-target relationship, is the mapping function that projects the feature-target attention weights into the shared interaction space, is the feature-target attention weight, is the matrix used to linearly transform the feature-target attention weights, is the projection vector of relation-relation weight, is the multi-head relationship weight matrix, is the mapping function that projects the relationship-relationship weights into the shared interaction space, is the matrix used to linearly transform the multi-head relationship weight matrix, is the adjustment factor calculated through cross-modal interaction, The final output cross-modal interaction relationship weight is composed of the multi-head relationship weight matrix and the adjustment factor Element-wise multiplication ( )get, is a linear rectification function, Interactive core.
[0035] First, the feature-target attention weight and relationship-relationship weights are projected into the shared interaction space, and then through the tensor product , that is, the relationship between feature-target relationship and relationship-relationship relationship calculates the interaction of the two attentions, and then uses the interaction kernel to adjust, and finally through Activate Generation Adjustment Factor , and multiply them step by step using the multi-head relationship weight matrix. The module models feature-target weights and relation-relation interactions.
[0036] The synergistic effect of color and texture features on grayscale values: When the mean value of the red channel is high, the model may enhance the influence of texture features (such as contrast) on grayscale values. This interaction is calculated using tensor products and interaction kernel matrices.
[0037] 4. Class Learning Behavior and Neural Networks The above conditional features, enhanced by the attention mechanism, are then input into the DynamicWeightNet neural network to map multiple nodes and hidden layers. Neural network model: Where: is the predicted output value of the neural network for the target ash value, is the activation function, is the input feature vector, The final layer. Neural networks can learn complex feature relationships and input multiple features. conduct The weight parameters of the final layer are weighted summed to activate the function The nonlinear characteristics of Matrix multiplication or vector dot product operation.
[0038] In addition to mapping the four enhanced connections mentioned above, the following three core algorithm mechanisms are added based on the characteristics of neural networks: 1. Experiential Memory and Heuristic Reasoning By calculating the similarity between the 18-dimensional feature vector (basic statistics, color, texture, edge) of the sampled image features and historical memory, the neural network is supplemented by the experience library function to achieve predictive reasoning.
[0039] Mathematical model: Where, For the current sample The comprehensive similarity score with historical samples, is the current sample vector feature, is the vector feature of the historical sample, is the feature-target relationship weight vector of the current sample, For historical samples The feature-target relationship weight vector, is the feature-feature relationship weight matrix of the current sample, For historical samples The feature-feature relationship weight matrix, is the cosine similarity function.
[0040] In the prediction stage, the memory bank obtains the weighted average fusion formula: Where, For historical memory The weighting coefficient of For current samples and historical memory The comprehensive similarity of For historical memory Confidence in predictions; is the total number of samples stored in the historical memory library, m is the historical sample sum index, is the comprehensive similarity between the current sample and the historical sample m, is the prediction confidence of historical sample m.
[0041] Use the MemoryBank class to store the 18-dimensional features of historical samples, target-feature attention weights (k_target_weights), feature-feature relationship weights (k_k_weights), and their corresponding gray values: When making the next prediction, the statistical features (mean, variance) correspond to the initial encoding of light intensity by the bionic retina, the GLCM texture features correspond to the selectivity of the bionic V1 area neurons for direction, and the attention mechanism with two attention weights as the feature measures the similarity between the predicted target and its two feature values, which together constitute a multi-dimensional clue for memory retrieval to retrieve empirical features. The closest historical training data is then added to the prediction process for prediction.
[0042] 2. Error feedback and synaptic plasticity By judging the difference between the predicted value and the verification target value, the weight of the sampled image features is adjusted. The principle of bionics adjusting the synaptic strength through electrical signal feedback is: Regularization constrains the attention weights of 18-dimensional features, simulates the "pruning" process of neuronal synapses, and retains key features (such as edges and textures). The model uses regularization with attention. The loss function is used for feedback adjustment.
[0043] Mathematical model: ;
[0044] The core of the attention mechanism is to calculate the importance weights of the input features: in, is the set of all learnable parameters of the neural network, To express the parameter Minimize the value of the subsequent expression for the optimization goal, is the loss function, is the function mapping, For the The feature vector of the training samples, represents the prediction function of the neural network model, For the The true gray value of the training samples, is the regularization coefficient, which controls the strength of the regularization term ( The larger the value, the stronger the constraint on the weight). is the weight matrix of the attention mechanism, which is used to measure the importance of input features (such as the attention score in Transformer, the attention mask in CNN, etc.), is the calculation function of the attention mechanism, is the query vector, is the key vector, is a value vector, is the normalized exponential function, is the transpose of the key vector matrix, is the dimension of the key vector.
[0045] Norm ( Regularization): , that is, the sum of the absolute values of the weight elements, is the weight element.
[0046] The reward and punishment measures of biomimetic training for the quality of the adjusted predicted value are: MAE method is used for judgment. MAE, namely Mean Absolute Error, is a regression task evaluation indicator. Its essence is to calculate the average value of the absolute deviation between the predicted value and the true value. It is a basic robust loss function, mainly used to intuitively measure the The optimization results of MAE are not directly involved in training and optimizing the model, but The loss is consistent with the core logic of MEA (both are based on absolute error). When the loss is detected, the neural network model is made to approach the model range evaluated by MAE, and then the various parameters of the neural network are adjusted so that the predicted value of the trained model gradually approaches the actual value.
[0047] MAE is the average difference between the N predicted values and the N actual target grayscale values. If the deviation is too large, it is necessary to use RMSELoss for adjustment. If the deviation is not large, it is necessary to use RegularizedSmoothL1Loss (SmoothL1 loss function with attention regularization) for adjustment. Finally, the state with the minimum MAE value is saved as the neural network and its eigenvalues.
[0048] In the DynamicWeightNet network, Dropout is added to the input layer to randomly discard some neurons to improve generalization. Using BatchNorm to normalize the input of each layer and ReLU to regularize the input of each layer helps alleviate the vanishing / exploding gradient problem.
[0049] For example, the model uses error feedback to weaken the connection of redundant features (such as background color) and strengthen the weight of key features (such as object edges). Or it uses error feedback to weaken reflective noise features (such as RGB fluctuations of water surface highlights) and strengthen the texture features of sediment particles (such as GLCM contrast).
[0050] 3. Dynamic Learning Rate and Meta-Learning By dynamically adjusting the learning rate by combining feature importance and feature stability, we simulate the mechanism of the biological brain adjusting neural plasticity according to task difficulty, and dynamically adjust the learning rate according to the stability of the sampled image features (such as texture variance) to adapt to samples of different concentrations.
[0051] Mathematical model: ; Bionic brain learning logic, The step length is The learning rate, is the initial learning rate, is the target association influence coefficient, is the average attention weight of feature importance, is the influence coefficient of the relationship between conditions, is the variance of the feature relationship.
[0052] 1) Similar to the brain's strengthening of key neuronal connections, for example, in coal slime water prediction, the attention weight of the "GLCM contrast" feature is high ( becomes larger), the learning rate increases to accelerate the learning of this feature.
[0053] 2) When the brain recognizes familiar patterns (such as stable coal slime water texture characteristics), the release of neurotransmitters (such as dopamine) increases, accelerating synaptic plasticity, corresponding to the formula When the learning rate increases, the learning rate increases; when faced with complex tasks (such as the fluctuation characteristics of turbid coal slime water), the release of neuromodulators decreases and the learning rate decreases to avoid overfitting, corresponding to Increasing α causes the learning rate to decay.
[0054] 5. Class-Aware Computing and Sample Prediction 1. Feature weighting mechanism In grayscale prediction, we use biological intelligence mapping and a focused awareness mechanism. We selectively enhance 18-dimensional features (such as grayscale mean, GLCM contrast, and edge energy) through fused_weights, mimicking the biological approach of the prefrontal cortex to regulating visual attention. The fused_weights here act as synaptic plasticity adjustments, enabling the model to automatically focus on key features.
[0055] For example, in turbid coal slime water, the weights of "texture contrast" and "edge clarity" will be significantly increased, while the weights of interfering features such as background reflections will be suppressed.
[0056] 2. Implementation of the memory-decision closed loop The memory bank mechanism forms a complete "perception-attention-memory-decision-correction" cycle that realizes the consciousness-like experience storage and retrieval of samples in ash content prediction.
[0057] For example, when encountering a new sample, the system will retrieve the historical sample with the most similar features in the memory library (such as texture feature similarity > 85%) and fuse its gray value as prior knowledge with the current model prediction.
[0058] The calculation steps for this solution are as follows: input data → feature extraction layer (convolution / recurrent, etc.) → multi-layer relational attention weighting → fully connected layer → memory fusion (prediction phase) → output prediction value → error calculation using loss function → backpropagation → network and parameter adjustment → output features → load features + samples to be predicted → predict gray value. The MAE error feedback process is: prediction error (MAE) → regression loss (SmoothL1) → neural network fc6 → fc5…fc1 → fused features; the fused features are then split into two branches: one for backpropagation calculation of the fused_weights feature and the other for backpropagation calculation of the kk-final_weights feature.
[0059] The generation of a raw sampled grayscale image is a physical conversion process from optical signal to electrical signal to digital signal. It directly carries the underlying visual information and provides the raw data foundation for subsequent feature extraction and model analysis. Color channel comparison statistically quantifies the "information richness" of a channel. In scenarios such as image optimization, quality assessment, and feature extraction, it provides a quantitative basis for data-driven decision-making based on sampled images. Color channel mean comparison converts the complex visual information of a sampled image into comparable numerical values, enabling automated feature extraction and decision-making. Color channel distribution is the "statistical fingerprint" of the sampled image's pixel values. The shape of the histogram allows intuitive judgment of the channel's brightness characteristics, detail richness, and potential issues. Edge images simplify the sampled image from visual information into structured contour features, supporting high-level object recognition and scene understanding. Texture feature comparison quantifies the "regularity" of surface patterns in the sampled image to determine the material, state, and category of an object. When color and edge information are insufficient to distinguish objects, texture features often provide more detailed differentiation. Edge intensity distribution represents the distribution of intensity variations in pixel intensity within a sampled image, describing the strength of edges at different locations within the image.
[0060] The distribution of training feature weights is the result of a feature weighting mechanism, which is used by the neural network to dynamically weight features. Driven by the attention mechanism, it increases the weights of features relevant to the training objective and decreases the weights of irrelevant or weakly relevant features. This dynamic strengthening mechanism makes the neural network more intelligent, mimicking the thinking of a biological brain. This section describes the transformation of feature weights in each echo, as well as the magnitude of the MAE estimate. A large MAE indicates a significant discrepancy between the predicted value and the target value, and confidence in this data is reduced or discarded. A small or near-zero MAE increases confidence in this data, and the training features with the lowest MAE are saved after training. The training-side human-computer interface includes a training storage module, a save module, and a camera display module. The prediction-side display interface includes a window displaying the predicted grayscale value curve and a camera display module.
[0061] The training end operations are as follows: (1) First, set the sampling camera's shooting time interval and sampling image storage location, and observe the situation of the sampling tailings coal slurry water platform to see whether the image is overexposed and whether there are stains on the observation platform.
[0062] (2) Then, enter the actual ash value in the entry, and use the name of the coal preparation plant to create the sample name under this coal preparation plant, and enter the complete entry into this sample name directory.
[0063] (3) Enter the directory where the training results will be stored under "Output Training Directory" and name the training result .pth file.
[0064] (4) Then the neural network training begins, and the program automatically takes the result of the best MAE value and saves it.
[0065] The input terminal operates as follows: (1) Open the software and observe the camera taking pictures of the sample.
[0066] (2) Soft trigger the feed valve, discharge valve and submersible sewage pump switch as needed.
[0067] (3) Set the sampling interval and load the neural network just trained.
[0068] (4) Click Start Prediction, and the predicted ash value will be displayed in the upper right corner of the software interface and displayed in the real-time ash chart in the form of a broken line.
[0069] This scheme proposes a feedback-regulated bionic neural network based on a multi-layer connected attention mechanism, bionic "meta-learning" and "minimum consciousness system" principles, and uses a four-layer attention mechanism and feedback regulation method to simulate the way brain neuron clusters collaboratively process information.
[0070] Traditional neural networks are trained based on a single relationship and are trained only at a fixed learning rate, resulting in inaccurate training results. This patent introduces a dynamic learning rate adjustment mechanism based on the construction of a four-layer attention mechanism consisting of feature-target (k-target), feature-feature (kk), relationship-relationship (kk-kk), and cross-modal (k-target-kk-kk) to simulate the brain's process of adjusting neural plasticity according to task difficulty. The learning rate is automatically adjusted according to the task load rate, and combined with Dropout, ReLU, and SmoothL1 loss functions, it effectively prevents the "pseudo-death" state during training, effectively improves training accuracy, avoids repeated ineffective training, and improves the prediction accuracy of tailings ash value and training stability.
[0071] Traditional neural networks have fixed weights and no memory mechanism. Training relies solely on programming, with no empirical data to draw upon during each training session. This results in inconsistent training results and progress, and slow training. This patented bionic brain-inspired synaptic strength / pruning mechanism utilizes the gradient of the loss function to update weights, driving the immediate adjustment and optimization of all weights (including attention weights). This patented bionic brain-inspired model mimics the brain's "perception-memory-action" consciousness cycle and its ability to rapidly generalize from limited experience. It memorizes empirical data from each training session, enabling heuristic reasoning and improving model generalization, allowing the model to adapt to a wider range of training scenarios. Furthermore, it incorporates a MAE evaluation mechanism to save the best training feature results. This gives neural networks intelligence and makes the training process more flexible.
[0072] Compared to complex and potentially health-threatening detection methods such as X-ray fluorescence spectroscopy and radioisotope analysis, as well as the precision limitations of traditional image analysis methods, this method is entirely based on camera images and computational models, is safe and simple to operate, and can be easily deployed on-site at coal preparation plants. Through the combined action of the aforementioned bionic neural network structure and algorithm, a high-level, online, real-time prediction of the water-ash content of tailings coal slurry is ultimately achieved. This provides a critical prerequisite for the precise proportioning of flotation reagents, which is expected to significantly improve clean coal recovery, reduce resource waste, and guide the further resource utilization of tailings, with important industrial application value and economic benefits.
[0073] Example 2: The tailings ash value prediction system based on the bionic neural network can implement the tailings ash value prediction method based on the bionic neural network described in Example 1, including: Image acquisition module: used to acquire tailings slime water collection images; Feature extraction module: used to extract multi-dimensional features from tailings slime water images; Dynamic weighting module: used to dynamically weight the extracted multi-dimensional features through a multi-layer relational attention mechanism; Gray value prediction module: used to input the dynamically weighted features into the pre-trained neural network to predict the gray value; Ash value output module: used to output ash value prediction results.
[0074] Example 3: The embodiment of the present invention further provides a tailings ash value prediction device based on a bionic neural network, which can implement the tailings ash value prediction method based on a bionic neural network described in the first embodiment, including a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the following method: Acquire tailings slime water collection images; Perform multi-dimensional feature extraction on tailings slime water collection images; Dynamically weighted based on the extracted multi-dimensional features through a multi-layer relational attention mechanism; The dynamically weighted features are input into the pre-trained neural network to predict the gray value; Output the ash value prediction result.
[0075] Example 4: The embodiment of the present invention further provides a computer-readable storage medium, which can implement the tailings ash value prediction method based on a bionic neural network described in Example 1, and stores a computer program thereon. When the program is executed by a processor, the steps of the following method are implemented: Acquire tailings slime water collection images; Perform multi-dimensional feature extraction on tailings slime water collection images; Dynamically weighted based on the extracted multi-dimensional features through a multi-layer relational attention mechanism; The dynamically weighted features are input into the pre-trained neural network to predict the gray value; Output the ash value prediction result.
[0076] It is understood from common technical knowledge that the present invention may be implemented by other embodiments that do not depart from its spirit or essential features. Therefore, the embodiments disclosed above are, in all respects, merely illustrative and not exclusive. All modifications within the scope of the present invention or equivalent to the scope of the present invention are intended to be encompassed by the present invention.
[0077] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0078] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0079] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. The tailings ash value prediction method based on bionic neural network is characterized by: include: Acquire tailings slime water collection images; Perform multi-dimensional feature extraction on tailings slime water collection images; Dynamically weighted based on the extracted multi-dimensional features through a multi-layer relational attention mechanism; The dynamically weighted features are input into the pre-trained neural network to predict the gray value; Output the ash value prediction result.
2. The method for predicting tailings ash content based on a bionic neural network according to claim 1, wherein The multidimensional features include grayscale features, color features, texture features and edge features, wherein: The grayscale feature includes grayscale mean; The color features include R / G / B channel mean, R / G / B channel standard deviation, and R / G / B channel variance; The edge feature is a Sobel edge feature, including edge mean, variance, and energy; The texture features are GLCM texture features, including contrast, homogeneity, correlation, energy, and entropy.
3. The tailings ash value prediction method based on bionic neural network according to claim 1, wherein The multi-layer relational attention mechanism includes: a feature-target relational attention module, a feature-feature relational attention module, a feature-feature relational attention module, and a cross-modal multi-interaction relational attention module, wherein: The feature-target relationship attention module is: Where: is the feature-target attention weight, is the Sigmoid function, is the first linear layer, is the second linear layer, is the activation function, is the input feature vector.
4. The method for predicting tailings ash value based on a bionic neural network according to claim 3, wherein The feature-feature relationship attention module is: Where: is the relationship weight, is the Sigmoid function, is the input feature vector, is the projection matrix, is the relationship matrix, is the number of heads of multi-head attention.
5. The tailings ash value prediction method based on bionic neural network according to claim 3, wherein The mathematical model of the feature-feature relationship attention module is: Where: is the new relationship matrix after the interaction between heads, is the normalized exponential function, is the multi-head relationship weight matrix, is element-wise multiplication, is a learnable bilinear matrix, is the Sigmoid function, is the gating weight, is the gating weight matrix used to Mapping to gating weights Dimensions, is the mean value function, is the final result of the feature-feature relationship attention module.
6. The method for predicting tailings ash value based on a bionic neural network according to claim 3, wherein: The mathematical model of the cross-modal multi-interaction attention module is: Where: is the projection vector of the feature-target relationship, is the mapping function that projects the feature-target attention weights into the shared interaction space, is the feature-target attention weight, is the matrix used to linearly transform the feature-target attention weights, is the projection vector of relation-relation weight, is the multi-head relationship weight matrix, is the mapping function that projects the relationship-relationship weights into the shared interaction space, is the matrix used to linearly transform the multi-head relationship weight matrix, is the adjustment factor calculated through cross-modal interaction, is the cross-modal interaction relationship weight of the final output, is a linear rectification function, Interactive core.
7. The method for predicting tailings ash value based on a bionic neural network according to claim 1, wherein The neural network includes a neural network model, a similarity calculation model, a weight adjustment model and a learning rate adjustment model, wherein: The neural network model is: Where: is the predicted output value of the neural network for the target ash value, is the activation function, is the input feature vector, is the final layer; The similarity calculation model is: Where: For the current sample The comprehensive similarity score with historical samples, is the current sample vector feature, is the vector feature of the historical sample, is the feature-target relationship weight vector of the current sample, For historical samples The feature-target relationship weight vector, is the feature-feature relationship weight matrix of the current sample, For historical samples The feature-feature relationship weight matrix, is the cosine similarity function, For historical memory The weighting coefficient of is the comprehensive similarity between the current sample and the historical memory k, is the prediction confidence of historical memory k; is the total number of samples stored in the historical memory library, m is the historical sample sum index, is the comprehensive similarity between the current sample and the historical sample m, is the prediction confidence of historical sample m; By judging the difference between the predicted value and the verification target value, the weight of the sampled image feature is adjusted. The weight adjustment model is: Where: is the set of all learnable parameters of the neural network, To express the parameter Minimize the value of the subsequent expression for the optimization goal, is the loss function, is the function mapping, For the The feature vector of the training samples, represents the prediction function of the neural network model, For the The true gray value of the training samples, is the regularization coefficient; is the calculation function of the attention mechanism, is the weight matrix of the attention mechanism, is the query vector, is the key vector, is a value vector, is the normalized exponential function, is the transpose of the key vector matrix, is the dimension of the key vector, is the weight element; The learning rate adjustment model is: Where: The step length is The learning rate, is the initial learning rate, is the target association influence coefficient, is the average attention weight of feature importance, is the influence coefficient of the relationship between conditions, is the variance of the feature relationship.
8. The tailings ash value prediction system based on bionic neural network is characterized by: include: Image acquisition module: used to acquire tailings slime water collection images; Feature extraction module: used to extract multi-dimensional features from tailings slime water images; Dynamic weighting module: used to dynamically weight the extracted multi-dimensional features through a multi-layer relational attention mechanism; Gray value prediction module: used to input the dynamically weighted features into the pre-trained neural network to predict the gray value; Ash value output module: used to output ash value prediction results.
9. The tailings ash value prediction device based on bionic neural network is characterized by: including processors and storage media; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Depth feature fusion and optimization method and system for multi-modal data
CN117909922A
Data prefetching method based on multi-head attention mechanism and RNN-LSTM network
CN120196565A
Home abnormal state signal detection method and system based on multi-mode sensing
CN120216965A
Catalytic cracking unit key index modeling method based on time sequence feature extraction
WO2024021536A1
Cited By
Method and device for predicting coal slime water treatment dosage based on granulation attention mechanism
CN122245529A