Inorganic particle characteristic prediction method based on deep learning and multi-scale characteristic extraction
By using deep learning and multi-scale feature extraction methods, combined with image and numerical data, the problems of low efficiency and large errors in inorganic particle detection were solved, and accurate prediction and analysis of the characteristics of inorganic particles in tunnel construction wastewater were achieved.
Patent Information
- Application Number
- CN202510777690.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
Existing inorganic particle detection methods are inefficient and have large errors when processing large amounts of data, and fail to effectively integrate multi-source data and multi-scale features, resulting in errors in prediction results.
A multi-scale feature extraction method based on deep learning is adopted to extract features through image branch, numerical branch and guidance branch, and the multi-scale feature fusion unit and cross-modal attention mechanism are used to realize the joint modeling and feature fusion of image data and numerical data.
It has achieved accurate prediction of the characteristics of inorganic particles in tunnel construction wastewater, improved the level of automation and prediction accuracy, especially the efficient prediction of particle size and surface characteristics.
Smart Images

Figure CN120673903A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cross-application technology of environmental engineering and deep learning, and more specifically, to a method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction. Background Art
[0002] With increasing attention to water pollution during tunnel construction, the analysis and monitoring of the properties of inorganic particulate matter, particularly in wastewater generated during construction, has become increasingly important. The morphology, particle size, and surface characteristics of inorganic particles in wastewater influence the degree of water contamination, directly impacting construction quality and environmental protection.
[0003] Traditional methods for detecting inorganic particles include manual observation and image processing technology, but these methods have defects such as low efficiency, large errors, and low degree of automation when processing large amounts of data. In recent years, with the rapid development of deep learning and image processing technology, especially the application of technologies such as convolutional neural networks (CNN) and U-Net architecture, new solutions have been provided for particle feature extraction and water quality prediction. However, existing systems usually focus on the extraction of a single feature and fail to effectively integrate multi-source data. In addition, most methods do not fully consider the fusion of multi-scale features and numerical data, resulting in certain errors in the prediction effect.
[0004] In view of this, this application is hereby made. Summary of the Invention
[0005] The purpose of the present invention is to provide an inorganic particle characteristics prediction method based on deep learning and multi-scale feature extraction. Through the joint modeling of image data and numerical data, combined with multi-branch network and feature fusion technology, accurate prediction and analysis of the characteristics of inorganic particles in tunnel construction wastewater are achieved.
[0006] The above technical objectives of the present invention are achieved through the following technical solutions: In one aspect, the present application provides a method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction, comprising the following specific steps: Acquire image data and numerical data of inorganic particles in the wastewater to be tested, and preprocess the image data and numerical data; The image branch is used to extract image features from the preprocessed image data, the numerical branch is used to extract numerical features from the preprocessed numerical data, and the guiding branch is used to guide and correct the image features and the numerical features. The multi-scale feature extraction unit includes the image branch, the numerical branch and the guiding branch. The guided and corrected image features and numerical features are fused through a multi-scale feature fusion unit until the fusion result meets the preset fusion conditions; The fusion results that meet the fusion conditions are input into the trained prediction model for processing to obtain the characteristic data of the inorganic particles in the wastewater to be tested.
[0007] On the basis of the above technical solution, the present invention can also be improved as follows.
[0008] Furthermore, the above fusion condition includes that the alignment loss function and the consistency constraint function do not exceed the threshold, wherein the alignment loss function is specifically:
[0009] Where, represents the alignment loss function, Indicates the image data i The embedded features of samples, Indicates the first j The embedded features of samples, represents the feature mapping function, represents the number of embedded features in the image data, represents the number of embedded features in numerical data, is the reproducing kernel Hilbert space.
[0010] Furthermore, the above consistency constraint function is specifically:
[0011] Where, represents the consistency constraint function, The fused feature vector representing the image features, Fused feature vector representing the numerical features.
[0012] Furthermore, the above fusion results are specifically as follows:
[0013] Where, Represents the fusion result, which is the multimodal unified feature representation after the final fusion. represents the image features extracted by the image branch, represents the numerical features extracted by the numerical branch, Represents the features extracted by the guide branch, and guides and corrects through the extracted features, 、 、 Represent the corresponding dynamic weight coefficients, and + + =1.
[0014] Furthermore, the above dynamic weight coefficient is determined by the following method:
[0015] Where, 、 、 are the corresponding dynamic weight coefficients, is a trainable weight scalar.
[0016] Furthermore, the above prediction model is trained in the following way: Acquire sample data, and construct training samples through the sample data processed by the multi-scale feature extraction unit and the multi-scale feature fusion unit, wherein the sample data includes image samples and numerical samples of inorganic particles in the wastewater; The training samples are input into the initial model for processing, and the loss function value of the initial model corresponding to each data in the training sample is calculated. If multiple loss function values meet the preset training end conditions, the initial model that meets the training end conditions is determined as the prediction model. If multiple loss function values do not meet the training end conditions, the model parameters of the initial model are adjusted, and the initial model is continued to be trained based on the adjusted model parameters until multiple loss function values meet the training end conditions.
[0017] Furthermore, the above loss function value is specifically:
[0018] Where, represents the loss function value, represents the prediction weight of the indicator SSA, represents the predicted value of the indicator SSA, represents the true value of the indicator SSA, Represents the forecast weight of the indicator PV, Indicates the predicted value of the indicator PV, Indicates the true value of the indicator PV, represents the prediction weight of the indicator AP, represents the predicted value of the indicator AP, Represents the true value of the indicator AP, N Indicates the number of data in the training sample.
[0019] In a second aspect, the present application provides an inorganic particle property prediction system based on deep learning and multi-scale feature extraction, which is applied to any of the inorganic particle property prediction methods based on deep learning and multi-scale feature extraction in the first aspect, including: A data processing module is used to obtain image data and numerical data of inorganic particles in the wastewater to be tested, and to preprocess the image data and numerical data; A multi-scale feature extraction module is used to extract image features from pre-processed image data using an image branch, extract numerical features from pre-processed numerical data using a numerical branch, and guide and correct the image features and numerical features using a guide branch. The multi-scale feature extraction unit includes an image branch, a numerical branch, and a guide branch. A multi-scale feature fusion module is used to fuse the guided and corrected image features and numerical features through a multi-scale feature fusion unit until the fusion result meets the preset fusion conditions; The characteristic data acquisition module is used to input the fusion results that meet the fusion conditions into the trained prediction model for processing to obtain the characteristic data of the inorganic particles in the wastewater to be tested.
[0020] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the methods in the first aspect when executing the computer program.
[0021] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable a computer to execute any one of the methods in the first aspect.
[0022] Compared with the prior art, the present invention has at least the following beneficial effects: In this application, the multi-scale feature extraction unit adopts a three-layer architecture design, including an image branch, a numerical branch, and a guidance branch. It can accurately predict the characteristics of inorganic particles in tunnel construction wastewater through the collected image data and numerical data; at the same time, multi-scale features are extracted through a multi-branch network architecture, combined with a cross-modal attention mechanism and a feature fusion strategy, to achieve efficient prediction of parameters such as particle size (such as D50, D90) and surface characteristics (such as specific surface area, pore volume, average pore size); this method can improve the automation level of tunnel construction wastewater pollutant monitoring with higher accuracy and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings: Figure 1 A flowchart of a prediction method according to an embodiment of the present invention; Figure 2 A connection diagram of a prediction system according to an embodiment of the present invention; Figure 3 : is a comparison chart of the predicted results and the actual results of D50 in the embodiment of the present invention; Figure 4: is a comparison chart of the predicted result and the actual result of D90 in an embodiment of the present invention; Figure 5 1 is a comparison chart of the predicted results and the actual results of the specific surface area in the embodiment of the present invention; Figure 6 A comparison chart of the predicted and actual pore volume results in an embodiment of the present invention; Figure 7 3 is a comparison chart of the predicted results and the actual results of the average pore size in the embodiment of the present invention. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0026] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0027] In the description of the embodiments of the present invention, "a plurality of" means at least two.
[0028] Example 1: In order to solve the problems that the current prediction methods usually focus on the extraction of a single feature, fail to effectively integrate multi-source data, and do not fully consider the integration of multi-scale features and numerical data, resulting in certain errors in the prediction effect; this embodiment provides an inorganic particle property prediction method based on deep learning and multi-scale feature extraction, such as Figure 1 As shown, the following specific steps are included: S1, obtaining image data and numerical data of inorganic particles in the wastewater to be tested, and preprocessing the image data and numerical data.
[0029] Among them, image data of the characteristics of inorganic particles in tunnel construction wastewater are collected, including SEM characterization images, optical microscope images, TEM images, etc.; numerical data of the characteristics of inorganic particles in tunnel construction wastewater are collected, including construction methods, particle lithology, wastewater flow, water quality parameters (pH, turbidity, SS concentration, conductivity, anion and cation concentrations), etc.
[0030] Specifically, the above-mentioned preprocessing is performed by a multivariate data preprocessing unit, which includes an image preprocessing unit and a numerical preprocessing unit; wherein, the image preprocessing unit: the image input data includes an image segmentation network based on an improved U-Net architecture, which realizes image boundary recognition and accurate extraction of regions; the image preprocessing unit: the image input data includes an image segmentation network based on an improved U-Net architecture, which realizes image boundary recognition and accurate extraction of regions.
[0031] S2, using the image branch to extract image features from the preprocessed image data, using the numerical branch to extract numerical features from the preprocessed numerical data, and using the guide branch to guide and correct the image features and numerical features. The multi-scale feature extraction unit includes an image branch, a numerical branch and a guide branch.
[0032] Specifically, a three-branch parallel architecture is used in the multi-scale feature extraction unit to implement multi-scale feature extraction, as follows: 1) Image branch: The input extracts spatial structural features such as texture, edges, and morphology from the image. The core adopts an improved convolutional neural network (CNN) structure and introduces densely connected blocks (DenseBlock) and attention mechanism (CBAM) to enhance the ability to capture and express multi-scale features. The image branch adopts the following hierarchical structure: Input layer: accepts a single-channel SEM image of size 224×224×1; Initial convolution layer: Use a large 7×7 convolution kernel for preliminary feature extraction; Feature extraction module: contains two DenseBlock modules and a transition layer; Attention module: Introducing CBAM (Convolutional Block Attention Module) to enhance channel and spatial attention; Global pooling and output: Use Global Average Pooling (GAP) to obtain a fixed-length feature vector.
[0033] The DenseBlock module is composed of multiple layers of convolutional units. Each layer receives the output of all previous layers as its input to achieve maximum feature reuse. The structure of each convolutional unit is as follows: BatchNorm—ReLU—3×3 Conv (growth_rate=32) The schematic formula is as follows:
[0034] Where Hl represents the l-th layer convolution unit and · represents feature concatenation.
[0035] The CBAM attention mechanism module consists of two submodules: the Channel Attention Module and the Spatial Attention Module, which are sequentially cascaded to improve feature representation. The Channel Attention Module calculates weights using max pooling and average pooling followed by two shared fully connected layers, while the Spatial Attention Module performs channel pooling and then uses convolutional kernels to extract spatial attention weights.
[0036] 2) The numerical branch extracts key statistical features from structured data representing tunnel wastewater environmental information, including but not limited to water quality parameters (such as pH, conductivity, and SS concentration) and effluent flow rate. Due to the stable structure and moderate dimensionality of this data, it is well-suited for processing using a lightweight fully connected neural network (MLP), supplemented by normalization preprocessing and dropout to mitigate overfitting.
[0037] Its network structure adopts a three-layer fully connected network structure, and introduces Batch Normalization and Dropout between each layer to enhance the stability and generalization ability of the model.
[0038] The schematic formula is as follows:
[0039] Where m: current sample size (batch size); x i : The activation value of the i-th sample in this layer (note that this is calculated independently for each feature channel. For the convolutional layer, the mean and variance are calculated on each channel along the spatial dimension and batch dimension); μB, the mean vector of the current mini-batch (each feature channel has a mean) The structural features are: lightweight structure, fewer MLP structural parameters, and fast training speed; strong scalability: the input dimension and hidden layer width can be flexibly increased to accommodate more numerical variables; high interpretability, the middle layer nodes can be regarded as feature aggregation units, which is convenient for visual analysis; fusion-friendly, the output is a medium- and low-dimensional feature vector, which is convenient for feature fusion with the image branch and the guidance branch.
[0040] 3) The guidance branch inputs auxiliary information or prior knowledge to improve the overall performance of the multi-scale feature recognition model. In this system, the guidance branch primarily uses annotations, regular features, or auxiliary signals to guide and correct the features extracted by the image and numerical branches. The guidance branch structure consists of two fully connected layers, combined with batch normalization and ReLU activation, to maintain the same output feature dimensions as the numerical and image branches, facilitating fusion.
[0041] First layer mapping:
[0042] Second layer mapping:
[0043] Where: W1∈R d1×dg , b1∈R d1 is the first layer weight and bias; W2∈R df×dg , b1∈R df为 The second layer weights and biases ReLU()=max(0) are used as activation functions to increase nonlinear expression capabilities.
[0044] S3, fusing the guided and corrected image features and numerical features through a multi-scale feature fusion unit until the fusion result meets the preset fusion conditions.
[0045] Among them, in order to fully integrate the multi-scale and multi-modal features extracted by the image branch, numerical branch and guidance branch, a hierarchical fusion strategy can be adopted; the fusion module mainly includes: attention mechanism enhancement, cross-modal feature alignment and consistency constraints, and dynamic weight fusion strategy to achieve effective information integration and redundancy suppression.
[0046] 1) CBAM Attention Mechanism: CBAM (Convolutional Block Attention Module) is a lightweight attention mechanism consisting of two sub-modules: channel attention and spatial attention. It aims to enhance the expressiveness of key feature channels and spatial positions.
[0047] Channel attention mechanism formula:
[0048] where F∈R C×H×W: Input feature map; AvgPool(F) performs global average pooling on each channel to obtain a C×1×1 vector; MaxPool(F) performs global maximum pooling on each channel to obtain a C×1×1 vector; MLP two-layer perceptron: generally C→C / r→C, to achieve nonlinear modeling between channels; σ is the Sigmoid activation function, which normalizes the attention weight to [0,1]; Mc(F) outputs the channel attention map (one weight for each channel).
[0049] Spatial attention module formula:
[0050] Among them F c : Feature map after channel attention weighting (output of the previous stage); AvgPool(F c ) performs average pooling on all channels, resulting in a 1×H×W feature map; MaxPool(F c ) Perform maximum pooling on all channels, resulting in a 1×H×W feature map; [.]: represents splicing along the channel dimension to form a 2×H×W tensor; f 7×7 Is a 7x7 convolution kernel that extracts spatial attention; Mc(F c ) is the output spatial attention map.
[0051] 2) Cross-modal feature alignment and consistency constraints Since images, numerical values, and guiding features have different statistical distributions and feature spaces, an alignment mechanism needs to be introduced to improve the fusion quality. This system adopts two methods: Alignment loss function This loss ensures that the features output by the image branch and the numerical branch are "aligned" in the high-dimensional space, thereby achieving consistency in modal fusion. The specific formula is:
[0052] Where, represents the alignment loss function, Indicates the image data i The embedded features of samples, Indicates the first j The embedded features of samples, represents the feature mapping function, represents the number of embedded features in the image data, represents the number of embedded features in numerical data, is the reproducing kernel Hilbert space.
[0053] Consistency constraints This loss makes the two modalities more unified in representation and helps reduce the semantic differences between the modalities. The specific formula is:
[0054] Where, represents the consistency constraint function, The fused feature vector representing the image features, Fused feature vector representing the numerical features.
[0055] 3) Dynamic Weight Strategy Since the contribution of different modal information to the final task changes dynamically, this study introduces a learnable weight coefficient to adjust the contribution of each modality.
[0056] The fusion calculation formula is:
[0057] Where, Represents the fusion result, which is the multimodal unified feature representation after the final fusion. represents the image features extracted by the image branch, represents the numerical features extracted by the numerical branch, Represents the features extracted by the guide branch, and guides and corrects through the extracted features, 、 、 Represent the corresponding dynamic weight coefficients, and + + =1.
[0058] Optionally, the dynamic weight coefficient is determined by:
[0059] Where, 、 、 are the corresponding dynamic weight coefficients, is a trainable weight scalar.
[0060] S4, inputting the fusion results that meet the fusion conditions into the trained prediction model for processing to obtain characteristic data of the inorganic particles in the wastewater to be tested.
[0061] The prediction model adopts a multi-task collaborative prediction architecture, which consists of three independent sub-modules: feature decoding layer, parallel prediction particle size prediction unit, and surface property prediction unit: (1) Feature decoding layer. The core of the feature decoding layer is to reduce the dimension of high-dimensional fusion features to a shared feature space through a fully connected network (FC). The feature decoding layer maps the 512-dimensional fusion features to a 256-dimensional shared feature space through a fully connected network. The formula is as follows:
[0062] Shared Features h shared The input is sent to the parallel prediction analysis branch (particle size prediction unit, surface property prediction unit), and each branch generates the final prediction result through an independent fully connected layer or convolutional layer.
[0063] (2) Particle size prediction unit 1) Input features and preprocessing Input features F∈R d is the result of multimodal feature fusion, where d=256 represents the feature dimension; F integrates the morphological features extracted by the image branch, the environmental parameters of the numerical branch, and the auxiliary information of the guidance branch, and has the ability to express multi-scale and multi-dimensional information.
[0064] 2) Network structure design The network adopts a two-layer fully connected structure (Multi-Layer Perceptron, MLP), including: Hidden layer: extracts nonlinear feature representation and alleviates linear constraints between features.
[0065] Output layer: Independently outputs D50 and D90 prediction values, using activation functions to ensure physical meaning (non-negative).
[0066] The specific mathematical form is:
[0067]
[0068] where F∈R d is the input feature vector; W1∈R m×d is the hidden layer weight matrix, m=128 is the number of hidden units; b1∈R m is the hidden layer bias; h∈R m is the hidden layer activation output; σ is the hidden layer activation function, using ReLU, defined as:
[0069] W50,W90∈R 1×mis the output layer weight vector, corresponding to D50 and D90 respectively; b50, b90∈R is the output layer bias scalar; ϕ is the output activation function, using Softplus, defined as:
[0070] The smoothness of the Softplus function ensures that the predicted value is non-negative and the gradient is continuous, which is suitable for regression tasks.
[0071] 3) Network structure description Input layer: accepts multimodal fusion features, with a shape of (256); Dense is the hidden layer with 128 neurons, which performs the linear transformation W1F+b1; The ReLU activation function implements nonlinearity, maintains the positive gradient, and avoids gradient disappearance; The output layer is divided into two branches, predicting D50 and D90 respectively; The Softplus activation function ensures that the output value is non-negative and smooth, which is better than the "dead neuron" problem of simple ReLU; by simultaneously predicting D50 and D90, the model implements multi-task learning, improving feature utilization and prediction accuracy.
[0072] (3) Surface property prediction unit 1) Input feature description Similar to the particle size prediction branch, the input of this branch is the multimodal fusion feature vector:
[0073] Among them, F includes image morphological information, water quality parameters and geological construction parameters, and has good physical correlation and feature richness.
[0074] 2) Network structure design This branch adopts a three-output structure, corresponding to SSA, PV and AP respectively.
[0075] The overall structure is as follows: Input layer: fusion feature F Hidden layer: 2-layer MLP for learning complex nonlinear mapping Output layer: three independent branches, respectively regressing and predicting the target variable The specific function expression is as follows: ,
[0076] , ,
[0077] where F∈R 256 , multimodal fusion input; W1∈R128×256 ,b1∈R 128 is the first layer parameter; W2∈R 64×128 ,b2∈R 64 is the second layer parameter; σ is the activation function, using ReLU (non-linear enhancement, reducing gradient disappearance); ϕ is the output activation function, using Softplus (continuous non-negative); the predicted value of the output dimension of 1 represents SSA (unit m 2 / g), PV (unit: cm 3 / g), AP (unit: nm).
[0078] 3) Joint loss function Considering that all three indicators are continuous physical quantities, a weighted multi-task mean square error (MSE) loss function is used, specifically:
[0079] Where, represents the loss function value, represents the prediction weight of the indicator SSA, represents the predicted value of the indicator SSA, represents the true value of the indicator SSA, Represents the forecast weight of the indicator PV, Indicates the predicted value of the indicator PV, Indicates the true value of the indicator PV, represents the prediction weight of the indicator AP, represents the predicted value of the indicator AP, Represents the true value of the indicator AP, N Indicates the number of data in the training sample.
[0080] Optionally, the above prediction model is trained in the following way: Acquire sample data, and construct training samples through the sample data processed by the multi-scale feature extraction unit and the multi-scale feature fusion unit, wherein the sample data includes image samples and numerical samples of inorganic particles in the wastewater; The training samples are input into the initial model for processing, and the loss function value of the initial model corresponding to each data in the training sample is calculated. If multiple loss function values meet the preset training end conditions, the initial model that meets the training end conditions is determined as the prediction model. If multiple loss function values do not meet the training end conditions, the model parameters of the initial model are adjusted, and the initial model is continued to be trained based on the adjusted model parameters until multiple loss function values meet the training end conditions.
[0081] The system training adopts a phased progressive training strategy, sequentially optimizing the image segmentation network, feature extraction module, and prediction convergence performance to ensure the coordinated convergence of each module. The overall process is as follows: (1) Image segmentation network pre-training, key steps: 1) Data loading: Input: original image (with pixel annotations) 2) Model initialization: Encoder: Load ResNet34 pre-trained weights (ImageNet) Decoder: He normal distribution initialization Optimizer: AdamW (lr=1e-4, weight decay=1e-4) 3) Training strategy: Early stopping condition: The IoU (Intersection over Union) of the training set has not improved for 5 consecutive epochs Output: Segmentation mask (512×512 pixels, binary image) (2) Multi-scale feature extraction module training Key steps: 1) Data flow construction: Input: Image: Segmented image output from stage 1 (512×512); Value: Normalized data (time-aligned) 2) Module freeze: Fixed the image segmentation network weights and only trained the feature extraction and fusion modules 3) Branch training strategy: The learning rates of the image branch, numerical branch, and guidance branch can be 3e-4, 5e-4, and 1e-3, respectively.
[0082] The following example further illustrates the operation process and prediction results based on 50 sets of inorganic particle samples from tunnel construction wastewater. Step 1: Data collection and integration Input: 50 sets of data were collected from wastewater and inorganic particles at the tunnel construction site. Each set of data contained: Image data: includes 50 SEM characterization images in TIFF format with a resolution of 2000×2000.
[0083] Numerical data: including tunnel construction method, particle lithology, water flow rate, water quality parameters (pH, turbidity, SS concentration, conductivity, anion and cation concentrations, etc.).
[0084] Processing: Integrate each set of data into a unified format to ensure that the image data and numerical data are aligned on the time axis and ensure data synchronization.
[0085] Output: The integrated data set, including image data and numerical data, is ready for subsequent processing.
[0086] Step 2: Image data preprocessing Input: The input is each set of image data (SEM characterization images) integrated in step 1.
[0087] Processing: Image segmentation is performed using a modified U-Net architecture. The encoder uses pre-trained ResNet34 weights, and the image segmentation network identifies image boundaries and extracts regions. Segmentation results are verified using Intersection over Union (IoU) to ensure accuracy.
[0088] Output: Segmentation results of each image, accurate extraction of image regions, providing a basis for subsequent feature extraction.
[0089] Step 3: Numerical Data Preprocessing Input: The input is each set of numerical data integrated in step 1 (such as construction method, granular rock properties, water flow rate, water quality parameters, etc.).
[0090] Processing: All numerical data is normalized, mapping the values to between 0 and 1 to eliminate scale differences between different data. At the same time, the consistency and synchronization of the data in time are ensured, and time series alignment is performed.
[0091] Output: Normalized numerical data. Time-series aligned data provides support for subsequent feature extraction and model input.
[0092] Step 4: Multi-scale feature extraction Input: The input is the image segmentation result of step 2 and the normalized numerical data of step 3.
[0093] Image feature extraction: The overall morphological features of the image are extracted through a 7×7 large convolution kernel and a dilated convolution; the detailed features of the image are extracted through a 3×3 standard convolution kernel, and the SE attention mechanism is used to optimize the weights of important features.
[0094] Numerical feature extraction: Extract key features from numerical data through 1D-CNN.
[0095] Output: The fusion result of the three feature branches generates a multi-dimensional feature vector as the input of the model.
[0096] Step 5: Feature Fusion Input: The input is the three feature branches in step 4.
[0097] Processing: Image and numerical features are fused using the Cross-Modal Attention Mechanism (CBAM), calculating attention weights in the spatial and channel dimensions to optimize the effect of feature fusion. The L2 norm is used to ensure semantic consistency between image and numerical features, and dynamic weighted fusion is performed.
[0098] Output: The fused feature vector, which serves as the input of the prediction model.
[0099] Step 6: Feature decoding and prediction output Input: The input is the feature vector fused in step 5.
[0100] deal with: Feature decoding: The fused features are reduced to a shared feature space (256 dimensions) through a fully connected network.
[0101] Particle size prediction: Decode particle size parameters (such as D50, D90) from shared features and use regression models to output prediction results.
[0102] Surface property prediction: Decode the surface property parameters of particles (such as specific surface area, pore volume, and average pore diameter) from shared features and output the corresponding prediction results.
[0103] Output: particle size prediction results (D50, D90) and surface characteristics prediction results (specific surface area, pore volume, average pore diameter); The prediction results based on 50 groups of inorganic particle samples from tunnel construction wastewater are shown in Table 1 and Figure 3-Figure 7 shown.
[0104] Table 1
[0105] Example 2: The present application provides an inorganic particle characteristic prediction system based on deep learning and multi-scale feature extraction, which is applied to the inorganic particle characteristic prediction method based on deep learning and multi-scale feature extraction in Example 1, such as Figure 2 As shown, including: A data processing module is used to obtain image data and numerical data of inorganic particles in the wastewater to be tested, and to preprocess the image data and numerical data; A multi-scale feature extraction module is used to extract image features from pre-processed image data using an image branch, extract numerical features from pre-processed numerical data using a numerical branch, and guide and correct the image features and numerical features using a guide branch. The multi-scale feature extraction unit includes an image branch, a numerical branch, and a guide branch. A multi-scale feature fusion module is used to fuse the guided and corrected image features and numerical features through a multi-scale feature fusion unit until the fusion result meets the preset fusion conditions; The characteristic data acquisition module is used to input the fusion results that meet the fusion conditions into the trained prediction model for processing to obtain the characteristic data of the inorganic particles in the wastewater to be tested.
[0106] Example 3: The embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method of Example 1 is implemented.
[0107] Example 4: The embodiment of the present application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable a computer to execute the method of Example 1.
[0108] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0109] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0110] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0112] Those skilled in the art will understand that all or part of the steps in implementing the above facts and methods can be completed by instructing relevant hardware through a program, and the program involved or the program can be stored in a computer-readable storage medium. When the program is executed, it includes the following steps: the corresponding method steps are then brought out, and the storage medium can be ROM / RAM, a disk, an optical disk, etc.
[0113] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. Inorganic particle property prediction method based on deep learning and multi-scale feature extraction, characterized in that: The specific steps include: Acquiring image data and numerical data of inorganic particles in the wastewater to be tested, and preprocessing the image data and the numerical data; An image branch is used to extract image features from the preprocessed image data, a numerical branch is used to extract numerical features from the preprocessed numerical data, and a guiding branch is used to guide and correct the image features and the numerical features, wherein the multi-scale feature extraction unit includes an image branch, a numerical branch, and a guiding branch; fusing the guided and corrected image features and the numerical features through a multi-scale feature fusion unit until a fusion result satisfies a preset fusion condition; The fusion results that meet the fusion conditions are input into the trained prediction model for processing to obtain characteristic data of the inorganic particles in the wastewater to be tested.
2. The method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction according to claim 1, characterized in that: The fusion condition includes that the alignment loss function and the consistency constraint function do not exceed a threshold, wherein the alignment loss function is specifically: Where, represents the alignment loss function, Indicates the image data i The embedded features of samples, Indicates the first j The embedded features of samples, represents the feature mapping function, represents the number of embedded features in the image data, represents the number of embedded features in numerical data, is the reproducing kernel Hilbert space.
3. The method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction according to claim 1, characterized in that: The consistency constraint function is specifically: Where, represents the consistency constraint function, The fused feature vector representing the image features, Fused feature vector representing the numerical features.
4. The method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction according to claim 1, characterized in that: The fusion results are specifically: Where, Represents the fusion result, which is the multimodal unified feature representation after the final fusion. represents the image features extracted by the image branch, represents the numerical features extracted by the numerical branch, Represents the features extracted by the guide branch, and guides and corrects through the extracted features, 、 、 Represent the corresponding dynamic weight coefficients, and + + =1.
5. The method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction according to claim 4, characterized in that: The dynamic weight coefficient is determined by: Where, 、 、 are the corresponding dynamic weight coefficients, is a trainable weight scalar.
6. The method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction according to claim 1, characterized in that: The prediction model is trained in the following way: Acquire sample data, and construct training samples through the sample data processed by the multi-scale feature extraction unit and the multi-scale feature fusion unit, wherein the sample data includes image samples and numerical samples of inorganic particles in the wastewater; The training samples are input into the initial model for processing, and the loss function value of the initial model corresponding to each data in the training samples is calculated. If multiple loss function values meet the preset training end conditions, the initial model that meets the training end conditions is determined as the prediction model. If multiple loss function values do not meet the training end conditions, the model parameters of the initial model are adjusted, and the initial model is continued to be trained based on the adjusted model parameters until multiple loss function values meet the training end conditions.
7. The method for predicting inorganic particle characteristics based on deep learning and multi-scale feature extraction according to claim 6, characterized in that: The loss function value is specifically: Where, represents the loss function value, represents the prediction weight of the indicator SSA, represents the predicted value of the indicator SSA, represents the true value of the indicator SSA, Represents the forecast weight of the indicator PV, Indicates the predicted value of the indicator PV, Indicates the true value of the indicator PV, represents the prediction weight of the indicator AP, represents the predicted value of the indicator AP, Represents the true value of the indicator AP, N Indicates the number of data in the training sample.
8. Inorganic particle property prediction system based on deep learning and multi-scale feature extraction, characterized by: include: A data processing module is used to obtain image data and numerical data of inorganic particles in the wastewater to be tested, and pre-process the image data and the numerical data; a multi-scale feature extraction module, configured to extract image features from pre-processed image data using an image branch, extract numerical features from pre-processed numerical data using a numerical branch, and guide and correct the image features and numerical features using a guiding branch, wherein the multi-scale feature extraction unit includes an image branch, a numerical branch, and a guiding branch; a multi-scale feature fusion module, configured to fuse the guided and corrected image features and the numerical features through a multi-scale feature fusion unit until a fusion result satisfies a preset fusion condition; The characteristic data acquisition module is used to input the fusion results that meet the fusion conditions into the trained prediction model for processing to obtain the characteristic data of the inorganic particles in the wastewater to be tested.
9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method according to any one of claims 1 to 7 is implemented when the processor executes the computer program.
10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, which enable a computer to execute the method of any one of claims 1 to 7.