Floccule image processing and identification method based on multi-modal deep learning
Through multimodal deep learning method, combined with underwater cameras and sensors to collect data, a dual-branch model was established, which solved the problem of image quality and model interpretability in floc image recognition, and realized the accurate evaluation of floc status and efficient control of effluent water quality.
Patent Information
- Application Number
- CN202510542182.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, floc image recognition methods have problems such as image quality defects, single-modal recognition ignores key features of water quality and process, and insufficient interpretability of deep learning models, resulting in inaccurate coagulation process control, high drug consumption and large fluctuations in water quality.
The multimodal deep learning method is adopted to collect floc images and monitor water quality parameters through underwater cameras, and combine the dual-branch model of adaptive data augmentation and hybrid fusion strategies to achieve end-to-end identification and prediction of floc images and water quality parameters.
The accuracy of floc image recognition and the interpretability of the model are improved, and high-precision prediction and intelligent feedback of the effluent water quality are achieved, reducing drug consumption and stabilizing water quality.
Smart Images

Figure CN120318666A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water treatment, and particularly relates to a method for processing and recognizing floc images based on multimodal deep learning. Background Art
[0002] In the coagulation process, coagulants are added to water to aggregate colloidal particles and tiny suspended solids in the water, and flocs are formed and removed in subsequent processes. Among them, the coagulant dosing link is the core of drinking water treatment, which directly affects the quality and efficiency of the produced water. Traditional coagulant dosing control relies on manual experience and prior knowledge, and realizes closed-loop control based on feed-forward prediction or feedback from a single sensor (turbidity meter), which has problems such as strong subjectivity, response lag, and poor anti-interference ability, resulting in high chemical consumption and large water quality fluctuations.
[0003] With the development of artificial intelligence and image recognition technology, extracting physical properties from floc images can more efficiently and accurately evaluate the coagulation state. For example, CN117115438A discloses an image recognition and dosing control system for alum flowers. The image recognition module processes the sampled alum flower images and extracts features such as area, perimeter, minimum circumscribed circle, roundness, compactness, eccentricity, etc. On this basis, a fully connected neural network is constructed to determine the dosing state (extremely over-dosed, over-dosed, suitable, under-dosed, extremely under-dosed), making the dosing index quantifiable, standardizing manual experience, and enabling intelligent online monitoring, thereby reducing the water production cost. However, the feature space based on feature engineering is restricted by the design and selection of artificial features, which may lead to incomplete information extraction and information loss, thus hindering model recognition. Deep learning models represented by convolutional neural networks have shown unique advantages in end-to-end image recognition. They can automatically extract and learn features at different scales and levels in floc images, avoiding the limitations of artificial feature selection. For example, CN113688751A discloses a method and device for analyzing alum flower features using image recognition technology. Convolutional neural networks are used to extract high-level features of floc images, the detection frame (position and size) of alum flowers is obtained through regression, the type of alum flowers (fluffy, sheet-like, fuzzy alum flowers) is obtained through classification, and then a post-sedimentation water turbidity prediction model is established to evaluate the quantitative index of alum flowers.
[0004] However, due to the image quality defects caused by the deficiencies in the pretreatment method of floc images and the neglect of key non-image features such as water quality and process by the single-modal recognition model, the model performance is prone to bottlenecks and the generalization ability is limited. The unclear edges and blurred boundaries with other interference factors in the water in the floc image quality, as well as the imbalance in quantity, seriously hinder the model's ability to identify flocs in water. At the same time, the existing methods only consider the single-modal floc image features and cannot process the key information carried by water quality parameters and operation process parameters, thus lacking a comprehensive simulation of the complex coagulation process. In addition, the decision-making process of the deep learning model is opaque, and it is difficult to quantify the causal contribution of floc morphology to the effluent water quality, resulting in it being difficult for actual engineering personnel to trust. Summary of the Invention
[0005] In view of the problems existing in the prior art, the purpose of the present invention is to provide a floc image recognition method based on multi-modal deep learning, which is used to solve the problems such as image quality defects caused by lack of processing methods in floc image recognition, neglect of key non-image features such as water quality / process by single-modal recognition, and insufficient interpretability of deep learning models, and to realize the quantitative evaluation of floc images and coagulation states and the accurate prediction of effluent water quality.
[0006] To achieve the above object, the present invention provides a floc image processing and recognition method based on multi-modal deep learning, including the following steps:
[0007] S1. Multi-modal data collection: Use an underwater camera to collect floc images between the coagulation unit and the sedimentation unit, and integrate sensors to monitor the influent water quality and process parameters in real time;
[0008] S2. Floc image processing: Eliminate the underwater imaging noise of flocs, divide all pixels in the floc image into a series of independent and non-connected pixel regions, and improve the clarity of the floc image;
[0009] S3. Data classification, annotation and division: Classify the floc images and their corresponding water quality / process multi-modal data according to the effluent turbidity, and store the floc images in the local folder according to the file directory structure for automatic division and annotation;
[0010] S4. Adaptive data augmentation: Perform adaptive data augmentation on the established floc image database for multi-modal deep learning model training, dynamically generate synthetic data for minority class samples, and solve the problem of data imbalance;
[0011] S5. Dual-branch Modeling Based on Hybrid Fusion Strategy: Establish a dual-branch multimodal deep learning model based on the hybrid fusion strategy. The floc morphology recognition branch performs early fusion on floc image features, influent water quality, and process parameters, and the end-to-end recognition branch directly predicts the water quality classification probability using a pre-trained convolutional neural network. The weighted decision fusion method is used for late fusion of the probability predictions of the two branches;
[0012] S6. Training of Multimodal Deep Learning Model: Use the multimodal database constructed in step S4, which contains floc images and water quality / process parameters, and use the classification annotations in step S3 as supervision labels to train the established multimodal deep learning model to obtain a trained multimodal deep learning model;
[0013] S7. Integrating Grad-CAM++ to Visualize Model Decisions: Use an underwater camera and integrated sensors to collect multimodal data of floc images, influent water quality, and process parameters. Use the trained multimodal deep learning model to determine whether the effluent water quality meets the standards, and integrate the Grad-CAM++ visualization model's region of interest to locate the key floc regions that affect the decision.
[0014] Preferably, step S2 further includes:
[0015] S201. Convert the original RGB image to a grayscale image using the weighted average method to reduce the computational amount and speed up the processing speed. The expression of the weighted average method is as follows:
[0016] Grey = 0.299×R + 0.587×G + 0.114×B
[0017] In the formula, 0.299, 0.587, and 0.114 are the standard coefficients defined by the International Telecommunication Union (ITU).
[0018] S202. Use threshold segmentation to fix the pixel values of the grayscale image to 0 or 255 to distinguish the effluent background features and floc features. The expression of the threshold segmentation is as follows:
[0019]
[0020] Among them, f(i, j) is the image after threshold processing, I(i, j) is the part composed of bright objects on the background in the grayscale image, and T is the selected threshold. All pixel points that meet the condition I≥T are recorded as target points to form the floc image, and the remaining points are recorded as background points to form the background image.
[0021] S203. Use median filtering to remove salt-and-pepper noise in the floc image.
[0022] S204. Use opening operation to remove small objects in the floc image, smooth the floc contour, disconnect narrow connections without significantly changing its area, and eliminate the protruding parts on the floc surface, which helps to remove noise and isolated small flocs and makes the main floc structure clearer.
[0023] Preferably, in step S203, a 3×3 filter kernel is used to better retain the details of the image while removing noise.
[0024] Preferably, in step S3, the multi-modal data is divided into four categories. Samples with an effluent turbidity less than 0.15 NTU are labeled as "Class A", samples with an effluent turbidity of 0.15 to 0.3 NTU are labeled as "Class B", samples with a turbidity of 0.3 to 1.0 NTU are labeled as "Class C" (1 NTU is the drinking water standard limit in China), and samples with a turbidity greater than 1.0 NTU are labeled as "Class D" (not meeting the drinking water quality standard).
[0025] Preferably, in step S3, the file directory structure is as follows: the main directory is the floc image folder, the secondary directory is the training set or the test set, and the tertiary directory is the effluent classification label.
[0026] Preferably, in step S3, the data volume in the training set in the division ratio is 80% of all data, and the data volume in the test set is 20% of all data.
[0027] Preferably, in step S4, the adaptive data augmentation further includes the following steps:
[0028] S401. Identify the most numerous class.
[0029] S402. Randomly perform five geometric transformation operations on the minority class pictures, including 90° rotation, 180° rotation, 270° rotation, horizontal flipping, or vertical flipping, and record the operations performed on each picture.
[0030] S403. If the number of the minority class is still less than that of the most numerous class, continue to perform oversampling and repeat the geometric transformation that is not recorded in the operation in S402.
[0031] S404. Repeat steps S402 and S403 until the number of the minority class pictures is not less than that of the most numerous class.
[0032] S405. If the number of the minority class pictures is greater than that of the most numerous class, perform downsampling on the generated pictures to remove the redundant pictures until the number of the minority class pictures is the same as that of the most numerous class.
[0033] Preferably, in step S5, the floc morphology recognition branch uses a feature extractor based on the pixels of the floc image to obtain the floc morphology features, and further includes the following steps:
[0034] S501. Determine the connected domains of the floc body, find out the mutually independent connected domains in the image and take them as connected units, and determine the floc positions and the number of flocs.
[0035] S502. Count the number of pixels on the periphery of each connected domain to obtain the perimeter C of each floc, and count the number of pixels within the connected domain to obtain the floc area A. The calculation formula for the equivalent diameter d of the floc is as follows:
[0036]
[0037] S503. Define the floc image density as the ratio of the number of pixels occupied by the floc particles in the image to all the pixels of the actual image. This characteristic parameter can reflect the proportion of the floc image in the water background in the image, and its calculation formula is as follows:
[0038]
[0039] In the formula, p is the floc image density, A1 is the number of pixels occupied by the floc, and A2 is the total number of image pixels.
[0040] S504. Further determine the D of the floc using the calculated C and A of the floc f , and its correlation relationship is as follows:
[0041]
[0042] Take the logarithm and use the least squares method to perform a first-order polynomial fitting to find the slope, and the slope is the fractal dimension D f .
[0043] Preferably, in step S501, the 8-nearest pixels are used to determine the connected domain of the floc body.
[0044] Preferably, in step S5, the floc image features used for early fusion include but are not limited to the number of flocs, the average equivalent diameter, the fractal dimension, and the image density.
[0045] Preferably, in step S5, a pre-trained convolutional neural network is used in the end-to-end recognition branch and fine-tuned in the weighted decision fusion to reduce the training cost.
[0046] Compared with the existing technical solutions, the present invention proposes a method for floc image processing and recognition based on multi-modal deep learning. The present invention has the following beneficial effects:
[0047] (1) By obtaining the grayscale image of the flocs and through image calculation methods such as threshold segmentation, median filtering, and opening operation, the noise is effectively eliminated and the edges of the flocs are highlighted, making the floc targets in the image clearer. At the same time, the designed adaptive data augmentation algorithm dynamically generates synthetic data, solving the problem of insufficient minority class samples in the training image data.
[0048] (2) By fusing the floc image features, influent water quality, and process parameters, a dual-branch multi-modal deep learning model is established based on the hybrid fusion strategy, improving the model's recognition ability of flocs in water, thereby achieving high-precision prediction of the effluent water quality to provide intelligent predictive feedback.
[0049] (3) Integrate the key floc regions concerned by the Grad-CAM++ model decision. By highlighting the key flocs that affect the model's judgment of the effluent water quality, it helps engineering personnel intuitively understand the basis of the model decision and enhances the trust in the predictive feedback based on multi-modal deep learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic flow chart of the steps of a floc image processing and recognition method based on multi-modal deep learning provided by the present invention;
[0051] Figure 2 It is a display diagram of the floc image processing process in the embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of the directory structure of the floc image annotation in the embodiment of the present invention;
[0053] Figure 4 It is a dual-branch multi-modal deep learning model based on the hybrid fusion strategy provided by the present invention;
[0054] Figure 5 It is the performance of the multi-modal deep learning model in the embodiment of the present invention;
[0055] Figure 6 It is the receiver operating characteristic curve and calibration curve of the multi-modal deep learning model in the embodiment of the present invention;
[0056] Figure 7 It is the decision curve of the multi-modal deep learning model in the embodiment of the present invention;
[0057] Figure 8 It is the result of the ablation experiment in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0058] The following describes the embodiments of the present invention through specific examples in combination with the accompanying drawings. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific examples. Various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. The typical but non-limiting embodiments of the present invention are as follows:
[0059] Example 1:
[0060] This embodiment specifically provides a method for processing and recognizing floc images based on multimodal deep learning, as Figure 1 shown, which specifically includes the following steps:
[0061] S1. Use an underwater camera to collect floc images between the coagulation unit and the sedimentation unit, and integrate sensors to monitor the influent turbidity, influent temperature, and coagulant dosage (including aluminum salts and iron salts) in real time.
[0062] S2. Process the floc images as Figure 2 shown, eliminate the underwater imaging noise of the flocs, divide all pixels in the floc images into a series of independent and non-connected pixel regions, and improve the clarity of the floc images;
[0063] Specifically, step S2 further includes:
[0064] S201. Use the weighted average method to convert the original RGB image into a grayscale image, reduce the computational amount, and speed up the processing speed. The expression of the weighted average method is as follows:
[0065] Grey = 0.299×R + 0.587×G + 0.114×B
[0066] In the formula, 0.299, 0.587, and 0.114 are the standard coefficients defined by the International Telecommunication Union (ITU).
[0067] S202. Use threshold segmentation to fix the pixel values of the grayscale image to 0 or 255 to distinguish the water background features and floc features. The expression of the threshold segmentation is as follows:
[0068]
[0069] where f(i, j) is the image after threshold processing, I(i, j) is the part composed of bright objects on the background in the grayscale image, and T is the selected threshold. All pixel points that meet the condition I≥T are recorded as target points to form the floc image, and the remaining points are recorded as background points to form the background image.
[0070] S203. Median filtering is used to remove salt-and-pepper noise in the floc image. A 3×3 filter kernel is used to better preserve the details of the image while removing the noise.
[0071] S204. Opening operation is used to remove irrelevant structures while maintaining the shape of the target flocs in the image, making the floc image clearer.
[0072] S3. Communicate with water plant business experts and divide the floc images and their corresponding water quality / process multi-modal data into four categories according to the effluent turbidity. Samples with an effluent turbidity less than 0.15 NTU are labeled as "Class A" (0.15 NTU is the internal drinking water standard limit of a certain water plant), samples with an effluent turbidity of 0.15 to 0.3 NTU are labeled as "Class B" (0.3 NTU is the internal drinking water standard limit of a certain water service group), samples with a turbidity of 0.3 to 1.0 NTU are labeled as "Class C" (1 NTU is the drinking water standard limit in China), and samples with a turbidity greater than 1.0 NTU are labeled as "Class D" (not meeting the drinking water quality standard); Store the floc images in the local folder according to the file directory structure as shown in Figure 3 for automatic partitioning and labeling. Among them, the data volume in the training set is 80% of all data, and the data volume in the test set is 20% of all data.
[0073] S4. Perform adaptive data augmentation on the established floc image database for multi-modal deep learning model training, dynamically generate synthetic data for minority class samples, and solve the problem of data imbalance.
[0074] Specifically, in step S4, the adaptive data augmentation further includes the following steps:
[0075] S401. Identify the majority class.
[0076] S402. Randomly perform five geometric transformation operations on the minority class pictures, including 90° rotation, 180° rotation, 270° rotation, horizontal flipping, or vertical flipping, and record the operations performed on each picture.
[0077] S403. If the number of the minority class is still less than that of the majority class, continue to perform oversampling and repeat the geometric transformation that is not recorded in the operation of S402.
[0078] S404. Repeat steps S402 and S403 until the number of the minority class pictures is not less than that of the majority class.
[0079] S405. If the number of the minority class pictures is greater than that of the majority class, perform downsampling on the generated pictures to remove the redundant pictures until the number of the minority class pictures is the same as that of the majority class.
[0080] S5. Establish a dual-branch multi-modal deep learning model based on a hybrid fusion strategy, as shown in Figure 4 . The floc morphology recognition branch uses a feature extractor based on the pixels of the floc image to obtain floc morphology features, including the number of flocs, average equivalent diameter, fractal dimension, and image density. Then, these features are concatenated with water quality and process parameters, and early fusion features are generated through a one-dimensional convolutional layer. Multilayer autoencoders are used to extract high-level features, and then a probability prediction for water quality classification is obtained through a multi-layer perceptron with 8 hidden layers and 32 nodes in each layer. The end-to-end recognition branch uses a pre-trained convolutional neural network for fine-tuning to directly predict the probability of water quality classification. Finally, a weighted decision fusion method is used to perform late fusion on the probability predictions of the two branches.
[0081] Specifically, in step S5, the extraction of the floc morphology features further includes the following steps:
[0082] S501. Use 8-nearest pixels to determine the connected domain of the floc body, find out the mutually independent connected domains in the image and regard them as connected units, and determine the floc position and the number of flocs.
[0083] S502. Count the number of peripheral pixels of each connected domain to obtain the perimeter C of each floc, and count the number of pixels within the connected domain to obtain the floc area A. The calculation formula for the equivalent diameter d of the floc is as follows:
[0084]
[0085] S503. Define the floc image density as the ratio of the number of pixels occupied by the floc particles in the image to all the pixels of the actual image. This feature parameter can reflect the proportion of the floc image in the water background in the image. The calculation formula is as follows:
[0086]
[0087] In the formula, p is the floc image density, A1 is the number of pixels occupied by the floc, and A2 is the total number of image pixels.
[0088] S504. Further determine the D of the floc using the calculated C and A of the floc f , and its correlation relationship is as follows:
[0089]
[0090] Take the logarithm and use the least squares method to perform a first-degree polynomial fitting to find the slope, and the slope is the fractal dimension D f .
[0091] Specifically, in step S5, the pre-trained convolutional neural network selects the Xception model and adds a scaling layer to adjust the floc images to meet the input matrix size, adopts early stopping to prevent overfitting and returns the best result after training, uses sparse cross-entropy as the loss function, constructs an Adam optimizer to find the parameters with the minimum loss during the pre-training process, and converts the output into probability values through the Softmax activation function.
[0092] S6. Using the multi-modal database containing floc images and water quality / process parameters constructed in step S4, and using the classification annotations in step S3 as supervision labels, train the established multi-modal deep learning model to obtain a trained multi-modal deep learning model.
[0093] S7. Use an underwater camera and integrated sensors to collect multi-modal data of floc images, influent water quality and process parameters, use the trained multi-modal deep learning model to determine whether the effluent water quality meets the standard, integrate the Grad-CAM++ visualization model to focus on the area, and locate the key floc areas affecting the decision-making.
[0094] The prediction performance of the training model is evaluated using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic (ROC) curve (AUC), as Figures 5 - 6 shown. The calibration of the model is evaluated by comparing the observed probability and the predicted probability, as Figure 6 shown. Decision curve analysis is used to evaluate the application value of the proposed multi-modal model in the decision-making of effluent water quality, as Figure 7 shown. The proposed method for floc image processing and recognition based on multi-modal deep learning shows strong prediction performance, with accuracy, precision, recall, F1-score, and AUC all exceeding 0.99, effectively avoiding false negatives and false positives. The calibration curve shows that the multi-modal deep learning model proposed in the present invention has good calibration effect on the effluent water quality probability spectrum. The decision curve shows that the method proposed in the present invention has a strong positive net benefit in actual decision-making, and the threshold range of the net benefit is very wide. The above results can be attributed to the effective data processing, rich multi-modal feature representation, and robust model structure in the present invention.
[0095] The contribution of each component to the model performance is analyzed by ablation study of the structure of the multi-modal deep learning model proposed in the present invention, as Figure 8As shown. Without adaptive data augmentation, the performance of the multimodal model drops sharply. Since the model pays more attention to the dominant categories during training, this can lead to overfitting and local optima. When there is no autoencoder, the model performance decreases slightly, indicating that the autoencoder plays an important role in early data fusion and model performance improvement. Because the non-linear transformation ability of the autoencoder allows significant features to be found, from which the input can be reconstructed into a high-level representation, making the original features better adapted to the model's understanding. The performance of the multimodal deep learning model without the Xception branch also decreases. Since the proposed hybrid fusion strategy takes advantage of the merits of early and late fusion and overcomes the classification errors associated with a single fusion mode. Therefore, the multimodal model proposed by the present invention can exhibit near-perfect classifier performance and accurately infer the effluent water quality before the sedimentation unit in order to achieve predictive feedback.
[0096] The above preferred embodiments are only used to illustrate the technical solutions and their effects of the present invention, rather than to limit the present invention. Any person skilled in the art can make modifications and changes to the above embodiments in form and details without departing from the spirit and scope of the present invention. Therefore, the scope of the right to protection of the present invention shall be as listed in the claims.
Claims
1. A floc image recognition method based on a multi-modal deep learning model, characterized in that, It includes the following steps: S1. Multimodal data collection: An underwater camera is used to collect floc images between the coagulation unit and the sedimentation unit, and an integrated sensor is used to monitor the influent water quality and process parameters in real time; S2. Floc image processing: Eliminate the noise in the underwater imaging of flocs, segment all pixels in the floc image into a series of independent and unconnected pixel regions, and improve the clarity of the floc image; S3. Data classification, annotation and partitioning: Classify the floc images and their corresponding water quality / process multimodal data according to the effluent turbidity, and store the floc images in a local folder according to the file directory structure for automatic partitioning and annotation; S4. Adaptive data augmentation: Perform adaptive data augmentation on the established floc image database for training the multimodal deep learning model, dynamically generate synthetic data for minority class samples, and solve the problem of data imbalance; S5. Dual-branch modeling based on a hybrid fusion strategy: Establish a dual-branch multimodal deep learning model based on a hybrid fusion strategy. The floc morphology recognition branch performs early fusion on the floc image features, influent water quality and process parameters, and the end-to-end recognition branch directly predicts the water quality classification probability using a pre-trained convolutional neural network. The weighted decision fusion method is used to perform late fusion on the probability predictions of the two branches; S6. Multimodal deep learning model training: Use the multimodal database constructed in step S4 that contains floc images and water quality / process parameters, and use the classification annotation in step S3 as the supervision label to train the established multimodal deep learning model to obtain a trained multimodal deep learning model; S7. Integrate Grad-CAM++ to visualize the model decision: Use an underwater camera and an integrated sensor to collect multimodal data of floc images, influent water quality and process parameters, use the trained multimodal deep learning model to judge whether the effluent water quality meets the standard, and integrate the Grad-CAM++ visualization model's region of interest to locate the key floc regions that affect the decision.
2. The floc image recognition method based on a multi-modal deep learning model according to claim 1, characterized in that Step S2 further includes: S201. Convert the original RGB image to a grayscale image using the weighted average method to reduce the computational amount and speed up the processing speed; the expression of the weighted average method is as follows: Grey = 0.299×R + 0.587×G + 0.114×B In the formula, 0.299, 0.587, and 0.114 are the standard coefficients defined by the International Telecommunication Union; S202. Fix the pixel values of the grayscale image to 0 or 255 using threshold segmentation to distinguish the effluent background features and floc features; the expression of the threshold segmentation is as follows: Among them, f(i,j) is the image after threshold processing, I(i,j) is the part composed of bright objects on the background in the grayscale image, and T is the selected threshold; all pixel points that satisfy the condition I≥T are recorded as target points to form the floc image, and the remaining points are recorded as background points to form the background image; S203. Use median filtering to remove salt-and-pepper noise in the floc image; S204. Use opening operation to remove small objects in the floc image, smooth the floc contour, disconnect narrow connections without significantly changing its area, and eliminate the protruding parts on the floc surface, which helps to remove noise and isolated small flocs and makes the main floc structure clearer.
3. The method for identifying floc images based on a multi-modal deep learning model according to claim 2, wherein In step S203, a 3×3 filter kernel is used to better retain the details of the image while removing noise.
4. A floc image recognition method based on a multi-modal deep learning model according to claim 1, characterized in that, In step S3, the multi-modal data is divided into four categories. Samples with an effluent turbidity less than 0.15 NTU are labeled as "Class A", samples with an effluent turbidity of 0.15 to 0.3 NTU are labeled as "Class B", samples with a turbidity of 0.3 to 1.0 NTU are labeled as "Class C", and 1 NTU is the drinking water standard limit in China; samples with a turbidity greater than 1.0 NTU are labeled as "Class D", which do not meet the drinking water quality standard.
5. A floc image recognition method based on a multi-modal deep learning model according to claim 1, characterized in that, In step S3, the file directory structure is as follows: the main directory is the floc image folder, the secondary directory is the training set or the test set, and the tertiary directory is the effluent classification label.
6. The method for identifying floc images based on a multimodal deep learning model according to claim 1, wherein In step S3, in the division ratio, the data volume in the training set is 80% of all the data, and the data volume in the test set is 20% of all the data.
7. A floc image recognition method based on a multi-modal deep learning model according to claim 1, characterized in that, In step S4, the adaptive data augmentation further includes the following steps: S401. Identify the most numerous class. S402. Randomly perform five geometric transformation operations on the minority class pictures, including 90° rotation, 180° rotation, 270° rotation, horizontal flipping, or vertical flipping, and record the operations performed on each picture. S403. If the number of the minority class is still less than that of the most numerous class, continue to perform oversampling and repeat the geometric transformation that is not recorded in the operation in S402. S404. Repeat steps S402 and S403 until the number of the minority class pictures is not less than that of the most numerous class. S405. If the number of the minority class pictures is greater than that of the most numerous class, perform downsampling on the generated pictures to remove the redundant pictures until the number of the minority class pictures is the same as that of the most numerous class.
8. A floc image recognition method based on a multi-modal deep learning model according to claim 1, characterized in that In step S5, the floc morphology recognition branch uses a feature extractor based on the pixels of the floc image to obtain the floc morphology features, which further includes the following steps: S501. Determine the connected components of the floc body, find out the mutually independent connected components in the image and regard them as connected units, and determine the floc position and the number of flocs. S502. Count the number of pixels on the periphery of each connected component to obtain the perimeter C of each floc, and count the number of pixels inside the connected component to obtain the floc area A. The calculation formula for the equivalent diameter d of the floc is as follows: S503. Define the floc image density by counting the ratio of the number of pixels occupied by the floc particles in the image to all the pixels of the actual image. This characteristic parameter can reflect the proportion of the floc image in the water background in the image, and its calculation formula is as follows: In the formula, p is the floc image density, A1 is the number of pixels occupied by the floc, and A2 is the total number of pixels of the image. S504. Further determine D of the flocs using the calculated C and A of the flocs f , and the associated relational expression is as follows: Take the logarithm and use the least squares method to perform a first-order polynomial fitting to find the slope, and this slope is the fractal dimension D f .
9. A floc image recognition method based on a multi-modal deep learning model according to claim 1, characterized in that In step S5, the features for early fusion of the floc image include but are not limited to the number of flocs, the average equivalent diameter, the fractal dimension, and the image density; in step S5, a pre-trained convolutional neural network is used in the end-to-end recognition branch and fine-tuned in the late fusion.
10. A floc image recognition method based on a multi-modal deep learning model according to claim 8, characterized in that, In step S501, the connected region of the floc body is determined using 8-neighboring pixels.
Citation Information
Patent Citations
Method and device for analyzing alumen ustum characteristics by using image recognition technology
CN113688751A
Colloidal pattern image recognition and dosing control system
CN117115438A