Micro-video-based activated sludge condition assessment method, device, and medium
By using metal sulfide technology, the technical problems that could not be effectively solved in the prior art have been solved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-16
Smart Images

Figure CN121904757B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent detection technology for wastewater treatment, and in particular to a method, equipment and medium for assessing the state of activated sludge based on microscopic video. Background Technology
[0002] Currently, activated sludge plays an important role in urban sewage and industrial wastewater treatment processes, and its condition directly affects the treatment effect. Microscopic examination is an important method for evaluating the condition of activated sludge.
[0003] In related technologies, single-frame microscopic images are usually used as the basic analytical unit to analyze activated sludge. However, activated sludge has significant spatial heterogeneity in the same reactor or biological tank, and the analysis scheme of single-frame microscopic images only reflects the microscopic structural state under local field of view, which is difficult to represent the overall characteristics of the sample. As a result, the activated sludge assessment results are difficult to reflect the actual state of the overall sample. Summary of the Invention
[0004] The main objective of this application is to provide a method, device, and medium for assessing the state of activated sludge based on microscopic video, aiming to solve the technical problem that activated sludge assessment results cannot reflect the actual state of the overall sample due to the analysis of activated sludge using single-frame microscopic images.
[0005] To achieve the above objectives, this application provides a method for assessing the state of activated sludge based on microscopic video. The method includes:
[0006] Control the sludge sample to move in a preset manner and acquire a microscopic video stream of the sludge sample;
[0007] Based on a preset frame extraction method, multiple frames of microscopic video images corresponding to the microscopic video stream are obtained.
[0008] Based on the feature vectors of each frame of microscopic video image, the comprehensive features of the sludge sample are obtained by graph learning.
[0009] The comprehensive features of the samples are input into a pre-trained evaluation model, and the state parameters of the sludge samples are determined based on the model output of the evaluation model.
[0010] In one embodiment, the step of acquiring multiple frames of microscopic video images corresponding to a microscopic video stream based on a preset frame extraction method includes:
[0011] According to a set time interval or a set frame interval, using the first frame of the microscopic video stream as the first frame microscopic video image, multiple frames of microscopic video images corresponding to the microscopic video stream are acquired; or...
[0012] Based on the frame extraction method corresponding to the shape of the sludge sample, multiple frames of microscopic video images corresponding to the microscopic video stream are obtained.
[0013] In one embodiment, the step of obtaining the comprehensive features of the sludge sample based on the feature vectors of each frame of microscopic video image using graph learning includes:
[0014] Determine the feature vectors of each frame of the microscopic video image. The feature vectors characterize the morphological features, texture features, grayscale and color distribution features of the sludge in that frame and / or the high-order structural features obtained by implicit learning.
[0015] By aggregating the feature vectors of sludge samples, a feature set is obtained.
[0016] Based on the graph structure construction method corresponding to the preset method, a graph structure is constructed based on the feature set;
[0017] The fused node features obtained by fusing features from multiple frames based on graph neural networks are used to obtain the comprehensive features of the sludge sample by using a graph-level readout strategy.
[0018] In one embodiment, the step of constructing a graph structure based on a feature set according to a graph structure construction method corresponding to a preset method includes:
[0019] If the preset method is continuous field-of-view acquisition, each frame of microscopic video image is used as a node, and the feature vector corresponding to the node is used as the node feature.
[0020] Based on the frame order indicated by the feature set, the edges between each node are determined to construct a fully connected graph as the graph structure.
[0021] In one embodiment, the step of constructing a graph structure based on a feature set according to a graph structure construction method corresponding to a preset method includes:
[0022] If the preset method is feature point acquisition, each frame of microscopic video image is used as a node, and the feature vector corresponding to the node is used as the node feature.
[0023] Based on the relative relationships between the feature points of each frame of microscopic video image, the edges between each node are determined to construct an adjacency graph as the graph structure.
[0024] In one embodiment, before the step of inputting the comprehensive features of the samples into a pre-trained evaluation model and determining the state parameters of the sludge samples based on the model output of the evaluation model, the following steps are included:
[0025] Multiple training sludge samples are controlled to move in a preset manner, and training microscopic video streams for the training sludge samples are acquired.
[0026] Obtain multidimensional analytical indicators from training sludge samples;
[0027] The training microscopic video streams of each training sludge sample are associated with multidimensional analysis indicators to form the training samples of the training sludge samples.
[0028] The original model is trained based on the training samples to obtain the evaluation model.
[0029] In one embodiment, the step of training the original model based on training samples to obtain the evaluation model includes:
[0030] For each training sample, based on a preset frame extraction method, multiple frames of training microscopic video images corresponding to the training microscopic video stream are obtained.
[0031] The feature vectors of each frame of training microscopic video image are aggregated into the comprehensive features of the training samples corresponding to the training sludge samples, thus obtaining the comprehensive features of the training samples corresponding to each training sample.
[0032] The comprehensive features of each training sample are used as input to the original model, and the multidimensional analysis indicators corresponding to each training sample are used as supervision signals to train and obtain the evaluation model.
[0033] In one embodiment, the step of training an evaluation model by using the comprehensive features of each training sample as input to the original model and the multidimensional analysis indicators corresponding to each training sample as supervision signals includes:
[0034] The comprehensive features of the training samples are used as input to the original model to obtain the predicted value.
[0035] Using the multidimensional analysis indicators corresponding to the training samples as supervision signals, the prediction bias of the predicted values is obtained;
[0036] The model parameters of the original model are optimized in reverse based on the prediction bias, and the steps of taking the comprehensive features of the training samples corresponding to the training samples as input to the original model are executed based on the next training sample to obtain the predicted value, so as to iteratively obtain the evaluation model.
[0037] In addition, to achieve the above objectives, this application also provides a microscopic video-based activated sludge state assessment device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the microscopic video-based activated sludge state assessment method described above.
[0038] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps of the activated sludge state assessment method based on microscopic video described above.
[0039] This application provides a method for assessing the state of activated sludge based on microscopic video. The method involves controlling the movement of a sludge sample according to a preset pattern and acquiring a microscopic video stream of the sludge sample. Based on a preset frame extraction method, multiple frames of microscopic video images corresponding to the video stream are acquired. The feature vectors of each frame of the microscopic video image are aggregated to form a comprehensive feature corresponding to the sludge sample. This comprehensive feature is then input into a pre-trained assessment model, and the state parameters of the sludge sample are determined based on the model's output. By acquiring microscopic video streams covering different random fields of view of the sludge sample, uniformly extracting frames to retain effective visual information across the entire field, and aggregating feature vectors from multiple frames to form a comprehensive feature representing the overall state, this method inputs the comprehensive feature into a pre-trained model to achieve accurate mapping from global features to overall state parameters. This process avoids interference from local features throughout, enabling quantitative assessment of the overall state of activated sludge and overcoming the limitations of single-frame image representation. Attached Figure Description
[0040] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating Example 1 of the activated sludge state assessment method based on microscopic video provided in this application.
[0043] Figure 2 This is a flowchart illustrating Example 6 of the activated sludge state assessment method based on microscopic video provided in this application.
[0044] Figure 3 A schematic diagram of the process for the activated sludge condition assessment method based on microscopic video provided in this application;
[0045] Figure 4 This is a schematic diagram of the activated sludge condition assessment device based on microscopic video, as described in an embodiment of this application.
[0046] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0047] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0048] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0049] Currently, microscopic examination is an important method for assessing the state of activated sludge. Existing techniques typically use single-frame microscopic images or a small number of field-of-view images as the basic analytical unit, reflecting only the microstructural state under local fields of view, and are difficult to represent the overall characteristics of the entire sample. To improve the accuracy of single-frame prediction, images often need to be manually or algorithmically screened to remove blank, blurred, and occluded frames. This not only increases manual or computational costs but also easily introduces related statistical biases, and the model results are highly dependent on the screening strategy. Multi-frame microscopic information is not jointly modeled. Even with high-quality microscopic video acquisition capabilities, existing methods usually still treat each frame of the extracted video as an independent sample, performing feature extraction and prediction separately, and then summarizing the results only through simple statistical methods. This fails to treat multiple frames of microscopic images obtained under the same sampling conditions as a whole sample for modeling, making it impossible to form a comprehensive representation of a single sample and to fully utilize the complementary information contained in different fields of view.
[0050] The main solution of this application is as follows: Controlling the sludge sample to move according to a preset method and acquiring a microscopic video stream of the sludge sample; acquiring multiple frames of microscopic video images corresponding to the microscopic video stream based on a preset frame extraction method; aggregating the feature vectors of each frame of microscopic video images into a comprehensive sample feature corresponding to the sludge sample; inputting the comprehensive sample feature into a pre-trained evaluation model, and determining the state parameters of the sludge sample based on the model output of the evaluation model. By acquiring microscopic video streams covering different random fields of view of the sludge sample, uniformly extracting frames to retain effective visual information across the entire domain, aggregating the feature vectors of multiple frames to form a comprehensive sample feature representing the overall state, and then inputting it into a pre-trained model to complete the accurate mapping from the global feature to the overall state parameters, the entire process avoids interference from local features, achieving the technical effect of improving the representativeness of the evaluation results to the overall state.
[0051] It should be noted that the executing entity in this embodiment can be an activated sludge condition assessment device based on microscopic video, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an activated sludge condition assessment device based on microscopic video capable of performing the above functions. This embodiment does not specifically limit it in this way. The following uses an activated sludge condition assessment device based on microscopic video as an example to describe this embodiment and the following embodiments.
[0052] Based on this, Embodiment 1 of this application proposes a method for assessing the state of activated sludge based on microscopic video. Please refer to... Figure 1 , Figure 1This is a flowchart illustrating an embodiment of the activated sludge state assessment method based on microscopic video of this application. The activated sludge state assessment method based on microscopic video includes steps S10-S40:
[0053] Step S10: Control the sludge sample to move in a preset manner and acquire a microscopic video stream of the sludge sample.
[0054] Sludge samples were collected from the biological treatment tank. In wastewater treatment plants, pollutant degradation mainly occurs in the biological treatment tank, and the state of the activated sludge in the tank directly affects the wastewater treatment effect. The activated sludge in the biological treatment tank can be divided into three stages: aerobic, anoxic, and anaerobic. Activated sludge mixtures collected at the same sampling time and under the same operating conditions are primarily composed of microbial communities, organic and inorganic suspended particles, and water. They are the same sample for microscopic observation and the determination of indicators such as MLSS (Mixed Liquor Suspended Solids), SVI (Sludge Volume Index), viscosity, SOUR (Specific Oxygen Uptake Rate), and KLa (Oxygen Transfer Coefficient). It is essential to ensure consistency between the microscopic observations and indicator measurements after sampling to avoid detection bias. Microscopic video streams are dynamic video data sequences obtained by continuously imaging moving sludge samples through a microscope. They include imaging dimensions, which are the microscopic structures of the sludge samples captured by the magnified field of view of the microscope; image correspondence, which means that each frame of the video corresponds to different random fields of view of the same sludge sample during its movement; and data correlation, which means that it is bound to the sludge sample evaluation indicators one by one, together forming a complete sample.
[0055] In this embodiment, activated sludge samples are obtained from the biological tank under the same sampling time and operating conditions, and observed under a microscope. The sludge samples are moved according to a preset method, and a microscopic video stream of a certain duration is acquired through continuous imaging to characterize the overall microscopic structure of the activated sludge under the sampling conditions. Each frame of the video corresponds to a different random field of view under the same sample. While the microscopic video stream is being acquired, evaluation indicators are measured on the same activated sludge sample to obtain one or more evaluation indicator values such as MLSS, SVI, viscosity, SOUR, and KLa, which are stored in a one-to-one correspondence with the corresponding microscopic video stream samples. Each segment of the microscopic video stream and its corresponding evaluation indicator are defined as a complete sample, establishing a one-to-one correspondence data structure between microscopic video streams and evaluation indicators. By controlling the sludge samples to move according to a preset method and synchronously acquiring microscopic video streams, continuous visualization capture of different random fields of view of the activated sludge is achieved, fully characterizing the overall microscopic structure of the sludge under the sampling conditions.
[0056] As one implementation method, the preset movement method can include tracked movement, random movement, and production line movement. There can be multiple sludge samples. For multiple samples, the same preset movement method can be used for different samples, or different preset movement methods can be used. For example, tracked movement can be used for one sample, and random movement can be used for another sample.
[0057] For example, for a single sludge sample, a preset movement mode can be selected as needed to execute single-sample movement control, simultaneously activating microscope imaging to acquire an independent microscopic video stream corresponding to that sample. Sample movement modes can include: tracked movement, where the sample moves at a uniform speed and continuously linearly with the stage track, achieving continuous field-of-view scanning; random movement, where the sample moves with the stage without a fixed trajectory, achieving random field-of-view acquisition, covering different random observation angles of the sample; and production line movement, where the sample moves along a fixed production line trajectory in a fixed sequence, achieving regularized and standardized field-of-view acquisition, obtaining a microscopic video stream for the sludge sample. Selecting a preset mode as needed for a single sludge sample and simultaneously acquiring the microscopic video stream allows for precise adaptation of the single-sample movement mode to the detection requirements, providing a high-quality microscopic visual data foundation for subsequent sludge condition assessment.
[0058] As another implementation method, sludge sample collection can include conveyor belt collection and sample box collection. For the testing needs of multiple sludge samples, the same collection method can be used for different samples, or different collection methods can be used. Based on the testing requirements of each sample, the collection method is determined, and each sample is sequentially controlled to complete its corresponding movement operation on the microscope stage, while the microscopic video stream of each sample is simultaneously acquired, thus completing the movement control and microscopic video stream acquisition of all samples.
[0059] For example, in the conveyor belt acquisition method, activated sludge suspension is dripped onto the conveyor belt, and the microscopic video of the activated sludge is captured as the conveyor belt moves, while the microscope remains stationary. In the sample box method, a certain volume of activated sludge suspension is placed in a sample box, which has a thickness. Microscopic video of specific points within the sample box is captured by moving the sample box and the microscope. For the conveyor belt acquisition method, continuous dynamic acquisition is achieved while the microscope remains stationary through conveyor belt movement; for the sample box method, microscopic video of specific points is acquired by moving the sample box and the microscope, achieving precise coverage of points in suspensions with thickness. These two acquisition methods are adapted to the microscopic video capture needs of sludge samples in different scenarios, ensuring the stability of the acquisition process and the accuracy of the acquisition points.
[0060] The above are only two feasible implementations of step S10 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S10.
[0061] Step S20: Based on a preset frame extraction method, acquire multiple frames of microscopic video images corresponding to the microscopic video stream.
[0062] Multi-frame microscopic video images are a collection of independent static microscopic images extracted from a microscopic video stream acquired from sludge samples using a preset frame extraction method. Each image is a single frame from a specific moment in the microscopic video stream and reflects the microscopic field-of-view characteristics of the sludge sample under the microscope at that moment.
[0063] In this embodiment, the microscopic video is periodically sampled according to a preset frame extraction method, extracting several frames of microscopic images from each video segment to obtain corresponding multi-frame microscopic video images. By extracting multi-frame microscopic video images from the microscopic video stream using the preset frame extraction method, effective microscopic visual data is efficiently screened, providing static image samples that can be quantified and analyzed.
[0064] As one implementation method, the preset frame extraction method may include performing frame extraction according to a preset time interval.
[0065] For example, a microscopic video stream is acquired, and frame extraction parameters are preset at equal time intervals, including the single frame extraction time interval, frame extraction start time, and total frame extraction duration. This ensures that the parameter settings match the acquisition duration of different microscopic video streams, guaranteeing a consistent number of images after frame extraction for all samples. An automatic frame extraction command is initiated, and the system performs timed and fixed-point frame extraction on the microscopic video stream according to preset time intervals. Starting from the start time of the microscopic video stream, a single frame of microscopic video image is extracted at each preset time interval node. Finally, multiple frames of microscopic video images corresponding to the microscopic video stream are acquired. Automated frame extraction of the microscopic video stream is completed through preset equal time intervals, accurately matching the temporal dimension characteristics of video acquisition. This effectively captures the random state of the sludge microstructure in the temporal dimension, ensuring the effectiveness of the extracted images and the consistency of data scale among different samples.
[0066] As another implementation method, the preset frame dropping method may include dropping frames at preset frame intervals.
[0067] For example, the total number of frames in each segment of the microscopic video stream is obtained. Based on a preset fixed total number of frames to be extracted, the uniform frame interval corresponding to each segment of the microscopic video stream is calculated to ensure that the number of images extracted from microscopic video streams of different durations and total frame numbers is consistent. The frame extraction program is started, and uniform frame extraction is performed on each segment of the microscopic video stream according to the calculated frame interval. Starting from the first frame of the microscopic video stream, one microscopic video image is extracted every preset number of frames. After the frame extraction is completed, the validity of the extracted images is checked, and a small number of blurry or abnormal frames without effective sludge particles are removed. At the same time, effective frames with the same frame interval are extracted from the corresponding microscopic video stream to ensure that the number of extracted images for each sample remains consistent. Calculating the uniform frame interval based on the total number of video frames and completing the frame extraction adapts to microscopic video streams of different acquisition durations and total frame numbers. This maximizes the preservation of the spatial random characteristics of the sludge microstructure. At the same time, the frame extraction and supplementation method can further remove abnormal frames and supplement effective frames, improving the overall quality of the extracted images while ensuring that the number of extracted images for different samples is consistent.
[0068] As another implementation method, the preset frame extraction method may include performing frame extraction based on a preset field of view overlap rate threshold.
[0069] For example, field-of-view region matching is performed on the microscopic images of consecutive frames in the microscopic video stream. The field-of-view overlap rate between adjacent frames is calculated and compared with a preset field-of-view overlap rate threshold. If the field-of-view overlap rate of adjacent frames is higher than the threshold, the next frame is discarded; if the field-of-view overlap rate of adjacent frames is lower than or equal to the threshold, the next frame is retained as the extracted frame image. This rule is followed, starting from the first frame of the microscopic video stream, to sequentially traverse all consecutive frames until the entire microscopic video stream is traversed. The final retained images are the extracted multi-frame microscopic video images. The preset field-of-view overlap rate threshold can be flexibly adjusted according to the detection requirements of sludge samples. A low threshold can be set for scenarios requiring large field-of-view coverage to reduce field-of-view duplication, while a high threshold can be set for scenarios requiring detailed acquisition of key areas to retain more local detail frames. Frame extraction based on the preset field-of-view overlap rate threshold, by filtering out redundant frames with high overlap rates and retaining effective frames with low overlap rates, achieves low-redundancy full-coverage acquisition of the sludge microscopic field of view. The threshold can be flexibly adjusted according to detection requirements to adapt to different scenarios of large field-of-view coverage or detailed acquisition of key areas.
[0070] Step S30: Based on the feature vectors of each frame of microscopic video image, obtain the comprehensive features of the sludge sample corresponding to the sample using graph learning.
[0071] The comprehensive feature of the sample is a single high-dimensional feature vector formed by the multi-frame microscopic video image features of the sludge sample after feature extraction, set construction, graph learning and multi-frame feature fusion, which can comprehensively characterize the overall state of the sludge sample.
[0072] In this embodiment, feature extraction is performed on each extracted frame of microscopic image, converting the image into a high-dimensional feature vector to characterize the morphological features, texture features, grayscale or color distribution features, and high-order structural features obtained through implicit learning of the activated sludge in that frame. The feature vectors corresponding to all frames in the same microscopic video sample are aggregated to form a multi-frame feature set for that sample. Each frame's feature vector is considered a constituent unit, constructing a feature set for joint modeling. Based on this feature set, the multiple frames of microscopic images extracted from the same microscopic video sample are constructed as graph-structured data, and a graph neural network (GNN) is used to jointly model the features of the multiple frames. After joint modeling, the features of the multiple frames are aggregated into a single comprehensive feature for the sample. This achieves multi-view fusion and high-order association learning of sludge features, effectively integrating single-frame local features, and thus comprehensively characterizing the unified features of the overall state of the sludge sample.
[0073] As one implementation method, feature extraction is performed using at least one of CNN and automated image omics methods to obtain feature vectors.
[0074] For example, for each extracted frame of microscopic video image, features are extracted using CNN and automated imagemics methods respectively, resulting in two types of high-dimensional feature vectors. These two types of feature vectors are then concatenated and fused to form a fused single-frame high-dimensional feature vector. All frame fused feature vectors corresponding to the same sludge sample are collected to form a multi-frame fused feature set for that sample. Based on this multi-frame fused feature set, the multi-frame microscopic video images of the sample are constructed as graph-structured data. A graph neural network is used to jointly model the multi-frame fused features. Through feature normalization and aggregation operations, the multi-frame fused features are aggregated into a single comprehensive feature representation of the sample, obtaining the comprehensive feature of the sludge sample. By extracting and concatenating single-frame features using CNN and automated imagemics methods, and then aggregating them into comprehensive features after joint modeling of the multi-frame fused features using a graph neural network, multi-dimensional complementary extraction of single-frame features of sludge is achieved, comprehensively capturing both shallow visual features such as sludge morphology and texture, as well as high-order structural features.
[0075] As another implementation method, the feature vectors of each frame are initially aggregated by temporal fusion or spatial pooling methods to generate an intermediate feature representation.
[0076] For example, average pooling or max pooling is performed along the temporal dimension on the feature vectors of all frames corresponding to the sludge sample to obtain a global feature vector. Alternatively, recurrent neural networks, such as LSTM, GRU, or self-attention mechanisms, can be used to model the temporal dependencies between frames and output a fused feature sequence. Then, based on the graph structure corresponding to the preset method, the comprehensive features of the sludge sample are further constructed. Directly aggregating temporal features through pooling operations can efficiently compress the feature dimensions of multiple frames and quickly obtain a global feature vector that can characterize the global distribution of the temporal features of the sludge sample. Furthermore, using recurrent neural networks or self-attention mechanisms to model the temporal dependencies between frames can accurately uncover the temporal correlation patterns of features across multiple frames and capture the dynamic changing trends of the microscopic features of the sludge.
[0077] Step S40: Input the comprehensive features of the sample into the pre-trained evaluation model, and determine the state parameters of the sludge sample based on the model output of the evaluation model.
[0078] In this embodiment, the comprehensive features of the sludge sample are input into the evaluation model, the model inference calculation is initiated, and the feature analysis results output by the model are obtained. Based on the model's preset output mapping rules, the analysis results are converted into the corresponding state parameters of the sludge sample, completing the sludge state assessment. This achieves accurate mapping of the vectorized state parameters of the sludge's microscopic visual features, and efficiently completes the sludge state assessment based on the learning rules of the model's pre-training.
[0079] As one implementation method, the evaluation model includes regression or classification models.
[0080] For example, if the state parameters of the sludge sample to be determined are continuous numerical indicators, a regression model is selected as the evaluation model. The comprehensive characteristics of the sludge sample are input into the pre-trained regression model. Based on the learned mapping rules between the comprehensive characteristics of the sample and continuous state parameters, the regression model outputs a continuous numerical result matching the sample. This result is the corresponding state parameter of the sludge sample, such as specific values for MLSS, SVI, SOUR, viscosity, KLa, etc. If the state parameters of the sludge sample to be determined are grade or category indicators, a classification model is selected as the evaluation model. The comprehensive characteristics of the sludge sample are input into the pre-trained classification model. Based on the learned mapping rules between the comprehensive characteristics of the sample and category state parameters, the classification model outputs a category or grade determination result matching the sample. This result is the corresponding state parameter of the sludge sample, such as high, medium, or low sludge activity level. By selecting the pre-trained regression or classification model according to the type of sludge state parameters required, targeted and accurate prediction of continuous numerical indicators and grade or category indicators is achieved, adapting to the detection needs of different sludge state assessments.
[0081] As another implementation method, the model output may include predicted values of state parameters of sludge samples and confidence scores of each parameter.
[0082] For example, the comprehensive features of the samples are input into the evaluation model to obtain stratified model output results with confidence levels, such as predicted values of core state parameters and confidence scores for each parameter. The stratified output results are then validated using preset confidence level filtering rules, removing predicted values that do not meet the confidence level criteria. The model output results are then converted into precise sludge state parameters, generating a complete sludge state assessment report with parameter confidence levels. By incorporating confidence level filtering, the quality of the input data to the evaluation model is effectively improved, the prediction bias of sludge state parameters is significantly reduced, and sludge state assessment becomes more accurate.
[0083] This embodiment provides a method for assessing the state of activated sludge based on microscopic video. First, a sludge sample is moved according to a preset method, and a microscopic video stream of the sludge sample is acquired. Based on a preset frame extraction method, multiple frames of microscopic video images corresponding to the video stream are acquired. The feature vectors of each frame of the microscopic video image are aggregated to form a comprehensive sample feature corresponding to the sludge sample. The comprehensive sample feature is input into a pre-trained assessment model, and the state parameters of the sludge sample are determined based on the model output. By acquiring microscopic video streams covering different random fields of view of the sludge sample, uniformly extracting frames to retain effective visual information across the entire field, and aggregating the feature vectors of multiple frames to form a comprehensive sample feature representing the overall state, this feature is then input into a pre-trained model to complete the accurate mapping from global features to overall state parameters. This process avoids interference from local features throughout, enabling quantitative assessment of the overall state of activated sludge and avoiding the limitations of single-frame image representation.
[0084] Based on Embodiment 1, in Embodiment 2 of this application, the content that is the same as or similar to that in Embodiment 1 can be referred to the above description and will not be repeated hereafter. Based on this, the step of obtaining multiple frames of microscopic video images corresponding to the microscopic video stream based on a preset frame extraction method includes step S21 or step S22:
[0085] Step S21: According to a set time interval or a set frame interval, take the first frame of the microscopic video stream as the first frame microscopic video image, and obtain multiple frames of microscopic video images corresponding to the microscopic video stream.
[0086] As one implementation method, a time interval or a frame interval is set, and the first frame of the microscopic video stream is used as the first frame microscopic video image. Subsequent frames are obtained according to a preset number of frame extractions at the set time interval or the set frame interval to obtain multiple frames of microscopic video images corresponding to the microscopic video stream.
[0087] Specifically, the first frame of the microscopic video stream is located and used as the first microscopic video image. Subsequent images are extracted sequentially from the microscopic video stream at preset time intervals, such as one frame every 2 seconds, or at frame intervals, such as one frame every 10 frames, until the preset number of frames is reached, resulting in multiple frames of microscopic video images. Image extraction is achieved through frame extraction according to fixed rules, ensuring the consistency of the frame extraction logic across different sludge samples and providing stable and representative image data for subsequent feature extraction.
[0088] As another implementation method, a set time interval or a set frame number interval is obtained, the first frame and the last frame are obtained, the total number of frames of the microscopic video stream is obtained, and the frame images are evenly extracted from the first frame and the last frame according to the set time interval or the set frame number interval. If the calculated number of extracted frames is not an integer, it is rounded up or down to determine the final number of extracted frames.
[0089] Specifically, the first and last frames of the microscopic video stream are located, and the total number of frames in the stream is calculated. The theoretical number of frames to be extracted is calculated based on a preset time interval or frame interval. If the theoretical number of frames to be extracted is not an integer, it is rounded up or down to an integer according to a preset rule. Then, extraction is performed evenly starting from the first frame until extraction is complete, resulting in multiple frames of microscopic video images. This method uses the first and last frames as boundaries to achieve even coverage of the entire microscopic video stream, avoiding sampling deviations caused by differences in features between the beginning and end segments of the stream, and maximizing the representativeness of the extracted frames for the entire microscopic video stream.
[0090] As another implementation method, a set time interval or a set frame number interval is obtained, and the nth frame of the microscopic video stream is used as the starting frame, where n is a positive integer greater than 1, and subsequent frames are obtained at the set time interval or the set frame number interval.
[0091] Specifically, the preset starting frame is frame n, where n is a positive integer greater than 1, such as frame 2 or frame 5. This starting frame is located and used as the first target microscopic video image. Then, following a preset time interval or frame interval, subsequent frames are extracted sequentially from the starting frame until the preset number of frames is reached or the last frame of the microscopic video stream is reached, resulting in multiple frames of microscopic video images. This avoids the problem of poor image quality in the first frame of the microscopic video stream due to issues such as device focusing or sample instability, ensuring that each extracted frame is high-quality and effective microscopic field-of-view data, thus improving the accuracy and effectiveness of subsequent feature extraction.
[0092] Step S22: Based on the frame extraction method corresponding to the shape of the sludge sample, obtain multi-frame microscopic video images corresponding to the microscopic video stream.
[0093] In this embodiment, image analysis is performed on the initial frames of the microscopic video stream to identify the coating shape and effective area boundary of the sludge sample on the glass slide. Based on the sample shape, such as circular, rectangular, or irregular, a targeted frame extraction strategy is designed. For example, for circular samples, radial interval frame extraction is used, that is, taking the geometric center of the circular sample coating area as the origin, dividing the radial sampling direction at preset angle intervals, and then extracting frames sequentially from the center to the edge in each direction at preset pixel or frame intervals. For irregular samples, layered frame extraction is performed along the edge and center regions, that is, first, the central and edge regions of the sample are divided by image recognition, then frames are extracted at a uniform number of frames or time intervals in the central region, and frames are extracted at a denser interval in the edge region. Simultaneously, the number of frames extracted from the center and edge is allocated according to a preset ratio to achieve full-area feature coverage of irregular samples without blind spots. Through adaptive frame extraction, the actual coating shape of the sludge sample is accurately matched, achieving full-area uniform coverage of the effective sample area and minimizing the loss of local features caused by the irregular shape of the sample.
[0094] In one feasible implementation, the long-duration microscopic video stream is divided into multiple equal-length stages, such as every 2 minutes. Within each stage, a preset number of images are extracted at fixed frame intervals to ensure the representativeness of features in each stage. The extracted frames from each stage are then aggregated to form a multi-frame microscopic video image set covering the entire video duration. This approach balances temporal uniformity with feature representativeness within each stage, avoiding the loss of local temporal features due to the excessive length of the video.
[0095] In one feasible implementation, the field-of-view overlap rate of adjacent frames is calculated based on the microscope's field of view size and the stage's movement step distance. A maximum allowable overlap rate threshold is set; when the overlap rate of adjacent frames exceeds the threshold, the current frame is automatically skipped, and the next frame is extracted. Extraction continues until a preset number of images is reached or the last frame is encountered, ensuring that the field-of-view overlap of the extracted images is within a reasonable range. This effectively reduces repeated sampling of the field of view, expands the effective field-of-view coverage of a single sample, and improves the information richness of the image data.
[0096] Based on any of the above embodiments of this application, Embodiment 3 of this application proposes a method for assessing the state of activated sludge based on microscopic video, which can be referred to the above description and will not be repeated hereafter. Based on this, the step of aggregating the feature vectors of each frame of microscopic video image into the comprehensive sample features corresponding to the sludge sample includes:
[0097] Step S31: Determine the feature vector of each frame of microscopic video image. The feature vector represents the morphological features, texture features, grayscale and color distribution features of the sludge in the frame and / or the high-order structural features obtained by implicit learning.
[0098] Feature vectors are high-dimensional numerical vectors formed by quantifying various visual features of sludge in a single frame of microscopic video image. They are extracted using CNN or automated imagemics methods. Each numerical dimension corresponds to a type of microscopic feature of the sludge, and multiple dimensions are combined to form a feature set that characterizes the microscopic state of the sludge in that frame of image. Morphological features are intuitive features that characterize the geometric shape and spatial distribution of sludge particles in that frame of microscopic image, such as flocs and filamentous bacteria. They mainly include quantitative indicators such as the size, shape, number, and arrangement of sludge particles. For example, the density of flocs and the length and number of branches of filamentous bacteria are all morphological features. Texture features are fine-grained features that characterize the surface texture of sludge particles and the distribution of texture between particles in that frame of microscopic image. They mainly include texture uniformity, roughness, contrast, texture direction, and the gray-level variation of local neighboring pixels. For example, the smooth texture of dense flocs, the porous texture of loose flocs, and the network texture formed by interwoven filamentous bacteria are all texture features used to distinguish the structural state of sludge particles. Grayscale features are fundamental visual characteristics used in sludge microscopic images to characterize the distribution and changes in pixel grayscale values. They primarily include quantitative indicators such as the mean, variance, median, grayscale histogram distribution, and grayscale entropy of pixel grayscale values. These reflect the differences in brightness between sludge particles and the background, and between different types of sludge particles within the microscopic field of view. For example, the grayscale difference between the floc region and the clear water background, and the difference in the mean grayscale value between activated sludge and aged sludge, are all reflected through grayscale features. Color distribution features are characteristics used in color microscopic sludge images to characterize the distribution, proportion, and combination patterns of pixels in each color channel. They primarily include the mean, variance, color proportion, color cluster center, and color uniformity of each color channel. For example, the light brown hue distribution of activated sludge flocs, the color deviation in the filamentous bacteria region during sludge expansion, and the color changes in localized oxidation areas are all quantitatively reflected through color distribution features. These are important supplementary characteristics to the sludge state under color imaging. The high-order structural features obtained through implicit learning are deep features that cannot be manually defined or quantified, automatically mined and extracted by deep learning models such as CNNs, beyond shallow basic features such as morphology and texture. These features do not require manual pre-setting of feature dimensions; they are implicitly generated by the model after learning from a large number of samples. For example, abstract patterns such as the symbiotic relationship between sludge flocs and filamentous bacteria, and the overall distribution patterns of active and inactive regions, can reflect the essential laws of sludge state that cannot be reflected by shallow features.
[0099] In this embodiment, for each frame of microscopic video image, feature extraction is performed using CNN or automated imagemics methods to learn the morphology, texture, grayscale, color distribution features, and / or higher-order structural features of sludge particles in the microscopic video image. The extracted features are encoded into high-dimensional vectors to form the feature vector corresponding to each frame. Through multimodal feature extraction, the visual information of a single frame of microscopic video image is transformed into a quantifiable high-dimensional feature vector, containing at least one interpretable explicit feature and implicitly learned higher-order structural features, providing basic data that accurately reflects the microscopic state of a single frame for subsequent aggregated modeling.
[0100] Step S32: Aggregate the feature vectors of the sludge samples to obtain the feature set.
[0101] In this embodiment, feature vectors from all frames of images for the same sludge sample are collected to form a multi-frame feature vector list for that sample. This feature vector list constitutes the feature set of the sludge sample. By aggregating feature vectors from multiple frames to form a feature set, local features from a single frame are integrated into a global feature data source for the sample, thus avoiding the limitations of single-frame features.
[0102] Step S33: Construct a graph structure based on the feature set according to the graph structure construction method corresponding to the preset method.
[0103] In this embodiment, different preset methods correspond to different graph structures, such as tracked movement, random movement, and production line movement, each corresponding to its own graph structure. The feature vector corresponding to each frame of image is considered as a constituent unit, constructing a feature set structure for joint modeling. Based on this feature set structure, multiple frames of microscopic images extracted from the same microscopic video sample are constructed as graph structure data, and a graph neural network is used to jointly model the features of multiple frames of images. That is, each frame of microscopic image is defined as a node in the graph, and its node features are composed of high-dimensional feature vectors obtained by feature extraction from that frame of image. Simultaneously, according to the acquisition method and sample organization of the microscopic video, the connection relationships between nodes are defined, thereby forming the input graph structure for graph neural network computation. The graph neural network can achieve information fusion and representation learning between features of multiple frames of images, enabling the model to comprehensively utilize complementary information contained in different fields of view to form a representative joint feature representation for a single microscopic video sample.
[0104] Step S34: The fused node features obtained by fusing multiple frames of features based on graph neural networks are used to obtain the comprehensive features of the sludge sample through a graph-level readout strategy.
[0105] In this embodiment, a graph neural network is used to perform information propagation and feature aggregation between nodes based on a graph structure. The graph neural network performs information fusion and representation learning between features from multiple frames of images through mechanisms such as message passing, neighbor feature aggregation, and node representation updates. Specifically, each node receives feature information from its neighboring nodes and aggregates it (e.g., by taking the mean or maximum value). Then, the aggregated neighbor information is combined with its own original features to generate a fused node feature containing local contextual relationships. Through a readout operation, the features from multiple frames are aggregated into a single sample comprehensive feature, used to characterize the overall activated sludge state corresponding to the microscopic video sample. This allows the model to comprehensively utilize complementary information contained in different fields of view to form a representative joint feature representation for a single microscopic video sample without relying on single-frame image filtering. Deep fusion of features from multiple frames is achieved through the graph neural network, and then the fused features are aggregated into a single sample comprehensive feature through a graph-level readout strategy, comprehensively characterizing the overall microscopic state of the sludge sample.
[0106] In this embodiment, by extracting multidimensional feature vectors from single-frame images, aggregating them to form a global feature set, and combining them with an adapted graph structure for GNN modeling, a comprehensive sample feature that can fully characterize the overall microscopic state of sludge is generated. This achieves dimensionality upgrade integration from local frame features to global sample features, providing a highly representative feature basis for subsequent accurate assessment of sludge state.
[0107] Based on any of the above embodiments of this application, Embodiment 4 of this application proposes a method for assessing the state of activated sludge based on microscopic video, which can be referred to the above description and will not be repeated hereafter. Based on this, the step of constructing the comprehensive sample characteristics corresponding to the sludge sample based on the feature set according to the graph structure corresponding to the preset method includes:
[0108] Step S331: If the preset method is continuous field-of-view acquisition, each frame of microscopic video image is used as a node, and the feature vector corresponding to the node is used as the node feature.
[0109] In this embodiment, when the microscopic video is obtained through tracked scanning or continuous field-of-view acquisition, each frame image originates from different random observation angles within the same sample. The preset method is continuous field-of-view acquisition, and there is no pre-defined spatial adjacency relationship between frames. For the microscopic video stream acquired through continuous field-of-view acquisition, each frame of the extracted microscopic video image is defined as an independent node in the graph structure. The high-dimensional feature vector of each frame is directly assigned as the node feature of the corresponding node, ensuring that the node feature is completely consistent with the dimension and data format of the original frame feature vector. After completing the node mapping and feature assignment for all frames, a set of nodes in the graph structure is formed. Directly mapping multiple frames of images and their features to the nodes and node features of the graph preserves the complete microscopic feature information of each frame image, providing a precise node data foundation for the subsequent construction of a fully connected graph.
[0110] Step S332: Based on the frame order indicated by the feature set, determine the edges between each node to construct a fully connected graph as the graph structure.
[0111] In this embodiment, based on the frame order indicated by the feature set and the defined node set, an edge connection is established between any two nodes. That is, for each node, it is connected to all other nodes in the node set except itself, forming a fully connected graph structure. The construction of the fully connected graph breaks the limitation of physical location between frames, allowing the model to fully integrate the microscopic features of different perspectives, and to more comprehensively and accurately represent the overall microscopic state of sludge.
[0112] In this embodiment, the conveyor belt acquisition method involves extracting activated sludge suspension and dropping it onto a conveyor belt. The belt moves with the sludge to capture microscopic video of the activated sludge, while the microscope remains stationary. Microscopic images and videos are acquired, and frames are periodically extracted to build a fully connected graph, completing graph learning and obtaining a comprehensive characterization. Alternatively, the activated sludge suspension can be dropped onto a slide placed on the conveyor belt. Through continuous field-of-view acquisition methods such as conveyor belt acquisition, dynamic capture of the activated sludge suspension and stable acquisition of microscopic video are achieved. For the microscopic video stream acquired through continuous field-of-view acquisition, by mapping each frame image to graph nodes, assigning feature vectors as node features, and constructing a fully connected graph, the physical position limitations between frames are broken, enabling the fusion learning of features from multiple frames under different fields of view. This effectively eliminates local field-of-view feature bias and can more comprehensively and accurately characterize the overall microscopic state of the sludge sample.
[0113] Based on any of the above embodiments of this application, Embodiment 5 of this application proposes a method for assessing the state of activated sludge based on microscopic video, which can be referred to the above description and will not be repeated hereafter. Based on this, the step of constructing the comprehensive sample characteristics corresponding to the sludge sample based on the feature set according to the graph structure corresponding to the preset method includes:
[0114] Step S333: If the preset method is feature point acquisition, each frame of microscopic video image is used as a node, and the feature vector corresponding to the node is used as the node feature.
[0115] In this embodiment, when the microscopic video is acquired through a sample box, fixed slide, or regular field of view, and each frame has a clear physical location or original acquisition adjacency relationship, the preset method is feature point acquisition. For the microscopic video stream acquired through feature point acquisition, each frame of the extracted microscopic video image is defined as an independent node in the graph structure. The high-dimensional feature vector of each frame is directly assigned to the corresponding node as node features. This achieves accurate mapping of microscopic video frames to graph structure nodes under feature point acquisition, fully preserving the morphology, texture, and other multi-dimensional features of each frame, laying a precise and unified node data foundation for the subsequent construction of an adjacency graph that conforms to the actual acquisition order of the samples.
[0116] Step S334: Based on the relative relationships between the feature points of each frame of microscopic video image, determine the edges between each node to construct an adjacency graph as a graph structure.
[0117] In this embodiment, the spatial relative positions, temporal order, or movement path relationships between each feature point are extracted to determine the sequential relationships between nodes in the graph structure. Edge connections are established only between nodes in adjacent sequences to construct an adjacency graph structure that conforms to the actual sampling patterns of the samples. This restores the physical sampling relationships between frames, effectively mining the continuous correlation information of sludge micro-features under sequential views, and improving the accuracy and fit of the overall sludge state characterization.
[0118] In this embodiment, the activated sludge suspension is placed in a sample box, and microscopic videos of the activated sludge are captured by moving the microscope up, down, left, and right. Frames are extracted from the video images of specific spatial points within the sample box, and a graph structure is built based on the original spatial positions to complete graph learning and obtain a comprehensive characterization. The activated sludge suspension can also be placed in containers such as reagent kits. By placing the activated sludge suspension in the sample box, combined with the precise positioning and dynamic capture of the microscope, stable acquisition of the microscopic morphology of the activated sludge is achieved. For the microscopic video stream acquired at feature points, each frame is mapped to a graph node, feature vectors are assigned as node features, and an adjacency graph is constructed based on the relative relationships between feature points. This accurately restores the physical adjacency relationships and inter-frame sequence associations of the sample acquisition, improving the accuracy and fit of the overall microscopic state characterization of the sludge sample. Through the coordinated application of multiple methods such as continuous field-of-view acquisition and feature point acquisition, panoramic coverage and precise focusing of the activated sludge sample are achieved, balancing acquisition efficiency and feature integrity, and ensuring the systematicity and accuracy of the sample data.
[0119] Based on any of the above embodiments of this application, Embodiment Six of this application proposes a method for assessing the state of activated sludge based on microscopic video, which can be referred to the above description and will not be repeated hereafter. Based on this, please refer to... Figure 2 , Figure 2 This is a flowchart illustrating Example 6 of the activated sludge state assessment method based on microscopic video provided in this application. Before the step of inputting the comprehensive features of the sample into a pre-trained assessment model and determining the state parameters of the sludge sample based on the model output, the following steps are included:
[0120] Step S41: Control multiple training sludge samples to move in a preset manner, and acquire training microscopic video streams for the training sludge samples.
[0121] In this embodiment, multiple groups of activated sludge in different states are selected as training samples, and each training sludge sample is controlled to move on the microscope stage in a preset manner. Training microscopic video streams corresponding to each group of training sludge samples are simultaneously acquired to ensure that the microscopic video streams cover the microscopic features of different fields of view of the samples, and that the acquisition rules are consistent with the actual evaluation stage. The acquired training microscopic video streams are standardized in format, unifying parameters such as resolution and frame rate to form a standardized training video dataset. Multiple groups of microscopic video streams of sludge samples in different states provide the model with microscopic feature data covering the entire range of sludge states.
[0122] Step S42: Obtain multidimensional analysis indicators of the training sludge samples.
[0123] In this embodiment, while acquiring training microscopic video streams, standard detection methods are used to measure multidimensional analytical indicators for each group of training sludge samples, including core state indicators such as MLSS, SVI, SOUR, viscosity, and KLa. The acquired multidimensional analytical indicators serve as labels for model training, and these indicators cover multiple dimensions of sludge physical and biochemical states, comprehensively reflecting the true state of the sludge.
[0124] Step S43: Associate the training microscopic video streams and multidimensional analysis indicators of each training sludge sample with the training samples of the training sludge sample.
[0125] In this embodiment, the training sludge sample is used as a unique identifier, and the training microscopic video stream corresponding to each group of samples is associated and matched with the multidimensional analysis index label. The training microscopic video stream and multidimensional analysis index are combined and encapsulated into an independent training sample. This achieves precise binding between visual data and state labels during the training phase, allowing the model to establish a correspondence between sludge features and state indicators.
[0126] Step S44: Train the original model based on the training samples to obtain the evaluation model.
[0127] In this embodiment, frame extraction, feature extraction, and graph structure modeling operations, consistent with the actual evaluation stage, are performed on the training microscopic video stream in the training samples to generate comprehensive sample features. These comprehensive features are then input into the original GNN regression or classification model, and the model parameters are optimized using multidimensional analysis indicators as supervisory signals. The model is iteratively trained using methods such as cross-validation and hyperparameter tuning until the error between the model's prediction results and the true labels reaches a preset threshold, resulting in a converged pre-trained evaluation model. This allows the model to fully learn the mapping rules of the entire process—feature extraction, graph structure fusion, and state indicator prediction—during the training phase, ensuring that the trained evaluation model can accurately identify the sludge state corresponding to different microscopic features. Iterative training further improves the model's prediction accuracy and generalization ability, enabling the model to adapt to the sludge state evaluation needs under different operating conditions.
[0128] In this embodiment, by collecting and aligning multi-state sludge microscopic video streams with the evaluation scenario, obtaining accurate multi-dimensional state labels, and constructing structured training samples, the trained evaluation model can accurately establish the mapping relationship between the microscopic visual features of sludge and macroscopic state indicators, and has high generalization and prediction accuracy.
[0129] Based on any of the above embodiments of this application, Embodiment Seven of this application proposes a method for assessing the state of activated sludge based on microscopic video, which can be referred to the above description and will not be repeated hereafter. Based on this, the step of training the original model based on training samples to obtain the assessment model includes:
[0130] Step S441: For each training sample, based on a preset frame extraction method, obtain multiple frames of training microscopic video images corresponding to the training microscopic video stream.
[0131] In this embodiment, for each training sample, a preset frame extraction method, identical to that used in the actual sludge state assessment stage, is employed in the training microscopic video stream. This method involves extracting frames from the training microscopic video stream according to a fixed time or frame interval, resulting in multiple training microscopic video images. This achieves full alignment of frame extraction rules between the training and actual assessment stages, preventing discrepancies between the model's learned features and the actual application scenario due to differences in frame extraction methods, and ensuring the consistency and effectiveness of model training.
[0132] Step S442: Aggregate the feature vectors of each frame of training microscopic video image into the training sample comprehensive features corresponding to the training sludge sample, and obtain the training sample comprehensive features corresponding to each training sample.
[0133] In this embodiment, high-dimensional feature vectors are extracted from each frame of the training microscopic video image. These feature vectors encompass sludge morphology, texture, grayscale, color distribution, and implicit high-order structural features. The feature vectors from all frames of the training sample are aggregated into a feature set, and a corresponding graph structure, such as a fully connected graph or an adjacency graph, is constructed based on the preset acquisition method of the training microscopic video stream. Through feature fusion and readout operations using a graph neural network, the features from multiple frames are integrated into a single, comprehensive feature that characterizes the overall state of the training sludge sample. This comprehensive feature, generated by fusing features from multiple frames using a graph neural network, comprehensively characterizes the full-domain microscopic state of the training sludge sample, providing highly representative input features for establishing feature-index mapping relationships in the model.
[0134] Step S443: Use the comprehensive features of each training sample as the input to the original model, and use the multidimensional analysis indicators corresponding to each training sample as the supervision signal to train and obtain the evaluation model.
[0135] In this embodiment, the comprehensive features of all training samples are used as the input data of the original model, and the multidimensional analysis indicators bound to each training sample, such as MLSS, SVI, SOUR, viscosity, and KLa, are used as supervision labels for model training. A model loss function, such as mean squared error or cross-entropy loss, is set, and the network parameters of the original model are iteratively optimized using the backpropagation algorithm, allowing the model to learn the mapping relationship between the comprehensive features of the training samples and the multidimensional analysis indicators. The model training process is optimized using methods such as cross-validation, hyperparameter tuning, and early stopping mechanisms until the error between the model's predicted value and the true supervision label reaches a preset threshold. After model convergence, an evaluation model is obtained. These multiple model optimization techniques effectively improve the model's prediction accuracy, generalization ability, and robustness. The final evaluation model can directly adapt to the sludge state parameter prediction needs of real-world scenarios.
[0136] In this embodiment, training sample comprehensive features are generated through frame extraction and feature aggregation methods consistent with the actual evaluation process. The original model is trained using multidimensional analysis indicators as supervision signals, enabling the trained evaluation model to accurately learn the mapping law between the microscopic features of the entire sludge domain and the macroscopic state indicators, ensuring the model's prediction accuracy, generalization ability, and adaptability to actual application scenarios.
[0137] Based on any of the above embodiments of this application, Embodiment 8 of this application proposes a method for assessing the state of activated sludge based on microscopic video, which can be referred to the above description and will not be repeated hereafter. Based on this, the steps of training the assessment model by using the comprehensive features of each training sample as input to the original model and the multidimensional analysis indicators corresponding to each training sample as supervision signals include:
[0138] Step S431: Use the comprehensive features of the training samples corresponding to the training samples as input to the original model to obtain the predicted value.
[0139] In this embodiment, the comprehensive features of the training samples corresponding to the training samples are input into the untrained original model. The original model performs feature parsing and mapping on the input comprehensive features of the training samples, and outputs the predicted values of the sludge multidimensional analysis indicators corresponding to the training samples. This achieves a preliminary mapping from the comprehensive features of the training samples to the predicted values of sludge state indicators, allowing the original model to complete feature interpretation and indicator prediction based on the current parameters, providing an intuitive reference for subsequent model parameter optimization.
[0140] Step S432: Using the multidimensional analysis index corresponding to the training sample as a supervision signal, the prediction bias of the predicted value is obtained.
[0141] In this embodiment, multidimensional analysis indicators corresponding to the training samples are obtained and used as supervision signals for model training. A preset loss function, such as mean squared error or mean absolute error, is used to calculate the prediction deviation between the predicted value output by the original model and the supervision signal, quantifying the degree of difference between the model's current prediction result and the true value. The quantitative calculation of the prediction deviation for a single set of samples is completed through real multidimensional analysis indicators, providing a clear optimization basis for model back-optimization.
[0142] Step S433: Based on the prediction bias, the model parameters of the original model are optimized in reverse. Then, based on the next training sample, the comprehensive features of the training samples corresponding to the training samples are used as the input of the original model to obtain the predicted values. This process is repeated to obtain the evaluation model.
[0143] In this embodiment, based on the calculated prediction bias, the gradient update values of the parameters of each layer are calculated by backpropagation along the model network layers. The network parameters of the original model are then optimized and adjusted according to a preset learning rate to reduce the model's prediction bias. The next training sample is retrieved, and the comprehensive features of the corresponding training sample are used as input to the original model. The prediction value acquisition operation is repeated, and a new prediction value is obtained based on the optimized parameters. This process of prediction value acquisition, bias calculation, and parameter optimization is continuously looped, traversing the training samples and iterating multiple times until the model's prediction bias converges to a preset threshold, resulting in a converged evaluation model. Backpropagation enables precise iterative optimization of model parameters, allowing the model to gradually correct prediction bias. Multiple rounds of iterative training allow the model to fully learn the feature patterns of the entire training sample set, effectively improving the model's prediction accuracy, robustness, and generalization ability, ultimately resulting in a converged evaluation model adapted to the actual evaluation scenario.
[0144] In this embodiment, the predicted value is obtained by inputting the comprehensive features of the training samples, the prediction deviation is calculated by combining the measured multidimensional analysis indicators, and the model parameters are optimized in reverse based on the deviation and iteratively trained, so that the model gradually corrects the prediction error and continuously learns the mapping law between features and indicators, and finally obtains a convergent evaluation model with high prediction accuracy and strong generalization ability.
[0145] For example, to help understand the technical concept or principle of the activated sludge state assessment method based on microscopic video after combining this embodiment with the above embodiments, please refer to... Figure 3 , Figure 3 The flowchart illustrating the activated sludge state assessment method based on microscopic video provided in this application is as follows:
[0146] Activated sludge samples are obtained from biological tanks or reactors in wastewater treatment systems, and these samples are subjected to microscopic observation, with corresponding activated sludge microscopic videos continuously acquired. The acquired microscopic videos are used to characterize the overall microscopic structure of the activated sludge under the same sampling conditions. Subsequently, the microscopic videos are timed frame-by-frame according to preset rules to obtain multiple frames of microscopic images. For each frame, features are extracted using a convolutional neural network or automated imageomics method to generate a corresponding high-dimensional image feature vector. Based on this, the features from multiple frames obtained from the same microscopic video are constructed into graph-structured data, with each frame of microscopic image as a node. Connections between nodes are established according to the acquisition method of the microscopic video. A graph neural network is used to jointly model the features of multiple frames, achieving information fusion and expression updates between image features from different viewpoints. After completing the joint modeling of multiple frame features, graph-level readout operations are used to aggregate the node features into a single comprehensive sample representation. Based on this comprehensive representation, the evaluation results of the activated sludge are output, achieving rapid quantitative assessment of the activated sludge state.
[0147] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the activated sludge state assessment method based on microscopic video of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0148] This application provides a microscopic video-based activated sludge condition assessment device. The microscopic video-based activated sludge condition assessment device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the microscopic video-based activated sludge condition assessment method in the above embodiment 1.
[0149] The following is for reference. Figure 4 The diagram illustrates a structural schematic of a microscopic video-based activated sludge condition assessment device suitable for implementing embodiments of this application. The microscopic video-based activated sludge condition assessment device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablets, and vehicle-mounted terminals, as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The activated sludge condition assessment device based on microscopic video shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0150] like Figure 4As shown, the activated sludge condition assessment device based on microscopic video may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the activated sludge condition assessment device based on microscopic video. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the microscopic video-based activated sludge condition assessment equipment to wirelessly or wiredly communicate with other devices to exchange data. Although the figure shows a microscopic video-based activated sludge condition assessment equipment with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0151] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0152] The activated sludge state assessment device based on microscopic video provided in this application, employing the activated sludge state assessment method based on microscopic video in the above embodiments, can solve the technical problem that activated sludge assessment results are difficult to reflect the actual state of the entire sample due to analysis of activated sludge using single-frame microscopic images. Compared with the prior art, the beneficial effects of the activated sludge state assessment device based on microscopic video provided in this application are the same as those of the activated sludge state assessment device based on microscopic video provided in the above embodiments, and other technical features in this activated sludge state assessment device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0153] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0154] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0155] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the activated sludge state assessment method based on microscopic video in the above embodiments.
[0156] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), or any suitable combination thereof.
[0157] The aforementioned computer-readable storage medium may be included in the microscopic video-based activated sludge condition assessment device; or it may exist independently and not assembled into the microscopic video-based activated sludge condition assessment device.
[0158] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the activated sludge state assessment device based on microscopic video, the activated sludge state assessment device based on microscopic video performs the following actions: controls the sludge sample to move in a preset manner and acquires a microscopic video stream of the sludge sample; acquires multiple frames of microscopic video images corresponding to the microscopic video stream based on a preset frame extraction method; aggregates the feature vectors of each frame of microscopic video images into a comprehensive sample feature corresponding to the sludge sample; inputs the comprehensive sample feature into a pre-trained assessment model, and determines the state parameters of the sludge sample based on the model output of the assessment model.
[0159] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0161] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0162] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described activated sludge state assessment method based on microscopic video. This solves the technical problem that activated sludge assessment results based on single-frame microscopic images are difficult to reflect the actual state of the entire sample. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the activated sludge state assessment method based on microscopic video provided in the above embodiments, and will not be repeated here.
[0163] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the activated sludge state assessment method based on microscopic video as described above.
[0164] The computer program product provided in this application can solve the technical problem that activated sludge assessment results based on single-frame microscopic images are difficult to reflect the actual state of the overall sample. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the activated sludge state assessment method based on microscopic video provided in the above embodiments, and will not be repeated here.
[0165] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for assessing the state of activated sludge based on microscopic video, characterized in that, The activated sludge state assessment method based on microscopic video includes: Control the sludge sample to move in a preset manner and acquire a microscopic video stream of the sludge sample; Based on a preset frame extraction method, multiple frames of microscopic video images corresponding to the microscopic video stream are obtained. Determine the feature vector of each frame of the microscopic video image, wherein the feature vector characterizes the morphological features, texture features, grayscale, color distribution features and / or implicitly learned high-order structural features of the sludge in that frame; The feature vectors of the sludge samples are aggregated to obtain a feature set; If the preset method is continuous field-of-view acquisition, each frame of microscopic video image is used as a node, and the feature vector corresponding to the node is used as the node feature; based on the frame order indicated by the feature set, the edges between each node are determined to construct a fully connected graph as the graph structure. If the preset method is feature point acquisition, each frame of microscopic video image is used as a node, and the feature vector corresponding to the node is used as the node feature; according to the relative relationship between the feature points of each frame of microscopic video image, the edges between each node are determined to construct an adjacency graph as a graph structure. The fused node features obtained by fusing multi-frame features based on the graph structure using a graph neural network are used to obtain the comprehensive features of the sample representing the global features of the sludge sample through a graph-level readout strategy. The comprehensive features of the sample are input into a pre-trained evaluation model, and the state parameters of the sludge sample are determined based on the model output of the evaluation model.
2. The activated sludge state assessment method based on microscopic video as described in claim 1, characterized in that, The step of acquiring multiple frames of microscopic video images corresponding to the microscopic video stream based on a preset frame extraction method includes: According to a set time interval or a set frame interval, using the first frame of the microscopic video stream as the first frame microscopic video image, multiple frames of microscopic video images corresponding to the microscopic video stream are obtained; or... Based on the frame extraction method corresponding to the shape of the sludge sample, multiple frames of microscopic video images corresponding to the microscopic video stream are obtained.
3. The activated sludge state assessment method based on microscopic video as described in claim 1, characterized in that, Before the step of inputting the comprehensive features of the sample into a pre-trained evaluation model and determining the state parameters of the sludge sample based on the model output of the evaluation model, the following steps are included: Multiple training sludge samples are controlled to move in a preset manner, and training microscopic video streams for the training sludge samples are acquired. Obtain multidimensional analysis indicators of the training sludge samples; The training microscopic video streams of each of the training sludge samples are associated with the multidimensional analysis indicators to form the training samples of the training sludge samples. The evaluation model is obtained by training the original model based on the training samples.
4. The activated sludge state assessment method based on microscopic video as described in claim 3, characterized in that, The step of training the original model based on the training samples to obtain the evaluation model includes: For each training sample, based on a preset frame extraction method, multiple frames of training microscopic video images corresponding to the training microscopic video stream are obtained. The feature vectors of each frame of the training microscopic video image are aggregated into the training sample comprehensive features corresponding to the training sludge sample, thus obtaining the training sample comprehensive features corresponding to each training sample. The comprehensive features of each training sample are used as input to the original model, and the multidimensional analysis indicators corresponding to each training sample are used as supervision signals to train the evaluation model.
5. The activated sludge state assessment method based on microscopic video as described in claim 3, characterized in that, The step of training the evaluation model by using the comprehensive features of each training sample as input to the original model and the multidimensional analysis indicators corresponding to each training sample as supervision signals includes: The comprehensive features of the training samples corresponding to the training samples are used as the input of the original model to obtain the predicted value; Using the multidimensional analysis index corresponding to the training sample as a supervision signal, the prediction bias of the predicted value is obtained; Based on the prediction bias, the model parameters of the original model are optimized in reverse, and based on the next training sample, the step of using the comprehensive features of the training samples corresponding to the training samples as input to the original model to obtain the predicted value is performed to iteratively obtain the evaluation model.
6. A device for assessing the condition of activated sludge based on microscopic video, characterized in that, The activated sludge state assessment device based on microscopic video includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the activated sludge state assessment method based on microscopic video as described in any one of claims 1 to 5.
7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the steps of the activated sludge state assessment method based on microscopic video as described in any one of claims 1 to 5.