Digital twin waterworks data modeling and analysis method and system based on artificial intelligence

By using quantum state entanglement entropy generation adversarial networks to expand data in water plant data processing, dynamic population evolution optimization neural networks for feature extraction, and combining fuzzy logic and adaptive sparseness regulation random forest algorithm, the problems of low data processing efficiency and poor generalization capabilities in the existing technology are solved, and more efficient and accurate data analysis is achieved.

CN120065932AInactive Publication Date: 2025-05-30YUNNAN AGRICULTURAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510131007.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing water plant sensor data, it is difficult for the prior art to deal with noise, abnormal data and multi-dimensional data correlation, resulting in insufficient accuracy, low efficiency, poor generalization ability of the model, and easy to overfit.

Method used

Data augmentation network based on quantum state entanglement entropy is used to generate simulated data with high diversity and fidelity; feature extraction is performed using neural network models based on dynamic population evolution optimization; and fuzzy logic and adaptive sparseness regulation mechanism are introduced in the random forest algorithm to improve the robustness and accuracy of the classifier.

Benefits of technology

It improves the processing capacity of the tap water plant data, enhances the generalization ability and training efficiency of the model, reduces the overfitting phenomenon, and can more accurately indicate the operating status of the tap water plant.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065932A_ABST
    Figure CN120065932A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and discloses a digital twin waterworks data modeling and analysis method and system based on artificial intelligence, and the method comprises the steps: collecting the data of a plurality of sensors of a waterworks, and carrying out the manual marking of the collected sensor data; expanding the manually labeled sensor data by adopting a quantum state entanglement entropy-based generative adversarial network to generate simulation data; inputting the simulation data into a neural network model based on dynamic group evolution optimization for feature extraction; classifying the extracted features by adopting a random forest algorithm based on fuzzy logic; indicating the state of the waterworks according to the classification result. According to the method, the waterworks complex data processing capability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method and system for digital twin waterworks data modeling and analysis based on artificial intelligence. Background Art

[0002] During the construction of a digital twin waterworks, with the continuous expansion of the operation scale and the improvement of the automation level of the waterworks, relying on a sensor network to monitor various parameters such as water quality, flow rate, and pressure in real time has become an important means to ensure the efficient and stable operation of the waterworks. However, with the increase in data volume and the improvement of monitoring requirements, existing data analysis methods are difficult to handle these complex and diverse data. Especially when facing noise, abnormal data, and the correlation of multi-dimensional data, traditional modeling and analysis techniques are prone to problems such as insufficient accuracy and low efficiency. In addition, due to the characteristics of time series and real-time nature of sensor data acquisition in the waterworks, the data has complex time-varying and non-linear characteristics, and traditional methods are difficult to capture the deep features in the data, resulting in inaccurate judgment of the operating state of the system. At the same time, the acquisition cost of training data is relatively high, the labeled samples are limited, and the model is prone to overfitting and cannot be well generalized to new data scenarios.

[0003] The Chinese invention patent with the application number CN202410902972.3 proposes a classification and grading storage method for reservoir digital twin data. It relates to the field of data processing for oil and gas resource management, aiming to improve data access efficiency and reduce data management costs. The present invention maps reservoir digital twin data to corresponding categories according to the initialized data dictionary respectively. If the mapping is empty, the data dictionary is updated using an adaptive rule based on a decision tree and then mapped again; then the classified reservoir digital twin data is graded for heat using a cold and hot data grading method based on micro-clustering to obtain corresponding heat scores; the reservoir digital twin data is marked as cold data or hot data according to the heat scores; finally, the cold data and hot data are stored separately. The data classification of the present invention is fast and reliable, and can improve the matching degree between the dynamic grading of cold and hot data and the current user needs, and reduce the data storage cost while meeting the requirement of efficient data access.

[0004] The Chinese invention patent with the application number CN202410674238.6 proposes a method for judging the end point of copper converter blowing based on digital twin and deep learning algorithm. This method first collects the flame images during the converter smelting through an image acquisition device and classifies the flame images. Secondly, a lightweight model ConvNext algorithm is built to train the images, and accurate classification of the images is achieved through the deep learning algorithm. Finally, the algorithm is integrated into the digital twin system of the copper smelting plant to complete the end point judgment of the copper converter based on digital twin. In the present invention, the flame images of the entire converter cycle are used to extract features and classify them by using the ConvNext network as an important alternative means to the manual inspection method and the instrument measurement method. The method not only has a high accuracy rate, but also can optimize the entire copper smelting process and has a high application prospect in the field of digital twin.

[0005] The Chinese invention patent with the application number CN202410993525.3 proposes a method and system for substation digital twin early warning decision-making based on a knowledge graph. The method includes obtaining multi-modal full-scale data of the substation and performing data screening and processing to obtain multi-source heterogeneous data carrying effective operation information, analyzing the knowledge graph association relationships in the multi-source heterogeneous data, obtaining the mapping relationship between the substation equipment ontology and the twin body, performing data mapping and association matching processing on the substation equipment concept framework and the pre-constructed twin model to obtain the knowledge graph and digital twin mapping model of the substation, obtaining the cross-section data of the substation equipment operation in real time, and inputting the cross-section data into the knowledge graph and digital twin mapping model for fault and anomaly detection and decision-making analysis to obtain early warning decision-making suggestions integrating the knowledge graph and the digital twin model. The present invention has the effect of improving the accuracy of the early warning decision-making of the substation digital twin model.

[0006] The existing technologies have the following deficiencies:

[0007] 1. When traditional data augmentation methods are used to deal with the problem of insufficient samples, the diversity and fidelity of the generated data are limited, which easily leads to insufficient generalization ability of the model and cannot effectively increase the scale and diversity of the data set.

[0008] 2. Traditional neural networks are prone to problems such as gradient disappearance and gradient explosion during the feature extraction process, and the evolutionary algorithm lacks an adaptive adjustment mechanism, resulting in low model training efficiency, being easily trapped in local optimal solutions, and being difficult to fully extract the deep features of the data.

[0009] 3. The traditional random forest algorithm shows insufficient robustness in dealing with noise and abnormal data. Especially during the node splitting process, the single strategy relying on information gain is prone to overfitting and cannot accurately classify samples with high uncertainty. In addition, the decision tree construction algorithm does not consider the sparsity of data, resulting in inaccurate node splitting decisions when dealing with sparse data and affecting the overall performance of the classifier in sparse data scenarios. Summary of the Invention

[0010] To solve the problems existing in the prior art, the present invention provides a method and system for data modeling and analysis of a digital twin waterworks based on artificial intelligence, which can improve the processing ability of complex data of the waterworks.

[0011] To achieve the above object, the present invention provides the following solutions:

[0012] A method for data modeling and analysis of a digital twin waterworks based on artificial intelligence, the method comprising:

[0013] Collecting multiple sensor data of the waterworks and manually annotating the collected sensor data;

[0014] Using a generative adversarial network based on quantum state entanglement entropy to expand the manually annotated sensor data to generate simulated data;

[0015] Inputting the simulated data into a neural network model based on dynamic population evolution optimization for feature extraction;

[0016] Using a random forest algorithm based on fuzzy logic to classify the extracted features;

[0017] Indicating the state of the waterworks according to the classification result.

[0018] Preferably, the collected sensor data includes: flow rate, water temperature, water pressure, chlorine content, turbidity, pH value, conductivity, hardness, residual chlorine and lead content; the types of manual annotation include: normal operation, mild abnormality and severe abnormality.

[0019] Preferably, the generative adversarial network based on quantum state entanglement entropy includes: a generator and a discriminator;

[0020] Wherein, the generator generates data through quantum noise, and the discriminator discriminates the generated data from the real data;

[0021] Using a generative adversarial network based on quantum state entanglement entropy to expand the manually annotated sensor data to generate simulated data includes:

[0022]

[0023] In the formula, Xgen,c represents the data output by the generator, f c () is the non-linear transformation function of the generator, θ c are the parameters of the generator, ∈ c is the input quantum noise, Ψ c is the quantum layer state represents the coupling operation between the quantum state and traditional data.

[0024] Preferably, inputting the simulation data into the neural network model based on dynamic population evolution optimization for feature extraction includes:

[0025]

[0026] In the formula, P select (i) represents the probability that the i-th individual is selected; γ pse is the parameter controlling the selection pressure; L pk is the loss of the neural network corresponding to the k-th individual; L pi is the loss of the neural network corresponding to the i-th individual; N p is the population size.

[0027] Preferably, using the random forest algorithm based on fuzzy logic to classify the extracted features includes:

[0028]

[0029] In the formula, y t,k is the classification result of tree t for the data point X u ; Y u is the final classification result; T is the total number of trees; 1() is the indicator function, which takes the value of 1 when y t,k =k, otherwise 0.

[0030] The present invention also provides a data modeling and analysis system for a digital twin waterworks based on artificial intelligence. The system is used to implement any one of the above methods. The system includes: a collection module, an expansion module, an extraction module, a classification module, and an indication module;

[0031] The collection module is used to collect the data of multiple sensors in the waterworks and perform manual annotation on the collected sensor data;

[0032] The expansion module is used to expand the manually annotated sensor data by using a generative adversarial network based on quantum state entanglement entropy to generate simulation data;

[0033] The extraction module is used to input the simulation data into the neural network model based on dynamic population evolution optimization for feature extraction;

[0034] The classification module is used to classify the extracted features by using a random forest algorithm based on fuzzy logic;

[0035] The indication module is used to indicate the state of the waterworks according to the classification result.

[0036] Preferably, the collected sensor data includes: flow rate, water temperature, water pressure, chlorine content, turbidity, pH value, conductivity, hardness, residual chlorine, and lead content; the manually labeled types include: normal operation, mild abnormality, and severe abnormality.

[0037] Preferably, the generative adversarial network based on quantum state entanglement entropy includes: a generator and a discriminator;

[0038] Among them, the generator generates data through quantum noise, and the discriminator discriminates the generated data from the real data;

[0039] The generative adversarial network based on quantum state entanglement entropy is used to augment the manually labeled sensor data, and the generated simulated data includes:

[0040]

[0041] In the formula, X gen,c represents the data output by the generator, f c ( ) is the non-linear conversion function of the generator, θ c is the parameter of the generator, ∈ c is the input quantum noise, Ψ c is the quantum layer state, represents the coupling operation between the quantum state and the traditional data.

[0042] Preferably, inputting the simulated data into a neural network model based on dynamic population evolution optimization for feature extraction includes:

[0043]

[0044] In the formula, P select (i) represents the probability that the i-th individual is selected; γ pse is the parameter controlling the selection pressure; L pk is the loss of the neural network corresponding to the k-th individual; L pi is the loss of the neural network corresponding to the i-th individual; N p is the population size.

[0045] Preferably, using a random forest algorithm based on fuzzy logic to classify the extracted features includes:

[0046]

[0047] In the formula, y t,kis the classification result of tree t for data point X u ; Y u is the final classification result; T is the total number of trees; 1() is the indicator function, which takes the value 1 when y t,k = k and 0 otherwise.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] The present invention proposes a method and system for data modeling and analysis of a digital twin waterworks based on artificial intelligence, aiming to efficiently process and analyze multi-parameter data collected by sensors in the waterworks. The method includes the following steps: First, multiple sensors in the waterworks collect various parameters such as flow rate, water temperature, water pressure, chlorine content, etc., and these data are manually labeled, and the labeled categories include normal operation, mild abnormality, and severe abnormality. To address the problem of insufficient training data, the present invention uses a generative adversarial network algorithm based on quantum state entanglement entropy for data augmentation to generate simulated data with high diversity and fidelity. The augmented data is input into a neural network model based on dynamic population evolution optimization for feature extraction. This model can dynamically adjust the evolution rules, effectively avoid problems such as gradient disappearance, gradient explosion, and local optimal solutions, and improve the generalization ability and training efficiency of the model. In the data classification process, a random forest algorithm based on fuzzy logic is used to improve the robustness of the classifier by adjusting the information gain and fuzzy membership degree. Especially in the case of noise and abnormal data, it can effectively reduce the overfitting phenomenon. In addition, the present invention also adopts an adaptive sparsity adjustment mechanism to automatically adjust the sparsity during the decision tree node splitting process, optimize the decision-making process, and improve the accuracy and adaptability of the model. Finally, the data of the waterworks is classified through the feature extraction and classification model, and the classification result is used to indicate the normal operation, mild abnormality, or severe abnormality state of the waterworks. The present invention provides an efficient and robust method for data modeling and analysis of a waterworks, which can improve the processing ability of complex data of the waterworks. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0051] Figure 1 is a schematic flow chart of a method for data modeling and analysis of a digital twin waterworks based on artificial intelligence according to an embodiment of the present invention;

[0052] Figure 2 is a schematic flow chart of the training of the feature extraction model according to an embodiment of the present invention. Detailed Implementation Modes

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation modes.

[0055] Embodiment 1

[0056] As Figure 1 shown, the present invention proposes a method for data modeling and analysis of a digital twin water treatment plant based on artificial intelligence, and the main steps are as follows:

[0057] S1. Data collection and annotation

[0058] The data source of the present invention is multiple sensor networks of a water treatment plant. These sensors can monitor various parameters such as water quality, flow rate, pressure, and temperature in real time. To meet the high-efficiency processing requirements of the present invention, the collected data is stored in a time-series database with a high read-write rate, and the data format is a structured JSON format.

[0059] In one embodiment, the attributes of the data include:

[0060] Ra is the flow rate (cubic meters per hour), reflecting the flow rate of the water in the water treatment plant;

[0061] Ta is the water temperature (degrees Celsius), which is related to the efficiency and method of water quality treatment;

[0062] Pa is the water pressure (Pascals), indicating the pressure state of the water in the pipe network;

[0063] Ca is the chlorine content (mg / L), showing the concentration of disinfectant in the water;

[0064] Da is the turbidity (NTU), an index of the clarity of water;

[0065] Ea is the pH value, reflecting the acidity and alkalinity of the water;

[0066] Fa is the conductivity (microSiemens per centimeter), an index for measuring the ion concentration in the water;

[0067] Ga is the hardness (mg / LCaCO3), representing the total amount of calcium and magnesium ions in the water;

[0068] Ha is the residual chlorine (mg / L), which is a key indicator for evaluating the disinfection effect;

[0069] Ia is the lead content (μg / L), which is an important attribute for measuring heavy metal pollution in water.

[0070] It should be noted that this embodiment is only to illustrate a data format and type of the present invention. In actual applications, the attributes of data are usually more than 10, and the number of data attributes may reach dozens or even hundreds.

[0071] Furthermore, the collected data is labeled. The labeling method of the present invention is manual labeling. In one embodiment, the labeling categories include:

[0072] Normal operation: All indicators are within the specified safety and efficiency ranges;

[0073] Slight abnormality: One or more indicators deviate slightly from the normal values, but do not affect the overall operation;

[0074] Severe abnormality: One or more indicators deviate severely from the normal range and require immediate intervention.

[0075] S2. Data augmentation

[0076] It can be understood that in the task of the present invention, the acquisition, labeling, and preprocessing of training data are time-consuming and laborious, and insufficient training samples are likely to lead to poor generalization ability of the model and affect the accuracy of the model at the same time. The present invention uses a generative adversarial network algorithm based on quantum state entanglement entropy for data augmentation. Based on the traditional generative adversarial network, the superposition and entanglement characteristics of quantum states are used to enhance the information processing ability between the generator and the discriminator.

[0077] Specifically, the training process of the generative adversarial network algorithm based on quantum state entanglement entropy is as follows:

[0078] S201. Initialize the quantum layers in the generator and the discriminator. The initialization of the quantum layer involves the superposition state of quantum bits, and the initialization method of the quantum state is expressed as:

[0079]

[0080] In the formula, Ψ c represents the initial state of the quantum layers in the generator and the discriminator, α i,c represents the complex amplitude of the quantum state corresponding to the ground state, |i> represents the ground state, and n c represents the number of quantum bits.

[0081] S202. The generator receives a random noise signal, generates data after being processed by the quantum layer, and uses the superposition and entanglement of quantum states to simulate complex data distributions to generate highly realistic simulated data, which is expressed as:

[0082]

[0083] Wherein, X gen,c represents the data output by the generator, and f c ( ) is the non-linear conversion function of the generator, and θ c is the parameter of the generator, ∈ c is the input quantum noise, and Ψ c is the quantum layer state, represents the coupling operation between the quantum state and the traditional data.

[0084] In one embodiment, the non-linear conversion function of the generator is defined as the composition of a quantum-enhanced convolution operation and an activation function, expressed as:

[0085]

[0086] Wherein, Sig( ) is the Sigmoid activation function; represents the convolutional layer with parameter θ k,c for processing the input combined with the quantum layer and random noise; α k,c is the coefficient of each convolutional kernel, adjusting the contribution of each convolution operation; b k,c is the bias term of the convolution operation; K c is the number of convolutional kernels.

[0087] Furthermore, an adaptive quantum noise injection strategy is adopted to improve the diversity and quality of the generated data by dynamically adjusting the level of quantum noise injected into the generator. Different from the traditional noise injection method, the adaptive quantum noise injection utilizes the characteristics of quantum computing to dynamically adjust the noise parameters according to the feedback of the discriminator and optimize the output of the generator. Then, the calculation method of the quantum noise is expressed as:

[0088]

[0089] Wherein, U noise (φ c ) is a quantum gate operation dependent on the noise parameter φ c , represents n c qubits initialized to zero.

[0090] S203. The discriminator evaluates the data from the generator and the data in the real dataset. The discriminator also uses the embedded quantum layer to enhance its judgment accuracy and improve the ability to distinguish between generated data and real data, expressed as:

[0091] D c (X) = Sig(W c ·gc (X, Φ c ) + b c )

[0092] In the formula, D c (X) represents the output of the discriminator's evaluation of the authenticity of data X, W c is the weight of the discriminator, b c is the bias, g c () is the feature transformation function after quantum layer processing, Φ c is the quantum layer parameter of the discriminator.

[0093] In one embodiment, the feature transformation function after quantum layer processing includes a composite function of quantum gate operations and traditional neural network layers, expressed as:

[0094] g c (X, Φ c ) = Re(V c ·(U c (Φ c )·X) + c c )

[0095] In the formula, U c () represents quantum gate operations, Φ c is the parameter of the quantum gate, responsible for converting the result of quantum computing into a form that can be processed by the traditional layer; V c is the weight matrix of the network layer of the discriminator; c c is the bias vector of the network layer of the discriminator; Re() is the ReLU activation function.

[0096] S204. According to the evaluation results of the discriminator, adjust the parameters of the quantum layers in the generator and the discriminator, and optimize these parameters to reduce the difference between the generated data and the real data. Through continuous iteration, gradually improve the quality of the generated data. In one embodiment, the parameter update of the generator and the discriminator adopts the gradient descent method in adversarial training, and the update method is expressed as:

[0097]

[0098]

[0099] In the formula, and respectively represent the parameters of the generator and the discriminator in the t-th iteration; and respectively represent the parameters of the generator and the discriminator in the (t + 1)-th iteration; η c and μ c are the learning rates of the generator and the discriminator respectively, and L c is the loss function of the generative adversarial network. Preferably, ηc and μ c are both set to 0.01.

[0100] Furthermore, the calculation method of the loss function of the generative adversarial network is expressed as:

[0101]

[0102] In the formula, p data is the distribution of real data, and p z is the distribution of input noise; ~ means following a specific distribution; x gan represents the input of the discriminator; z gan represents the input of the generator; D c () represents the discriminator function; G c () represents the generator function.

[0103] S205. Repeat and iterate the above steps until the preset stop iteration condition is satisfied, which means the model training is completed. In one embodiment, the preset stop iteration condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0104] After the data augmentation model training is completed, use the trained data augmentation model to increase the number of samples. In one embodiment, assume the original collected samples are 800, and the data augmentation model generates 200 samples through augmentation. Then the augmented dataset contains 1000 samples.

[0105] S3. Feature extraction model training

[0106] Input the augmented data into the feature extraction model for training the feature extraction model. The present invention uses a 6-layer fully connected neural network for feature extraction. In the prior art, some solutions use neural networks for feature extraction. In some neural network structures, problems such as gradient disappearance, gradient explosion, or getting stuck in local optimal solutions may occur, affecting the training stability and model performance. The present invention uses a neural network algorithm based on dynamic population evolution optimization as the feature extraction model. In the traditional population evolution algorithm, the evolution of all individuals is based on fixed rules. The present invention uses a self-correction mechanism to automatically adjust the evolution rules according to the characteristics of the current training data, and adjusts the probabilities of crossover and mutation according to the change trend of the loss function, so that the algorithm can more flexibly adapt to different data distributions, improving the generalization ability and training efficiency of the model.

[0107] Specifically, as Figure 2 shown, the training process of the neural network algorithm based on dynamic population evolution optimization is as follows:

[0108] S301. According to the initialization method of the bionic algorithm, in the initialization stage, an initial population is generated, and each individual represents a configuration of network weights. Specifically, let the population size be N p , initialize the weights and biases for the i-th individual, which is expressed as:

[0109]

[0110]

[0111] In the formula, W pi is the weight of the neural network corresponding to the i-th individual, and b pi is the bias of the neural network corresponding to the i-th individual, represents the weight matrix of the i-th individual in the initial state; represents the bias of the i-th individual in the initial state; σ 2 represents the initialized variance; represents a normal distribution with a mean of 0 and a variance of σ 2 ; is the normal distribution. Preferably, σ 2 is set to 0.01.

[0112] S302. For each individual in the population, use its corresponding neural network configuration to process the input training data, calculate the output of the model, and evaluate its performance according to a predetermined loss function. Specifically, for the i-th individual, calculate the loss on the training data set using its weights and biases, which is expressed as:

[0113]

[0114] In the formula, L pi is the loss of the neural network corresponding to the i-th individual; mps represents the number of samples input in the current batch; represents the composite loss function, and fsig() represents the neural network model function; represents the feature of the j-th sample; represents the label of the j-th sample.

[0115] In one embodiment, the composite loss function includes a regularization term, which can increase the generalization ability of the model, and the calculation method is expressed as:

[0116]

[0117] In the formula, MSE() is the mean square error function, Reg(W pi ) is the regularization term, and λ ps is the regularization parameter.

[0118] Furthermore, the calculation method of the regularization term is expressed as:

[0119]

[0120] Wherein, W pi,k is the k-th weight of the neural network corresponding to the i-th individual.

[0121] S303. According to the fitness of the individuals, select the best-performing individuals from the current population for retention as candidate solutions for the next generation. Specifically, selection is based on the fitness of the individuals, and excellent individuals have a higher probability of being selected. The calculation method of the probability of an individual being selected is expressed as:

[0122]

[0123] Wherein, P select (i) represents the probability that the i-th individual is selected; γ pse is a parameter that controls the selection pressure; L pk is the loss of the neural network corresponding to the k-th individual. Preferably, γ pse is set to 2.

[0124] S304. Generate new individuals through crossover and mutation operations. The crossover operation allows two excellent individuals to exchange part of their genes to generate new offspring; the mutation operation randomly changes part of the genes in an individual to increase the diversity of the population. Specifically, the crossover operation randomly selects two individuals for gene exchange, expressed as:

[0125] W′ pi = α pcs W p1 +(1 - α pcs )W p2

[0126] b′ pi = α pcs b p1 +(1 - α pcs )b p2

[0127] Wherein, α pcs is the crossover rate, W p1 is the weight of the neural network corresponding to the first selected individual, b p1 is the bias of the neural network corresponding to the first selected individual, W p2 is the weight of the neural network corresponding to the second selected individual, b p2 is the bias of the neural network corresponding to the second selected individual, W′ pi is the weight of the neural network corresponding to the individual after the crossover operation, b′ pi is the bias of the neural network corresponding to the individual after the crossover operation. Preferably, α pcsSet to 0.3.

[0128] Moreover, the mutation operation performs a small random perturbation on the weights of the newly generated individuals, expressed as:

[0129]

[0130]

[0131] In the formula, τ 2 represents the variance of mutation, W″ pi is the weight of the neural network corresponding to the individual after the mutation operation, and b″ pi is the bias of the neural network corresponding to the individual after the mutation operation. Preferably, τ 2 is set to 0.04.

[0132] S305. Repeat the above steps iteratively until the preset iteration stop condition is met, which indicates that the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0133] S4. Classifier model training

[0134] Input the data after feature extraction into the classifier for the training of the classifier model. The present invention uses a random forest algorithm based on fuzzy classification as the classifier. The core of this algorithm lies in combining the advantages of fuzzy logic and random forest to achieve efficient classification of the data of the water treatment plant after feature extraction. In the traditional random forest algorithm, when selecting split nodes, it is usually based on maximizing the improvement of purity. In the present invention, the judgment criterion for node splitting is adjusted based on fuzzy logic, so that the model can be more robust in the face of noise and abnormal data, thereby reducing the risk of overfitting and improving the generalization ability of the model.

[0135] Specifically, the training process of the random forest algorithm based on fuzzy classification is as follows:

[0136] S401. According to the input data after feature extraction, set fuzzy logic rules, which are used to preliminarily screen which data should be classified into the same pre-classification set. The defined fuzzy classification rules are expressed as:

[0137] R u = f u (X u , C u )

[0138] In the formula, X u represents the data vector after feature extraction; C u is the parameter set of the fuzzy classification rule; R u is the data point Xu Classification results under fuzzy logic; f u () is the fuzzy classification function

[0139] Furthermore, the calculation method of the fuzzy membership degree of each data point is expressed as:

[0140]

[0141] In the formula, Σ u is the covariance matrix in fuzzy logic, which controls the dispersion range of the fuzzy membership degree;

[0142] μ u (X u ) represents the membership degree of X u .

[0143] Furthermore, the composition of the parameter set of the fuzzy classification function is:

[0144] C u ={c u,i |c u,i =h u (θ u,i , ψ u,i )}

[0145] In the formula, c u,i is the i-th parameter of the fuzzy classification rule; θ u,i and ψ u,i represent the center and width parameters of the fuzzy classification respectively. These parameters are obtained through an optimization process, such as the grid optimization method, to adapt to different data characteristics; h u () is the parameter generation function based on the Gaussian function.

[0146] In one embodiment, the parameter generation function based on the Gaussian function characterizes the influence of the center offset and diffusion of the fuzzy logic, enabling the fuzzy classification to adaptively adjust its range. The calculation method is expressed as:

[0147]

[0148] S402. After performing fuzzy classification, construct a decision tree in the random forest for each pre-classified data set. The construction of each tree starts with randomly selecting a feature subset. Different from the traditional random forest, each split considers not only the information gain but also the weight defined by the fuzzy rule, ensuring that the growth of the tree pays more attention to the key features in the data set. Specifically, for the samples of the k u -th class, the way of randomly selecting a feature subset and constructing a decision tree is expressed as:

[0149] T u,k =g u (Du,k , F u,k )

[0150] In the formula, D u,k represents the k u -th class data set; F u,k is a randomly selected feature subset; T u,k is the decision tree corresponding to k u -th class; g u () represents the decision tree construction function.

[0151] Furthermore, the decision of node splitting of the decision tree is determined by information gain and fuzzy weight adjustment, expressed as:

[0152]

[0153] In the formula, H(D u ) is the entropy of the data set D u ; D u,v is the subset after splitting; |D u,v | and |D u | are the number of data points in the subset and the original set respectively; ω u is the weight adjusted according to the fuzzy membership degree; Γ u (s) represents the splitting decision function of the s-th node of the decision tree.

[0154] Furthermore, the entropy of the data set provides a measure to evaluate the uncertainty of the data set, which is used to evaluate the effects of different features when the decision tree splits nodes. The calculation method is expressed as:

[0155]

[0156] In the formula, p u,j is the probability distribution of the j-th class elements in the data set D u , and m u is the number of classes.

[0157] Furthermore, the calculation of the fuzzy weight ω u takes into account the fuzzy membership degree of the data points, enabling the information gain calculation to pay more attention to important and highly uncertain data. The calculation method is expressed as:

[0158]

[0159] In the formula, λ u,i is the importance weight of the i-th data point; μ u,i is the i-th element of the covariance matrix in fuzzy logic.

[0160] S403. During the node splitting process of the decision tree, the sparsity of each node is detected and adjusted through an adaptive sparsity adjustment mechanism. If the sparsity of a node is too high, the algorithm will adjust its splitting strategy, select more appropriate features for splitting, or adjust the weights of the fuzzy rules to optimize the structure of the tree. Specifically, the calculation method of the sparsity of each node is expressed as:

[0161]

[0162] In the formula, N u,n is the index set of data points in the nth node; is the mean vector of N u,n ; S u (n) represents the sparsity of the nth node; δ u,n represents the local density of the nth node; X u,i represents the ith feature in the data vector after feature extraction.

[0163] Furthermore, the nodes with sparsity higher than the threshold λ u will adjust the splitting strategy, which is expressed as:

[0164] If S u (n)>λ u , then adjust ω u

[0165] Furthermore, the local density allows the system to adjust the sparsity of each node according to the actual distance between data points, improving the accuracy and adaptability of the node splitting strategy. The calculation method is expressed as:

[0166]

[0167] In the formula, σ u controls the sensitivity of similarity, and X u,j represents the jth feature in the data vector after feature extraction.

[0168] S404. After all decision trees are constructed, the random forest determines the final classification result through a voting mechanism. Each tree provides a classification result, and the output of the entire model is the class obtained based on the majority voting principle. The final output of the random forest is expressed as:

[0169]

[0170] In the formula, y t,k is the classification result of tree t for the data point X u ; Y u is the final classification result; T is the total number of trees; 1() is the indicator function, which takes the value 1 when y t,k =k, and 0 otherwise.

[0171] S5. Tap Water Plant Data Classification

[0172] Using the trained model to process new samples. In one embodiment, the collected raw data is input into the trained feature extraction model for feature processing. Further, the processed features are input into the classifier model for classification, and then the classification results are obtained. In this embodiment, the classification categories include: normal operation, mild anomaly, and severe anomaly.

[0173] Embodiment 2

[0174] The present invention also provides a digital twin tap water plant data modeling and analysis system based on artificial intelligence. The system is used to implement any one of the methods described above. The system includes: a collection module, an expansion module, an extraction module, a classification module, and an indication module;

[0175] The collection module is used to collect the data of multiple sensors in the tap water plant and perform manual annotation on the collected sensor data;

[0176] The expansion module is used to expand the manually annotated sensor data by using a generative adversarial network based on quantum state entanglement entropy to generate simulated data;

[0177] The extraction module is used to input the simulated data into a neural network model based on dynamic population evolution optimization for feature extraction;

[0178] The classification module is used to classify the extracted features by using a random forest algorithm based on fuzzy logic;

[0179] The indication module is used to indicate the state of the tap water plant according to the classification results.

[0180] In this embodiment, the collected sensor data includes: flow rate, water temperature, water pressure, chlorine content, turbidity, pH value, conductivity, hardness, residual chlorine, and lead content; the types of manual annotation include: normal operation, mild anomaly, and severe anomaly.

[0181] In this embodiment, the generative adversarial network based on quantum state entanglement entropy includes: a generator and a discriminator;

[0182] Among them, the generator generates data through quantum noise, and the discriminator discriminates the generated data from the real data;

[0183] Using the generative adversarial network based on quantum state entanglement entropy to expand the manually annotated sensor data to generate simulated data includes:

[0184]

[0185] In the formula, X gen,cdenotes the data output by the generator, f c () is the non - linear transformation function of the generator, θ c are the parameters of the generator, ∈ c is the input quantum noise, Ψ c is the quantum layer state denotes the coupling operation between the quantum state and the traditional data.

[0186] In this embodiment, inputting the analog data into the neural network model based on dynamic population evolution optimization for feature extraction includes:

[0187]

[0188] In the formula, P select (i) represents the probability that the i - th individual is selected; γ pse is the parameter controlling the selection pressure; L pk is the loss of the neural network corresponding to the k - th individual; L pi is the loss of the neural network corresponding to the i - th individual; N p is the population size.

[0189] In this embodiment, adopting the random forest algorithm based on fuzzy logic to classify the extracted features includes:

[0190]

[0191] In the formula, y t,k is the classification result of the tree t for the data point X u ; Y u is the final classification result; T is the total number of trees; 1() is the indicator function, which takes the value of 1 when y t,k =k, and 0 otherwise.

[0192] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A digital twin water plant data modeling and analysis method based on artificial intelligence, characterized in that: The method comprises: Collect multiple sensor data from the water plant and manually annotate the collected sensor data; A generative adversarial network based on quantum state entanglement entropy is used to expand the manually annotated sensor data to generate simulated data; The simulated data is input into a neural network model based on dynamic population evolution optimization for feature extraction; The extracted features are classified using the random forest algorithm based on fuzzy logic; Indicates the status of the water plant based on the classification results.

2. The method according to claim 1, characterized in that The sensor data collected include: flow, water temperature, water pressure, chlorine content, turbidity, pH value, conductivity, hardness, residual chlorine and lead content; the manually labeled types include: normal operation, slight abnormality and severe abnormality.

3. The method according to claim 1, characterized in that The generative adversarial network based on quantum state entanglement entropy includes: a generator and a discriminator; Among them, the generator generates data through quantum noise, and the discriminator distinguishes the generated data from the real data; The generative adversarial network based on quantum state entanglement entropy is used to expand the manually annotated sensor data to generate simulated data including: Where, X gen,c Represents the data output by the generator, f c () is the nonlinear transformation function of the generator, θ c is the parameter of the generator, ∈ c is the input quantum noise, Ψ c is the quantum layer state, Represents the coupling operation between quantum states and traditional data.

4. The method according to claim 1, characterized in that Inputting simulated data into a neural network model based on dynamic population evolution optimization for feature extraction includes: Where P select (i) represents the probability of the i-th individual being selected; γ pse is the parameter that controls the selection pressure; L pk is the loss of the neural network corresponding to the kth individual; L pi is the loss of the neural network corresponding to the i-th individual; N p is the population size.

5. The method according to claim 1, characterized in that The extracted features are classified using the random forest algorithm based on fuzzy logic including: In the formula, y t,k is the tree t for the data point X u The classification result of Y u is the final classification result; T is the total number of trees; 1() is the indicator function, when y t,k =k, it takes the value 1, otherwise it takes the value 0.

6. A digital twin water plant data modeling and analysis system based on artificial intelligence, the system is used to implement the method described in any one of claims 1 to 5, characterized in that: The system comprises: a collection module, an expansion module, an extraction module, a classification module and an indication module; The acquisition module is used to collect multiple sensor data of the water plant and manually annotate the collected sensor data; The expansion module is used to expand the manually annotated sensor data using a generative adversarial network based on quantum state entanglement entropy to generate simulated data; The extraction module is used to input the simulation data into a neural network model based on dynamic population evolution optimization to extract features; The classification module is used to classify the extracted features using a random forest algorithm based on fuzzy logic; The indication module is used to indicate the status of the water plant according to the classification result.

7. The system according to claim 6, characterized in that The sensor data collected include: flow, water temperature, water pressure, chlorine content, turbidity, pH value, conductivity, hardness, residual chlorine and lead content; the manually labeled types include: normal operation, slight abnormality and severe abnormality.

8. The system according to claim 6, characterized in that The generative adversarial network based on quantum state entanglement entropy includes: a generator and a discriminator; Among them, the generator generates data through quantum noise, and the discriminator distinguishes the generated data from the real data; The generative adversarial network based on quantum state entanglement entropy is used to expand the manually annotated sensor data to generate simulated data including: Where, X gen,c Represents the data output by the generator, f c () is the nonlinear transformation function of the generator, θ c is the parameter of the generator, ∈ c is the input quantum noise, Ψ c is the quantum layer state, Represents the coupling operation between quantum states and traditional data.

9. The system according to claim 6, characterized in that Inputting simulated data into a neural network model based on dynamic population evolution optimization for feature extraction includes: Where P select (i) represents the probability of the i-th individual being selected; γ pse is the parameter that controls the selection pressure; L pk is the loss of the neural network corresponding to the kth individual; L pi is the loss of the neural network corresponding to the i-th individual; N p is the population size.

10. The system according to claim 6, characterized in that The extracted features are classified using the random forest algorithm based on fuzzy logic including: In the formula, y t,k is the tree t for the data point X u The classification result of Y u is the final classification result; T is the total number of trees; 1() is the indicator function, when y t,k =k, it takes the value 1, otherwise it takes the value 0.

Citation Information

Patent Citations

  • Classification and classification storage method for oil reservoir digital twin data

    CN118466852A

  • Copper converter blowing end point judgment method based on digital twinning and deep learning algorithms

    CN118506092A

  • Substation digital twin early warning decision method and system based on knowledge graph

    CN118521433B