Permanent magnet motor energy storage and optimized operation method and system based on electricity price
By combining technologies such as variational modal decomposition, deep learning and probability modeling, the problem that traditional control methods are difficult to adapt to complex working conditions and electricity price changes is solved, more accurate prediction and better control strategies are achieved, and the operating stability and adaptability of permanent magnet motors are improved.
Patent Information
- Application Number
- CN202510094956.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The traditional permanent magnet motor control method is difficult to adapt to complex operating conditions and dynamically changing electricity prices, and lacks an effective treatment mechanism for system uncertainty and random disturbances.
Power prediction, electricity price prediction and control strategy optimization are adopted using technologies such as variational modal decomposition, deep belief neural network, space-time graph convolution neural network, quantum genetic optimization algorithm and hybrid Gaussian-hidden Markov model.
More accurate power and electricity price predictions are achieved, better motor operation control strategies are formulated, the stability and adaptability of the system are improved, and the optimal effect of energy storage and optimization of permanent magnet motors is achieved.
Smart Images

Figure CN119514816B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of motor energy storage and optimized operation, and in particular to a method and system for energy storage and optimized operation of a permanent magnet motor based on electricity price. Background Art
[0002] Permanent magnet motors are widely used in various industrial and civil fields due to their advantages such as high efficiency, high power density and reliability. As the requirements of power systems for energy efficiency and stability continue to increase, higher requirements are placed on the energy storage and optimized operation of permanent magnet motors.
[0003] Traditional permanent magnet motor control methods are usually based on pre-set control strategies or simple feedback control, which are difficult to adapt to complex operating conditions and dynamically changing electricity prices. At the same time, traditional operation control methods lack effective processing mechanisms for system uncertainties and random disturbances.
[0004] Therefore, a solution is urgently needed to solve the problems existing in the prior art. Summary of the invention
[0005] The embodiments of the present invention provide a method and system for energy storage and optimized operation of a permanent magnet motor based on electricity price, which can at least solve some of the problems existing in the prior art.
[0006] A first aspect of an embodiment of the present invention provides a method for energy storage and optimized operation of a permanent magnet motor based on electricity price, comprising:
[0007] The operation data of the permanent magnet motor is collected and decomposed into multiple intrinsic mode functions through variational mode decomposition. The optimal number of modes is determined based on the signal mutual information and the center frequency of each mode is iteratively updated in combination with particle swarm optimization to obtain the operation feature component and add it to the deep belief neural network. The restricted Boltzmann machine is pre-trained layer by layer through the contrast divergence algorithm and a self-attention module is added to extract the time series feature weights, the power prediction result is output, and a spatiotemporal graph convolutional neural network is constructed to predict the electricity price. A dynamic adjacency matrix is constructed in combination with the spatiotemporal correlation of the historical electricity price data obtained in advance. Multi-scale features are extracted through parallel convolutional layers and long-term and short-term dependencies are determined to obtain the electricity price prediction sequence. The system error is quantified through a Gaussian mixture model, and the model parameters are estimated in combination with variational Bayesian inference. The number of Gaussian distributions is adaptively adjusted through the marginal likelihood function criterion to obtain a multidimensional probability distribution model.
[0008] The power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model are obtained, global optimization is performed through a quantum genetic optimization algorithm, optimization variables are represented by quantum bit encoding, the population is updated in combination with a quantum revolving gate evolution operator, and a quantum entanglement mechanism is introduced to obtain an initial optimization strategy set, the initial optimization strategy set is added to a meta-heuristic hybrid optimization model, strategy search is performed according to adaptive temperature regulation and a dynamic taboo table, experimental solutions are generated in combination with a differential strategy, a timing optimization network based on an attention mechanism is constructed, control sequence features are extracted and timing constraint relationships are established based on a causal attention mechanism, a discrete action layer based on Gobel-Soft maximum sampling is set, a timing association control strategy is obtained, system modeling is performed through a hybrid Gaussian-Hidden Markov model, model parameters are estimated through a variational expectation maximization algorithm, and a probability transfer matrix is obtained according to a particle filter prediction state transition process;
[0009] A two-layer action evaluation control network is constructed according to the timing association control strategy and the probability transfer matrix, and operation control instructions are generated and a control instruction sequence is obtained based on the state entropy. The control instruction sequence is optimized by the dominant action evaluation algorithm and the confidence interval constraint to obtain a smooth control instruction and the intelligent agent is collaboratively controlled through a distributed reinforcement learning framework to obtain a coordinated control strategy. An online learning module is constructed based on an adaptive dynamic programming algorithm and the system value function is approximated according to a radial basis function network. The update step size is determined in combination with the Lyapunov stability analysis to obtain optimized controller parameters. The control scenario features are extracted through a meta-learning-based strategy migration mechanism and a task embedding network. The coordinated control strategy is optimized according to the rapid gradient descent to obtain the optimal control strategy.
[0010] In an optional embodiment,
[0011] The operation data of the permanent magnet motor is collected and decomposed into multiple intrinsic mode functions through variational mode decomposition. The optimal number of modes is determined based on the signal mutual information and the center frequency of each mode is iteratively updated in combination with particle swarm optimization to obtain the operation feature component and add it to the deep belief neural network. The restricted Boltzmann machine is pre-trained layer by layer through the contrast divergence algorithm and a self-attention module is added to extract the time series feature weights, the power prediction result is output, and a spatiotemporal graph convolutional neural network is constructed to predict the electricity price. A dynamic adjacency matrix is constructed by combining the spatiotemporal correlation of the historical electricity price data obtained in advance. Multi-scale features are extracted through parallel convolutional layers and long-term and short-term dependencies are determined to obtain the electricity price prediction sequence. The system error is quantified through a Gaussian mixture model, and the model parameters are estimated in combination with variational Bayesian inference. The number of Gaussian distributions is adaptively adjusted through the marginal likelihood function criterion to obtain a multidimensional probability distribution model including:
[0012] Collecting permanent magnet motor operation data, the permanent magnet motor operation data includes a stator current signal, a stator voltage signal and a rotor speed signal, inputting the permanent magnet motor operation data into a variational mode decomposition module, and decomposing the permanent magnet motor operation data into a plurality of intrinsic mode functions through the variational mode decomposition module, wherein the variational mode decomposition module converts signal decomposition into an optimization problem by constructing an augmented Lagrangian function, iteratively updates each mode function by an alternating direction multiplier method, calculates the signal mutual information value between adjacent intrinsic mode functions, constructs an evaluation index based on the signal mutual information value, and if the difference of the signal mutual information value is less than a preset threshold, determines the corresponding mode number as the optimal mode number;
[0013] The center frequency of each intrinsic mode function is iteratively updated through the particle swarm optimization algorithm. The position and speed of the particle swarm are initialized, and the fitness function is constructed based on the mean square error between the reconstructed signal and the original signal. The position of each particle is expressed as a combination of the center frequencies. The individual optimal solution and the global optimal solution are calculated according to the fitness function. The particle position and speed are dynamically updated according to the inertia weight, the individual learning factor and the social learning factor to obtain the corresponding operating characteristic components of the permanent magnet motor.
[0014] Input the operation feature component into a deep belief neural network, wherein the deep belief neural network is composed of a plurality of stacked restricted Boltzmann machines, the restricted Boltzmann machines are pre-trained by a contrastive divergence algorithm, a network reconstruction sample is obtained by Gibbs sampling, forward propagation and back propagation are alternately performed based on the conditional probability distribution between the visible layer and the hidden layer, the gradients of the weight parameters and the bias parameters are calculated and the network parameters are updated, a self-attention module is constructed in the deep belief neural network, the self-attention module maps the operation feature component into a query matrix, a key matrix and a value matrix, an attention score is obtained by performing a dot product operation of the query matrix and the key matrix and performing softmax normalization, the attention score is multiplied by the value matrix to obtain a weighted feature representation, and a power prediction result is obtained;
[0015] Constructing a spatiotemporal graph convolutional neural network, calculating the correlation coefficient between historical electricity price data at different time points, constructing a dynamic adjacency matrix based on the correlation coefficient, setting a multi-scale parallel convolution layer in the spatiotemporal graph convolutional neural network, the multi-scale parallel convolution layer uses convolution kernels with different receptive fields to extract features, and fuses multi-scale information through feature splicing, setting a graph convolution layer to aggregate spatial features, the graph convolution layer calculates the Laplace matrix based on the dynamic adjacency matrix, and implements graph convolution operations in combination with Chebyshev polynomial approximation, constructing a long short-term memory unit, the long short-term memory unit controls the importance of current input information through an input gate, uses a forget gate to forget historical information, and uses an output gate to adjust the unit state output to achieve long-term dependency modeling, and inputs the historical electricity price data and the dynamic adjacency matrix into the spatiotemporal graph convolutional neural network to generate an electricity price prediction sequence;
[0016] The prediction error of the power prediction result and the prediction error of the electricity price prediction sequence are input into a Gaussian mixture model, and the distribution parameters of the Gaussian mixture model are estimated according to the variational Bayesian inference algorithm. The variational distribution is introduced to approximate the posterior distribution, and a variational lower bound is constructed as the optimization target. The coordinate ascent method is used to iteratively optimize the mean, covariance and mixing weight of the Gaussian components. The model complexity is calculated based on the marginal likelihood function. If the improvement of the marginal likelihood function after adding a new Gaussian component is less than a preset improvement threshold, the Gaussian component is stopped from being added, and a multidimensional probability distribution model of the system prediction error is established.
[0017] In an optional embodiment,
[0018] The distribution parameters of the Gaussian mixture model are estimated according to the variational Bayesian inference algorithm. The variational distribution is introduced to approximate the posterior distribution, and the variational lower bound is constructed as the optimization target. The mean, covariance and mixture weight of the Gaussian components are iteratively optimized using the coordinate ascent method. The model complexity is calculated based on the marginal likelihood function, including:
[0019] Constructing a Gaussian mixture model, taking power prediction error data and electricity price prediction error data as observation data input, normalizing the observation data to obtain standardized observation data, initializing the number of Gaussian components of the Gaussian mixture model, and setting an initial mean parameter, an initial covariance matrix parameter, and an initial mixing weight parameter for each Gaussian component, wherein the initial mean parameter is obtained by adding the mean of the standardized observation data to a random disturbance value, the initial covariance matrix parameter is obtained by multiplying a unit matrix by a preset positive number, and the initial mixing weight parameter is set using a uniform distribution;
[0020] Constructing a variational distribution model, the variational distribution model includes a latent variable Dirichlet distribution module, a mean Gaussian distribution module and a covariance Wishart distribution module, and constructing an optimization objective function based on the divergence between the variational distribution model and the true posterior distribution, the optimization objective function includes an observation data likelihood expectation term and a distribution relative entropy term;
[0021] The parameters of the Gaussian mixture model are iteratively optimized by using the coordinate ascent method, and the parameters are randomly selected in each iterative optimization to be fixed and the distribution parameters of the latent variable Dirichlet distribution module are updated, the posterior probability of the Gaussian component corresponding to each of the standardized observation data is calculated, and the distribution parameters of the mean Gaussian distribution module and the distribution parameters of the covariance Wishart distribution module are updated based on the posterior probability, wherein the distribution parameters of the mean Gaussian distribution module are updated by calculating sufficient statistics and a posterior precision matrix, and the distribution parameters of the covariance Wishart distribution module are updated by accumulating second-order statistics;
[0022] During the parameter updating process, the marginal likelihood function value is calculated, wherein the marginal likelihood function value is used to characterize the fitting degree and complexity of the Gaussian mixture model to obtain the model complexity.
[0023] In an optional embodiment,
[0024] The power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model are obtained, and global optimization is performed through a quantum genetic optimization algorithm. The optimization variables are represented by quantum bit encoding, and the population is updated by combining the quantum revolving gate evolution operator and the quantum entanglement mechanism is introduced to obtain an initial optimization strategy set. The initial optimization strategy set is added to the meta-heuristic hybrid optimization model, and strategy search is performed according to adaptive temperature regulation and dynamic taboo table. The experimental solution is generated by combining the differential strategy, and a timing optimization network based on the attention mechanism is constructed. The control sequence features are extracted and the timing constraint relationship is established based on the causal attention mechanism. A discrete action layer based on the Gobel-Soft maximum sampling is set to obtain a timing association control strategy. The system is modeled by a mixed Gaussian-hidden Markov model, and the model parameters are estimated by a variational expectation maximization algorithm. The probability transfer matrix obtained according to the particle filter prediction state transfer process includes:
[0025] Obtain a power prediction result data set, an electricity price prediction sequence, and a multidimensional probability distribution model, generate quantum bit encoding to represent optimization variables based on a quantum genetic optimization algorithm and initialize the population, describe the quantum state of each quantum bit by setting the amplitude value range to 0 to 1 and the phase angle range to 0 to 2π, and combine multiple quantum bits to construct a chromosome encoding structure;
[0026] Performing a quantum revolving door evolution operation to update the population, calculating the fitness value of each individual in the population and determining the rotation angle of the quantum revolving door based on the fitness value, sorting the individuals in the population according to the fitness value and selecting individuals with fitness values in the top 20% as dominant individuals, selecting local gene fragments of the dominant individuals to perform entangled state superposition operations with other individuals to form a new quantum state, performing a quantum measurement operation on the quantum state to collapse it into a classical solution to construct an initial optimization strategy set;
[0027] Input the initial optimization strategy set into the meta-heuristic hybrid optimization model, establish a temperature adaptive adjustment mechanism to dynamically adjust the temperature parameters, construct a dynamic taboo table to record and update the characteristic values of the visited solution space area, and at the same time screen feasible solutions according to the dynamic taboo table, perform a differential strategy generation operation based on population evolution, select the basis vector in the current population and extract the historical optimal solution, perform a differential operation on the basis vector and the historical optimal solution to generate a differential vector, apply the differential vector to the current solution to generate a test solution and perform boundary constraint correction;
[0028] Construct a multi-layer attention structured timing optimization network and perform sequence feature extraction. Perform embedded coding on the input original control sequence. Extract sequence features of different time scales through multiple attention heads set in parallel. Calculate the similarity scores between the query vector, key vector, and value vector and generate attention weights. Perform mask operations of the causal attention mechanism to filter information after the current moment. Construct an attention weight matrix to represent the temporal dependencies at different moments. Perform concatenation operations on the features output by each attention head and input them to the next network layer through linear transformation.
[0029] Based on the Gubel-Soft maximum sampling, a discretization mapping is performed, the score value of each possible action is calculated, and probability sampling is performed according to the score value to generate the final discrete control action sequence, a mixed Gaussian-hidden Markov model is constructed and multiple discrete hidden states are set, and a corresponding Gaussian distribution model is configured for each hidden state to describe the distribution of observation values, and a Markov transfer relationship between the hidden states is established. The variational expectation maximization algorithm is executed to iteratively optimize the parameters of the mixed Gaussian-hidden Markov model, and the posterior distribution calculation and the likelihood function maximization operation are alternately performed, and the state transition probability matrix and the Gaussian distribution parameters corresponding to each hidden state are output;
[0030] Execute the particle filter to predict the state transfer process, initialize the particle swarm and predict the motion trajectory of each particle based on the state transfer probability, calculate the weight value of each particle according to the observed data, perform the importance resampling operation to screen high-weight particles and update the particle swarm distribution, and summarize the particle swarm distribution to obtain the probability matrix of system state transfer.
[0031] In an optional embodiment,
[0032] Establish a Markov transfer relationship between the hidden states, execute the variational expectation maximization algorithm to iteratively optimize the parameters of the mixed Gaussian-hidden Markov model, alternately perform posterior distribution calculation and likelihood function maximization operations, and output the state transfer probability matrix and the Gaussian distribution parameters corresponding to each hidden state, including:
[0033] Divide multiple discrete hidden states according to system operation characteristics and establish a hidden state set, set the discrete hidden states to high power generation state, medium power generation state, low power generation state, extremely low power generation state and shutdown state, and configure a Gaussian distribution model for each discrete hidden state to describe the probability distribution characteristics of the observation value;
[0034] The data of the acquisition system operation is used to construct a sampling sequence, wherein the sampling interval of the sampling sequence is five minutes, and the sampling sequence includes 288 sampling points. Based on the sampling sequence, the transition frequencies between the hidden states are counted, and a state transition probability matrix is constructed according to the transition frequencies. The dimension of the state transition probability matrix is five times five, and the probability of the system transferring from the current hidden state to the next hidden state is represented by the state transition probability matrix. The state transition probability matrix is normalized so that the sum of the state transition probabilities is one.
[0035] Classify the sampling sequence according to the hidden state category to obtain state classification data, calculate the Gaussian distribution parameters corresponding to each hidden state based on the state classification data, the Gaussian distribution parameters include mean parameters and variance parameters, execute the variational expectation maximization algorithm to optimize the state transition probability matrix and the Gaussian distribution parameters, calculate the posterior probability of each hidden state corresponding to the observed value based on the current parameters, and calculate the marginal probability distribution of the state sequence through the forward-backward recursive algorithm;
[0036] The state transition probability matrix is updated according to the posterior probability, the number of transitions between each pair of hidden states is counted and normalized to obtain the updated transition probability, the Gaussian distribution parameters corresponding to each hidden state are updated based on the posterior probability and the observed value, the posterior probability calculation and parameter update operations are performed alternately until the model parameters converge, and the optimized state transition probability matrix and Gaussian distribution parameters are output.
[0037] In an optional embodiment,
[0038] A two-layer action evaluation control network is constructed according to the timing association control strategy and the probability transfer matrix, and an operation control instruction is generated and a control instruction sequence is obtained based on the state entropy. The control instruction sequence is optimized by the dominant action evaluation algorithm and the confidence interval constraint to obtain a stable control instruction and to coordinately control the intelligent agent through a distributed reinforcement learning framework to obtain a coordinated control strategy. An online learning module is constructed based on an adaptive dynamic programming algorithm and the system value function is approximated according to a radial basis function network. The update step size is determined in combination with Lyapunov stability analysis to obtain optimized controller parameters. The control scene features are extracted through a meta-learning-based strategy migration mechanism and a task embedding network. The coordinated control strategy is optimized according to the rapid descent of the gradient to obtain the optimal control strategy including:
[0039] A double-layer action evaluation control network is constructed according to a timing association control strategy and a probability transfer matrix, and the temperature information, pressure information, and flow information are normalized to obtain a normalized state quantity, and the valve opening information and the motor speed information are normalized to obtain a normalized action quantity. The bottom network of the double-layer action evaluation control network receives the normalized state quantity and the normalized action quantity, extracts features through multiple hidden layers and rectified linear activation functions, and the output layer uses a linear activation function to generate an operation control instruction. The top network of the double-layer action evaluation control network uses a long short-term memory network structure, receives the output sequence of the bottom network at multiple consecutive moments through a memory unit, extracts the timing correlation features of the action sequence, divides the system state space into a temperature sub-interval and a pressure sub-interval, and counts the state distribution frequency of the temperature sub-interval and the pressure sub-interval within a preset time period, calculates the state entropy value based on the state distribution frequency, and screens the operation control instructions according to the state entropy value to obtain a control instruction sequence;
[0040] The cumulative reward value of the current action is calculated by the dominant action evaluation algorithm in the dynamic evaluation window, the dominant value is obtained by performing a difference operation between the cumulative reward value and the average reward value of all actions in the dynamic evaluation window, and the upper and lower limits of the dominant value are set by a confidence interval constraint mechanism to obtain a stable control instruction;
[0041] Construct a distributed reinforcement learning framework, divide the system into a reactor subsystem, a heat exchanger subsystem, and a separator subsystem, configure an intelligent agent for each subsystem to perform state observation and control, and the intelligent agent interactively shares observation data and control experience through a communication network, and collaboratively executes the stable control instructions to obtain a coordinated control strategy;
[0042] An online learning module is constructed based on an adaptive dynamic programming algorithm. A radial basis function network is used to approximate the system value function. An experience replay pool is constructed to store the interactive data of the state vector, action vector, reward value, next state vector, and termination flag. The state samples are clustered by a clustering method to obtain the center point of the basis function. A value function approximation network is constructed based on the center point of the basis function. The system state change rate is obtained in real time and a Lyapunov stability analysis is performed. If the state change rate is greater than a preset change rate threshold, the network parameter update step size is increased, otherwise the network parameter update step size is reduced to obtain the optimized controller parameters.
[0043] A meta-learning-based strategy migration mechanism is constructed. The control scenario features are extracted through a task embedding network. The task embedding network extracts features of the input state through multiple convolutional layers, compresses features through pooling layers, and encodes the compressed features into task embedding vectors through fully connected layers. The coordinated control strategy is optimized using the fast gradient descent method. During the optimization process, the strategy network parameters are adjusted in combination with the task embedding vector, and the momentum term is superimposed to accelerate convergence. The optimal parameter configuration is saved by periodically evaluating the strategy performance to obtain the optimal control strategy.
[0044] In an optional embodiment,
[0045] The state samples are clustered by clustering method to obtain the center point of basis function, and the value function approximation network is constructed based on the center point of basis function, and the system state change rate is obtained in real time and Lyapunov stability analysis is performed, including:
[0046] A hybrid clustering strategy combining density peak clustering and adaptive K-means is used to process state samples. By calculating the local density and distance factor of the points corresponding to the state samples, state sample points whose number of adjacent sample points exceeds 20% of the total samples and whose mutual distance is greater than a preset distance threshold are selected as candidate center points, and the candidate center points are used as the initial clustering centers.
[0047] A dynamic category number adjustment mechanism is introduced into the hybrid clustering strategy. Based on the intra-category sample variance threshold, for each category, if the intra-category sample variance exceeds the intra-category sample variance threshold, the current category is split. If the center distance of the adjacent category corresponding to the current category is less than the preset distance, the adjacent category is merged with the current category to obtain the final cluster center point.
[0048] A multi-scale radial basis function network is constructed based on the final cluster center points, and a kernel width parameter, a small kernel width parameter, and a large kernel width parameter that are positively correlated with the distribution range of the current category samples are configured for each of the final cluster center points to obtain an initial value function network;
[0049] The state samples are predicted by the initial value function network, and the error sequence between the predicted value and the actual value is calculated and added to the sliding time window to obtain the system state change rate sequence, and the statistical characteristics of the system state change rate sequence are calculated according to the sliding time window, wherein the statistical characteristics include the change rate mean and the change rate variance;
[0050] A Lyapunov function is constructed based on the statistical characteristics, the change rate mean and the change rate variance are used as input variables of the Lyapunov function, and the ratio of the difference between the Lyapunov function values at two adjacent sampling moments and the sampling period is calculated to obtain the Lyapunov function change rate.
[0051] A second aspect of an embodiment of the present invention provides a permanent magnet motor energy storage and optimized operation system based on electricity price, comprising:
[0052] The first unit is used to collect the operation data of the permanent magnet motor and decompose the operation data of the permanent magnet motor into multiple intrinsic mode functions through variational mode decomposition, determine the optimal number of modes based on signal mutual information and iteratively update the center frequency of each mode in combination with particle swarm optimization, obtain the operation feature component and add it to the deep belief neural network, pre-train the restricted Boltzmann machine layer by layer through the contrast divergence algorithm and add the self-attention module to extract the time series feature weight, output the power prediction result, construct a spatiotemporal graph convolutional neural network for electricity price prediction, construct a dynamic adjacency matrix in combination with the spatiotemporal correlation of the historical electricity price data obtained in advance, extract multi-scale features through parallel convolutional layers and determine long-term and short-term dependencies to obtain the electricity price prediction sequence, quantify the system error through the Gaussian mixture model, estimate the model parameters in combination with variational Bayesian inference, and adaptively adjust the number of Gaussian distributions through the marginal likelihood function criterion to obtain a multidimensional probability distribution model;
[0053] The second unit is used to obtain the power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model, perform global optimization through a quantum genetic optimization algorithm, represent optimization variables through quantum bit encoding, update the population in combination with a quantum revolving gate evolution operator and introduce a quantum entanglement mechanism to obtain an initial optimization strategy set, add the initial optimization strategy set to a meta-heuristic hybrid optimization model, perform strategy search according to adaptive temperature regulation and a dynamic taboo table, generate experimental solutions in combination with a differential strategy, construct a timing optimization network based on an attention mechanism, extract control sequence features and establish timing constraint relationships based on a causal attention mechanism, set a discrete action layer based on Gobel-Soft maximum sampling, obtain a timing-related control strategy, perform system modeling through a hybrid Gaussian-Hidden Markov model, estimate model parameters through a variational expectation maximization algorithm, and obtain a probability transfer matrix based on a particle filter prediction state transition process;
[0054] The third unit is used to construct a two-layer action evaluation control network according to the timing association control strategy and the probability transfer matrix, generate operation control instructions and obtain a control instruction sequence based on state entropy, optimize the control instruction sequence through the dominant action evaluation algorithm and the confidence interval constraint, obtain smooth control instructions, and coordinate the intelligent body through a distributed reinforcement learning framework to obtain a coordinated control strategy, build an online learning module based on an adaptive dynamic programming algorithm and approximate the system value function according to a radial basis function network, determine the update step size in combination with Lyapunov stability analysis, obtain optimized controller parameters, extract control scenario features through a meta-learning-based strategy migration mechanism and a task embedding network, optimize the coordinated control strategy according to the rapid descent of the gradient, and obtain the optimal control strategy.
[0055] According to a third aspect of the embodiments of the present invention,
[0056] An electronic device is provided, comprising:
[0057] processor;
[0058] a memory for storing processor-executable instructions;
[0059] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0060] A fourth aspect of the embodiments of the present invention is:
[0061] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0062] In the present invention, by combining deep learning models and probability models, and introducing methods such as self-attention mechanism and variational Bayesian inference, the power output of the permanent magnet motor and future electricity price fluctuations can be predicted more accurately, providing a reliable data basis for optimizing the operation strategy. By using advanced optimization algorithms such as quantum genetic optimization algorithm, meta-heuristic hybrid optimization model and timing optimization network based on attention mechanism, combined with system modeling (hybrid Gaussian-hidden Markov model) and state prediction (particle filtering), a better motor operation control strategy can be formulated to achieve intelligent management of energy storage and release. Through the distributed reinforcement learning framework, adaptive dynamic programming algorithm and meta-learning-based strategy migration mechanism, the control system can learn online, adaptively adjust, and respond to changes in different operation scenarios, thereby improving the stability and adaptability of the system, and ultimately achieving the best effect of permanent magnet motor energy storage and optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1It is a flow chart of a method for energy storage and optimized operation of a permanent magnet motor based on electricity price according to an embodiment of the present invention;
[0064] Figure 2 It is a structural schematic diagram of a permanent magnet motor energy storage and optimized operation system based on electricity price according to an embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0067] Figure 1 FIG. 1 is a flow chart of a method for storing and optimizing the operation of a permanent magnet motor based on electricity price according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0068] S1. Collect the operation data of the permanent magnet motor and decompose the operation data of the permanent magnet motor into multiple intrinsic mode functions through variational mode decomposition, determine the optimal mode number based on the signal mutual information and iteratively update the center frequency of each mode in combination with particle swarm optimization, obtain the operation feature component and add it to the deep belief neural network, pre-train the restricted Boltzmann machine layer by layer through the contrast divergence algorithm and add the self-attention module to extract the time series feature weight, output the power prediction result, construct a spatiotemporal graph convolutional neural network for electricity price prediction, construct a dynamic adjacency matrix based on the spatiotemporal correlation of the historical electricity price data obtained in advance, extract multi-scale features through parallel convolutional layers and determine the long-term and short-term dependencies, obtain the electricity price prediction sequence, quantify the system error through the Gaussian mixture model, estimate the model parameters in combination with variational Bayesian inference, and adaptively adjust the number of Gaussian distributions through the marginal likelihood function criterion to obtain a multidimensional probability distribution model;
[0069] The signal mutual information is a measure of the correlation and shared information between two signals or random variables. The information transmission and dependency are evaluated by calculating the difference between the joint distribution and marginal distribution of the signal. The optimal modal number refers to the selection of the best distribution or number of categories in data analysis and pattern recognition to optimally describe the structure of the data set. The deep belief neural network is a deep learning model that combines multiple levels of neural networks and probabilistic models for automatic feature learning and unsupervised training. The contrastive divergence algorithm is an approximate algorithm in restricted Boltzmann machine training, which is used to estimate the log-likelihood of data. The restricted Boltzmann machine is an energy model used for unsupervised learning and feature extraction in deep learning. It consists of a visible layer and a hidden layer. The implicit features of the data are learned by adjusting the weights of the connections between the two layers. The Gaussian mixture model is a probabilistic model used to represent a data set with multiple Gaussian distribution components. It simulates complex distribution structures by mixing multiple Gaussian distributions and is widely used in cluster analysis, density estimation and data modeling. The variational Bayesian inference is an approximate inference method used to infer the posterior distribution in complex Bayesian models. By optimizing the variational lower bound, variational Bayesian inference can effectively approximate the probability distribution in high-dimensional space. The marginal likelihood function criterion is a criterion for model selection and evaluation. It evaluates the fit of the model by calculating the marginal likelihood of the data (that is, the likelihood value of the model to the data), and is often used in Bayesian models and generative models.
[0070] In an optional embodiment,
[0071] The operation data of the permanent magnet motor is collected and decomposed into multiple intrinsic mode functions through variational mode decomposition. The optimal number of modes is determined based on the signal mutual information and the center frequency of each mode is iteratively updated in combination with particle swarm optimization to obtain the operation feature component and add it to the deep belief neural network. The restricted Boltzmann machine is pre-trained layer by layer through the contrast divergence algorithm and a self-attention module is added to extract the time series feature weights, the power prediction result is output, and a spatiotemporal graph convolutional neural network is constructed to predict the electricity price. A dynamic adjacency matrix is constructed by combining the spatiotemporal correlation of the historical electricity price data obtained in advance. Multi-scale features are extracted through parallel convolutional layers and long-term and short-term dependencies are determined to obtain the electricity price prediction sequence. The system error is quantified through a Gaussian mixture model, and the model parameters are estimated in combination with variational Bayesian inference. The number of Gaussian distributions is adaptively adjusted through the marginal likelihood function criterion to obtain a multidimensional probability distribution model including:
[0072] Collecting permanent magnet motor operation data, the permanent magnet motor operation data includes a stator current signal, a stator voltage signal and a rotor speed signal, inputting the permanent magnet motor operation data into a variational mode decomposition module, and decomposing the permanent magnet motor operation data into a plurality of intrinsic mode functions through the variational mode decomposition module, wherein the variational mode decomposition module converts signal decomposition into an optimization problem by constructing an augmented Lagrangian function, iteratively updates each mode function by an alternating direction multiplier method, calculates the signal mutual information value between adjacent intrinsic mode functions, constructs an evaluation index based on the signal mutual information value, and if the difference of the signal mutual information value is less than a preset threshold, determines the corresponding mode number as the optimal mode number;
[0073] The center frequency of each intrinsic mode function is iteratively updated through the particle swarm optimization algorithm. The position and speed of the particle swarm are initialized, and the fitness function is constructed based on the mean square error between the reconstructed signal and the original signal. The position of each particle is expressed as a combination of the center frequencies. The individual optimal solution and the global optimal solution are calculated according to the fitness function. The particle position and speed are dynamically updated according to the inertia weight, the individual learning factor and the social learning factor to obtain the corresponding operating characteristic components of the permanent magnet motor.
[0074] Input the operation feature component into a deep belief neural network, wherein the deep belief neural network is composed of a plurality of stacked restricted Boltzmann machines, the restricted Boltzmann machines are pre-trained by a contrastive divergence algorithm, a network reconstruction sample is obtained by Gibbs sampling, forward propagation and back propagation are alternately performed based on the conditional probability distribution between the visible layer and the hidden layer, the gradients of the weight parameters and the bias parameters are calculated and the network parameters are updated, a self-attention module is constructed in the deep belief neural network, the self-attention module maps the operation feature component into a query matrix, a key matrix and a value matrix, an attention score is obtained by performing a dot product operation of the query matrix and the key matrix and performing softmax normalization, the attention score is multiplied by the value matrix to obtain a weighted feature representation, and a power prediction result is obtained;
[0075] Constructing a spatiotemporal graph convolutional neural network, calculating the correlation coefficient between historical electricity price data at different time points, constructing a dynamic adjacency matrix based on the correlation coefficient, setting a multi-scale parallel convolution layer in the spatiotemporal graph convolutional neural network, the multi-scale parallel convolution layer uses convolution kernels with different receptive fields to extract features, and fuses multi-scale information through feature splicing, setting a graph convolution layer to aggregate spatial features, the graph convolution layer calculates the Laplace matrix based on the dynamic adjacency matrix, and implements graph convolution operations in combination with Chebyshev polynomial approximation, constructing a long short-term memory unit, the long short-term memory unit controls the importance of current input information through an input gate, uses a forget gate to forget historical information, and uses an output gate to adjust the unit state output to achieve long-term dependency modeling, and inputs the historical electricity price data and the dynamic adjacency matrix into the spatiotemporal graph convolutional neural network to generate an electricity price prediction sequence;
[0076] The prediction error of the power prediction result and the prediction error of the electricity price prediction sequence are input into a Gaussian mixture model, and the distribution parameters of the Gaussian mixture model are estimated according to the variational Bayesian inference algorithm. The variational distribution is introduced to approximate the posterior distribution, and a variational lower bound is constructed as the optimization target. The coordinate ascent method is used to iteratively optimize the mean, covariance and mixing weight of the Gaussian components. The model complexity is calculated based on the marginal likelihood function. If the improvement of the marginal likelihood function after adding a new Gaussian component is less than a preset improvement threshold, the Gaussian component is stopped from being added, and a multidimensional probability distribution model of the system prediction error is established.
[0077] The augmented Lagrangian function is an optimization method that transforms a constrained optimization problem into an unconstrained problem by introducing Lagrangian multipliers and augmented terms so as to optimize and solve it under the constraints. The alternating direction multiplier method is an optimization algorithm used to solve optimization problems with constraints by decomposing the problem into more tractable sub-problems and solving each sub-problem alternately until it converges to a global solution. The Gibbs sampling is a Markov chain Monte Carlo (MCMC) method used to sample from complex probability distributions. The Chebyshev polynomial is an orthogonal polynomial used for function approximation and numerical calculation, which is widely used in approximation theory and numerical analysis. The variational lower bound is a concept in variational inference, which means that the lower bound approximation of the objective function is performed during the optimization process. The coordinate ascent method is an optimization algorithm used to solve high-dimensional optimization problems. It gradually optimizes one variable by fixing other variables until all variables reach the optimal solution. It is usually used to optimize objective functions with separated variables.
[0078] The operation data of the permanent magnet motor is collected, including the stator current signal, the stator voltage signal and the rotor speed signal. For example, the data of these three signals can be collected every 0.01 seconds to form a time series data set.
[0079] The collected permanent magnet motor operation data is input into the variational modal decomposition module. The signal decomposition problem is converted into an optimization problem, and each modal function is iteratively updated using the alternating direction multiplier method, so that the operation data is decomposed into multiple intrinsic mode functions. For example, the original signal can be decomposed into 8 intrinsic mode functions. In order to determine the optimal number of modes, it is necessary to calculate the signal mutual information value between adjacent intrinsic mode functions. If the difference between the signal mutual information values of adjacent intrinsic mode functions is less than a preset threshold (for example, 0.05), the corresponding number of modes is determined as the optimal number of modes.
[0080] The particle swarm optimization algorithm is used to iteratively update the center frequency of each intrinsic mode function, and a group of particles are initialized. The position of each particle represents a set of center frequencies. The fitness function is defined as the mean square error between the reconstructed signal and the original signal. By iteratively updating the position and speed of the particles, the best combination of center frequencies is found, thereby obtaining the operating characteristic components of the permanent magnet motor. For example, the position of each particle is an 8-dimensional vector, representing the center frequencies of 8 intrinsic mode functions.
[0081] The obtained running feature components are input into the deep belief neural network for power prediction. The deep belief neural network is composed of multiple stacked restricted Boltzmann machines. Each restricted Boltzmann machine is pre-trained using the contrastive divergence algorithm, and the network reconstruction samples are obtained through Gibbs sampling. Forward and backward propagation is performed through the conditional probability distribution between the visible layer and the hidden layer to calculate and update the network parameters. A self-attention module is added to the deep belief neural network, which maps the running feature components to the query matrix, key matrix, and value matrix. The attention score is obtained by calculating the dot product of the query matrix and the key matrix and performing softmax normalization. The attention score is multiplied by the value matrix to obtain the weighted feature representation, and the power prediction result is output. For example, the deep belief neural network can be composed of 3 stacked restricted Boltzmann machines, the number of visible layer units of each restricted Boltzmann machine is the dimension of the running feature component, and the number of hidden layer units can be set to 128, 64, and 32.
[0082] Construct a spatiotemporal graph convolutional neural network for electricity price forecasting. First, calculate the correlation coefficient between historical electricity price data at different time points, and build a dynamic adjacency matrix based on this. The spatiotemporal graph convolutional neural network contains multiple parallel convolutional layers, which use convolution kernels of different sizes to extract multi-scale features and splice and fuse these features. The graph convolution layer performs graph convolution operations based on the Laplacian matrix calculated by the dynamic adjacency matrix to aggregate spatial features. The spatiotemporal graph convolutional neural network also contains long short-term memory units, which control the flow of information through input gates, forget gates, and output gates to achieve long-term dependency modeling. Input the historical electricity price data and the dynamic adjacency matrix into the spatiotemporal graph convolutional neural network to generate an electricity price forecast sequence. For example, the historical electricity price data can be the electricity price data for the past 24 hours.
[0083] The prediction error of the power prediction result and the prediction error of the electricity price prediction sequence are input into the Gaussian mixture model. The distribution parameters of the Gaussian mixture model are estimated using the variational Bayesian inference algorithm, and the number of Gaussian components is adaptively adjusted through the marginal likelihood function criterion to establish a multidimensional probability distribution model of the system prediction error. For example, the initial setting of the Gaussian mixture model contains 3 Gaussian components. If the increase in the marginal likelihood function after adding a new Gaussian component is less than a preset threshold (for example, 0.01), stop adding Gaussian components.
[0084] In this embodiment, by combining the permanent magnet motor operation data and historical electricity price data and adopting advanced deep learning models, future power and electricity prices can be predicted more accurately. Through variational mode decomposition and self-attention mechanism, operation characteristics and timing characteristics can be effectively extracted to improve the robustness of the model. Through Gaussian mixture model and variational Bayesian inference, system errors can be quantified and more reliable probability distribution prediction results can be provided.
[0085] In an optional embodiment,
[0086] The distribution parameters of the Gaussian mixture model are estimated according to the variational Bayesian inference algorithm. The variational distribution is introduced to approximate the posterior distribution, and the variational lower bound is constructed as the optimization target. The mean, covariance and mixture weight of the Gaussian components are iteratively optimized using the coordinate ascent method. The model complexity is calculated based on the marginal likelihood function, including:
[0087] Constructing a Gaussian mixture model, taking power prediction error data and electricity price prediction error data as observation data input, normalizing the observation data to obtain standardized observation data, initializing the number of Gaussian components of the Gaussian mixture model, and setting an initial mean parameter, an initial covariance matrix parameter, and an initial mixing weight parameter for each Gaussian component, wherein the initial mean parameter is obtained by adding the mean of the standardized observation data to a random disturbance value, the initial covariance matrix parameter is obtained by multiplying a unit matrix by a preset positive number, and the initial mixing weight parameter is set using a uniform distribution;
[0088] Constructing a variational distribution model, the variational distribution model includes a latent variable Dirichlet distribution module, a mean Gaussian distribution module and a covariance Wishart distribution module, and constructing an optimization objective function based on the divergence between the variational distribution model and the true posterior distribution, the optimization objective function includes an observation data likelihood expectation term and a distribution relative entropy term;
[0089] The parameters of the Gaussian mixture model are iteratively optimized by using the coordinate ascent method, and the parameters are randomly selected in each iterative optimization to be fixed and the distribution parameters of the latent variable Dirichlet distribution module are updated, the posterior probability of the Gaussian component corresponding to each of the standardized observation data is calculated, and the distribution parameters of the mean Gaussian distribution module and the distribution parameters of the covariance Wishart distribution module are updated based on the posterior probability, wherein the distribution parameters of the mean Gaussian distribution module are updated by calculating sufficient statistics and a posterior precision matrix, and the distribution parameters of the covariance Wishart distribution module are updated by accumulating second-order statistics;
[0090] During the parameter updating process, the marginal likelihood function value is calculated, wherein the marginal likelihood function value is used to characterize the fitting degree and complexity of the Gaussian mixture model to obtain the model complexity.
[0091] The latent variable Dirichlet distribution module refers to the use of Dirichlet distribution as the prior distribution of latent variables in the Bayesian model. The mean Gaussian distribution module is a module used to describe data distribution, in which data points are assumed to come from a Gaussian distribution, whose mean and covariance are parameters to be estimated. The covariance Wishart distribution module refers to the use of Wishart distribution to model the prior distribution of the covariance matrix in Bayesian modeling. The Wishart distribution is the conjugate prior of the covariance matrix and is widely used in parameter inference of multivariate normal distribution. The distribution relative entropy term is a measure of the difference between two probability distributions and is typically used for model optimization in information theory. The second-order statistics are statistics that describe correlation in random processes or signals, typically referring to the variance, covariance, etc. of the signal, and are widely used in signal processing, time series analysis, and machine learning.
[0092] The power forecast error data and electricity price forecast error data over a period of time are collected and standardized, for example, using the Z-score standardization method, that is, subtracting the mean from the data and dividing it by the standard deviation, so that the mean is 0 and the standard deviation is 1.
[0093] Determine the number of Gaussian components in the Gaussian mixture model, for example, initially set to 3 components. Set initial parameters for each component: mean, covariance matrix, and mixture weights. The initial mean can be set to the mean of the standardized observations plus some small random perturbations to avoid all components starting from the same point. The initial covariance matrix can be set to the identity matrix multiplied by a preset positive number, such as 1. The initial mixture weights can be set to a uniform distribution, for example, each component weight is 1 / 3.
[0094] The variational distribution is introduced to approximate the true posterior distribution. The variational distribution consists of three modules: the latent variable Dirichlet distribution module, the mean Gaussian distribution module, and the covariance Wishart distribution module. The Dirichlet distribution is used to model the posterior distribution of the mixture weights, the Gaussian distribution is used to model the posterior distribution of the mean, and the Wishart distribution is used to model the posterior distribution of the covariance matrix. An optimization objective function is constructed, which includes: the observed data likelihood expectation term and the distribution relative entropy term. The observed data likelihood expectation term measures the degree of fit of the model to the observed data, and the distribution relative entropy term measures the difference between the variational distribution and the true posterior distribution. The goal is to maximize this optimization objective function, also known as the variational lower bound.
[0095] The parameters of the Gaussian mixture model are iteratively optimized using the coordinate ascent method. In each iteration, a parameter is randomly selected to be fixed, such as the mean parameter, and the distribution parameters of the latent variable Dirichlet distribution module are updated. The posterior probability, also known as the responsibility, of each standardized observation corresponding to each Gaussian component is calculated. Based on the calculated responsibility, the distribution parameters of the mean Gaussian distribution module and the distribution parameters of the covariance Wishart distribution module are updated. The distribution parameter update of the mean Gaussian distribution module is achieved by calculating the sufficient statistics and the posterior precision matrix. The distribution parameter update of the covariance Wishart distribution module is achieved by accumulating second-order statistics.
[0096] During the parameter update process, the marginal likelihood function value is calculated. The marginal likelihood function value is used to characterize the degree of fit and complexity of the Gaussian mixture model. By comparing the marginal likelihood function values under different numbers of Gaussian components, the most appropriate model complexity can be selected. For example, suppose there are 1000 power prediction error data and 1000 electricity price prediction error data. After standardization, the above method is used for iterative optimization. By comparing the marginal likelihood function values under different numbers of Gaussian components (such as 2, 3, and 4 components), the number of components with the largest marginal likelihood function value is selected as the number of Gaussian components of the final model.
[0097] In this embodiment, the variational Bayesian Gaussian mixture model can more accurately capture the complex distribution characteristics of power prediction error and electricity price prediction error, thereby improving prediction accuracy and reducing prediction risk. The marginal likelihood function provides a comprehensive indicator of model fit goodness and complexity, which can effectively avoid overfitting or underfitting and select the most appropriate model complexity. The coordinate ascent method can effectively optimize the parameters of the Gaussian mixture model and has higher computational efficiency than other optimization methods, especially when the amount of data is large.
[0098] S2. Obtain the power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model, perform global optimization through a quantum genetic optimization algorithm, represent optimization variables through quantum bit encoding, update the population in combination with a quantum revolving gate evolution operator and introduce a quantum entanglement mechanism to obtain an initial optimization strategy set, add the initial optimization strategy set to a meta-heuristic hybrid optimization model, perform strategy search based on adaptive temperature regulation and a dynamic taboo table, generate experimental solutions in combination with a differential strategy, construct a timing optimization network based on an attention mechanism, extract control sequence features and establish timing constraint relationships based on a causal attention mechanism, set a discrete action layer based on Gobel-Soft maximum sampling, obtain a timing-related control strategy, perform system modeling through a hybrid Gaussian-Hidden Markov model, estimate model parameters through a variational expectation maximization algorithm, and obtain a probability transfer matrix based on a particle filter prediction state transition process;
[0099] The quantum genetic optimization algorithm is an optimization method that combines quantum computing with genetic algorithms. It uses the parallelism of quantum bits and the advantages of quantum computing to accelerate the search process of traditional genetic algorithms. It enhances the exploration ability of genetic algorithms in complex search spaces through quantum gates and quantum superposition. The quantum bit encoding is a key step in the quantum genetic optimization algorithm, which is used to encode the genetic information in the classical genetic algorithm into the state of quantum bits. Quantum bits can represent multiple solutions at the same time, thereby improving search efficiency. The quantum rotation gate evolution operator is a quantum operation used to update the state of quantum bits in the quantum genetic optimization algorithm. By rotating the phase of the quantum bit, the evolution operator can The search process is guided to approach the optimization goal under the framework of quantum computing. The quantum entanglement mechanism is a phenomenon in quantum computing, which refers to the strong correlation between the states of multiple quantum bits. Through quantum entanglement, the quantum genetic algorithm can simultaneously consider the potential combinations of multiple solutions, thereby improving the search capability. The meta-heuristic hybrid optimization model is a hybrid optimization method that combines meta-heuristic algorithms with other optimization techniques. By flexibly combining different optimization strategies, it can achieve better search results in a variety of optimization tasks and has strong adaptability and global search capabilities. The dynamic taboo table is a component in the taboo search algorithm, which is used to prevent the algorithm from falling into a local optimal solution during the search process. The taboo table is dynamically updated to record the visited solutions or solution spaces to prevent repeated searches and enhance the diversity of searches. The causal attention mechanism is an attention mechanism used to capture causal relationships in deep learning models. It is particularly suitable for processing time series data or data with causal relationships. It helps the model make predictions or decisions more accurately by weighting the causal influence of different input information. The Gubel-Soft maximum sampling is a sampling method for generating samples, which is usually used in optimization and reasoning processes. The mixed Gaussian-hidden Markov model is a hybrid model that combines Gaussian distribution and hidden Markov model, which is often used in time series data analysis, speech recognition and other tasks. The variational expectation maximization algorithm is an inference algorithm for complex probability models. It combines variational inference and expectation maximization algorithms to approximate the true posterior distribution by maximizing the variational lower bound. It is widely used in unsupervised learning and Bayesian inference.
[0100] In an optional embodiment,
[0101] The power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model are obtained, and global optimization is performed through a quantum genetic optimization algorithm. The optimization variables are represented by quantum bit encoding, and the population is updated by combining the quantum revolving gate evolution operator and the quantum entanglement mechanism is introduced to obtain an initial optimization strategy set. The initial optimization strategy set is added to the meta-heuristic hybrid optimization model, and strategy search is performed according to adaptive temperature regulation and dynamic taboo table. The experimental solution is generated by combining the differential strategy, and a timing optimization network based on the attention mechanism is constructed. The control sequence features are extracted and the timing constraint relationship is established based on the causal attention mechanism. A discrete action layer based on the Gobel-Soft maximum sampling is set to obtain a timing association control strategy. The system is modeled by a mixed Gaussian-hidden Markov model, and the model parameters are estimated by a variational expectation maximization algorithm. The probability transfer matrix obtained according to the particle filter prediction state transfer process includes:
[0102] Obtain a power prediction result data set, an electricity price prediction sequence, and a multidimensional probability distribution model, generate quantum bit encoding to represent optimization variables based on a quantum genetic optimization algorithm and initialize the population, describe the quantum state of each quantum bit by setting the amplitude value range to 0 to 1 and the phase angle range to 0 to 2π, and combine multiple quantum bits to construct a chromosome encoding structure;
[0103] Performing a quantum revolving door evolution operation to update the population, calculating the fitness value of each individual in the population and determining the rotation angle of the quantum revolving door based on the fitness value, sorting the individuals in the population according to the fitness value and selecting individuals with fitness values in the top 20% as dominant individuals, selecting local gene fragments of the dominant individuals to perform entangled state superposition operations with other individuals to form a new quantum state, performing a quantum measurement operation on the quantum state to collapse it into a classical solution to construct an initial optimization strategy set;
[0104] Input the initial optimization strategy set into the meta-heuristic hybrid optimization model, establish a temperature adaptive adjustment mechanism to dynamically adjust the temperature parameters, construct a dynamic taboo table to record and update the characteristic values of the visited solution space area, and at the same time screen feasible solutions according to the dynamic taboo table, perform a differential strategy generation operation based on population evolution, select the basis vector in the current population and extract the historical optimal solution, perform a differential operation on the basis vector and the historical optimal solution to generate a differential vector, apply the differential vector to the current solution to generate a test solution and perform boundary constraint correction;
[0105] Construct a multi-layer attention structured timing optimization network and perform sequence feature extraction. Perform embedded coding on the input original control sequence. Extract sequence features of different time scales through multiple attention heads set in parallel. Calculate the similarity scores between the query vector, key vector, and value vector and generate attention weights. Perform mask operations of the causal attention mechanism to filter information after the current moment. Construct an attention weight matrix to represent the temporal dependencies at different moments. Perform concatenation operations on the features output by each attention head and input them to the next network layer through linear transformation.
[0106] Based on the Gubel-Soft maximum sampling, a discretization mapping is performed, the score value of each possible action is calculated, and probability sampling is performed according to the score value to generate the final discrete control action sequence, a mixed Gaussian-hidden Markov model is constructed and multiple discrete hidden states are set, and a corresponding Gaussian distribution model is configured for each hidden state to describe the distribution of observation values, and a Markov transfer relationship between the hidden states is established. The variational expectation maximization algorithm is executed to iteratively optimize the parameters of the mixed Gaussian-hidden Markov model, and the posterior distribution calculation and the likelihood function maximization operation are alternately performed, and the state transition probability matrix and the Gaussian distribution parameters corresponding to each hidden state are output;
[0107] Execute the particle filter to predict the state transfer process, initialize the particle swarm and predict the motion trajectory of each particle based on the state transfer probability, calculate the weight value of each particle according to the observed data, perform the importance resampling operation to screen high-weight particles and update the particle swarm distribution, and summarize the particle swarm distribution to obtain the probability matrix of system state transfer.
[0108] The dominant individual is a concept in genetic algorithms, which refers to the individuals with the best performance in the current population. These individuals are usually copied or mutated in the next generation to ensure the transmission of excellent genes. The local gene fragment is a part of the gene structure in the genetic algorithm, which is usually composed of a local fragment in the gene sequence of the individual. The collapse is a concept in quantum computing, which refers to the collapse of a quantum system from a superposition state to a certain state. The population evolution is a process in genetic algorithms, which means that the algorithm continuously updates the genetic information of the population in each generation through operations such as selection, crossover, and mutation, thereby guiding the search process towards the direction of the global optimal solution. The differential strategy generation operation is an operation used in genetic algorithms, which is usually used to generate new solutions or populations.
[0109] Collect power data and corresponding electricity price data for a period of time from historical data, such as data from the past year. The power data can be the power generated per hour, and the electricity price data can be the electricity price per hour. The multidimensional probability distribution model is used to describe the relationship between power and electricity price. A Gaussian mixture model or other suitable probability models can be used. For example, the historical data of power and electricity price can be input into the Gaussian mixture model for training to obtain model parameters. Assume that the power prediction result is [10, 12, 15, 13, 11] kW, and the electricity price prediction sequence is [0.5, 0.6, 0.7, 0.6, 0.5] yuan / kWh.
[0110] Based on the quantum genetic optimization algorithm, quantum bit encoding is generated to represent the optimization variables and the population is initialized. Each optimization variable, such as the output of a generator, can be encoded with a certain number of quantum bits. For example, using 4 quantum bits to encode an optimization variable can represent 16 different output levels. When the population is initialized, the amplitude and phase angle of each quantum bit are randomly set between 0 and 1 and 0 and 2π. For example, the initial state of a quantum bit can be represented as an amplitude of 0.7 and a phase angle of π / 2. Assuming the population size is 10, 10 such chromosome encoding structures need to be initialized.
[0111] Perform the quantum revolving gate evolution operation to update the population. According to the fitness value of each individual, adjust the rotation angle of the quantum revolving gate. The fitness value can be obtained by calculating the benefit of the strategy represented by the individual, for example, the benefit can be the profit of power generation. The higher the fitness value of the individual, the smaller the rotation angle of the corresponding quantum revolving gate is to maintain its dominant gene. Sort the individuals in the population according to the fitness value, and select the top 20% of the individuals as the dominant individuals. For example, if the fitness value of the individual with the highest fitness value is 100, its rotation angle can be set to 0.1 radians.
[0112] Perform entanglement superposition operations on the local gene fragments of the dominant individual and other individuals to form a new quantum state. For example, the second and third qubits of the individual with the highest fitness value are entangled with the corresponding qubits of another individual. Perform quantum measurement operations on the new quantum state and collapse it into a classical solution to construct an initial optimization strategy set. For example, after measurement, a quantum state collapses to [0, 1, 1, 0], indicating a specific output level. Assume that 5 initial optimization strategies are finally obtained.
[0113] The initial optimization strategy set is input into the metaheuristic hybrid optimization model. A temperature adaptive adjustment mechanism and a dynamic taboo table are established. The temperature parameter is dynamically adjusted according to the current search state. For example, if the search is trapped in a local optimum, the temperature parameter is increased. The dynamic taboo table records the eigenvalues of the visited solution space region to avoid repeated searches. For example, the eigenvalues of a solution, such as the combination of power generation and electricity price, are recorded in the taboo table.
[0114] Perform differential strategy generation based on population evolution. Select the basis vectors and historical optimal solutions in the current population, perform differential operations to generate differential vectors. Apply the differential vector to the current solution to generate a trial solution, and perform boundary constraint correction to ensure that the trial solution is within the feasible range. For example, if the current solution is 10kW and the differential vector is 2kW, the trial solution is 12kW.
[0115] Construct a multi-layer attention structured timing optimization network and perform sequence feature extraction. Perform embedded coding on the input control sequence and extract sequence features of different time scales through multiple attention heads. Calculate the similarity scores between the query vector, key vector, and value vector to generate attention weights. Execute the causal attention mechanism to filter information after the current moment. For example, if the current moment is t, only the control sequence information before and after t is considered. Concatenate the features output by each attention head and input them to the next network layer through linear transformation.
[0116] Discretization mapping is performed based on Gubel-Soft maximum sampling. The score value of each possible action is calculated, and probability sampling is performed based on the score value to generate the final discrete control action sequence. For example, if the action "increase output" has the highest score, it is more likely to be selected.
[0117] Construct a mixed Gaussian-hidden Markov model. Configure a corresponding Gaussian distribution model for each hidden state. Perform a variational expectation maximization algorithm to iteratively optimize the model parameters. For example, suppose there are 2 hidden states, representing "high demand" and "low demand", and each state corresponds to a Gaussian distribution that describes the distribution of power.
[0118] Execute the particle filter prediction state transition process. Initialize the particle swarm and predict the trajectory of each particle based on the state transition probability. Calculate the weight value of each particle based on the observed data. Perform importance resampling, filter high-weight particles and update the particle swarm distribution. For example, if the predicted power of a particle is closer to the actual power, it is given a higher weight. Summarize the particle swarm distribution and obtain the probability matrix of the system state transition. For example, the probability of transitioning from the "high demand" state to the "low demand" state can be obtained.
[0119] In this embodiment, the parallelism and quantum entanglement mechanism of the quantum genetic algorithm, combined with the meta-heuristic hybrid optimization model, can search the solution space more quickly and find a better control strategy. It is more efficient than the traditional optimization algorithm. The combination of the hybrid Gaussian-Hidden Markov model and the particle filter algorithm can more accurately capture the dynamic changes of the system state and improve the prediction accuracy of the state transition probability matrix, thereby making the control strategy more adaptable. The timing optimization network of the attention mechanism can extract the timing characteristics and causal relationships of the control sequence. Combined with the Gubel-Soft maximum sampling, the generated control strategy can better adapt to the complex power system environment and improve the robustness of the control strategy.
[0120] In an optional embodiment,
[0121] Establish a Markov transfer relationship between the hidden states, execute the variational expectation maximization algorithm to iteratively optimize the parameters of the mixed Gaussian-hidden Markov model, alternately perform posterior distribution calculation and likelihood function maximization operations, and output the state transfer probability matrix and the Gaussian distribution parameters corresponding to each hidden state, including:
[0122] Divide multiple discrete hidden states according to system operation characteristics and establish a hidden state set, set the discrete hidden states to high power generation state, medium power generation state, low power generation state, extremely low power generation state and shutdown state, and configure a Gaussian distribution model for each discrete hidden state to describe the probability distribution characteristics of the observation value;
[0123] The data of the acquisition system operation is used to construct a sampling sequence, wherein the sampling interval of the sampling sequence is five minutes, and the sampling sequence includes 288 sampling points. Based on the sampling sequence, the transition frequencies between the hidden states are counted, and a state transition probability matrix is constructed according to the transition frequencies. The dimension of the state transition probability matrix is five times five, and the probability of the system transferring from the current hidden state to the next hidden state is represented by the state transition probability matrix. The state transition probability matrix is normalized so that the sum of the state transition probabilities is one.
[0124] Classify the sampling sequence according to the hidden state category to obtain state classification data, calculate the Gaussian distribution parameters corresponding to each hidden state based on the state classification data, the Gaussian distribution parameters include mean parameters and variance parameters, execute the variational expectation maximization algorithm to optimize the state transition probability matrix and the Gaussian distribution parameters, calculate the posterior probability of each hidden state corresponding to the observed value based on the current parameters, and calculate the marginal probability distribution of the state sequence through the forward-backward recursive algorithm;
[0125] The state transition probability matrix is updated according to the posterior probability, the number of transitions between each pair of hidden states is counted and normalized to obtain the updated transition probability, the Gaussian distribution parameters corresponding to each hidden state are updated based on the posterior probability and the observed value, the posterior probability calculation and parameter update operations are performed alternately until the model parameters converge, and the optimized state transition probability matrix and Gaussian distribution parameters are output.
[0126] The hidden state refers to a state that cannot be directly observed in a statistical model such as a hidden Markov model, and is usually determined by the internal process of the model. The transition frequency refers to the transition probability or frequency between states in a Markov process or a hidden Markov model, reflecting the possibility of one state transferring to another state, and is used to estimate the dynamic law of state transition in model training.
[0127] According to the operating characteristics of the power system, the system operating state is divided into five discrete hidden states: high power generation state, medium power generation state, low power generation state, extremely low power generation state and shutdown state, which constitute a hidden state set. Each hidden state corresponds to a Gaussian distribution model, which is used to describe the probability distribution characteristics of the observation value in this state. For example, the Gaussian distribution model corresponding to the high power generation state can describe the probability distribution of the system output power under the high power generation state.
[0128] Collect system operation data and construct a sampling sequence. For example, collect the system output power once every five minutes, and collect 288 data points continuously to form a sampling sequence. Based on the sampling sequence, count the transition frequencies between each hidden state. For example, count the number of transitions from the high power generation state to the medium power generation state, and the number of transitions from the high power generation state to other states. Based on these transition frequencies, construct a state transition probability matrix. This matrix is a five-by-five matrix, in which each element represents the probability of the system transitioning from one hidden state to another. For example, the first row and second column elements of the matrix represent the probability of the system transitioning from the high power generation state to the medium power generation state. In order to ensure the validity of the probability, the state transition probability matrix is normalized so that the sum of the probabilities of each row is one.
[0129] The sampling sequence is classified according to the hidden state category to obtain state classification data. For example, the sampling points belonging to the high power generation state are extracted to form a data set of the high power generation state. Based on the state classification data, the Gaussian distribution parameters corresponding to each hidden state are calculated, including the mean parameter and the variance parameter. For example, the mean and variance of the system output power in the high power generation state are calculated.
[0130] The state transition probability matrix and Gaussian distribution parameters are iteratively optimized using the variational expectation maximization algorithm. The posterior probability of each hidden state corresponding to the observed value is calculated based on the current parameters, and the marginal probability distribution of the state sequence is calculated by the forward-backward recursive algorithm, and the state transition probability matrix is updated according to the posterior probability. For example, the number of transitions between each pair of hidden states is counted, and normalization is performed to obtain the updated transition probability, and the Gaussian distribution parameters corresponding to each hidden state are updated based on the posterior probability and the observed value. For example, the mean and variance of the system output power under the high power generation state are recalculated based on the posterior probability of the high power generation state and the corresponding observed value. The posterior probability calculation and parameter update operations are performed alternately until the model parameters converge, and the optimized state transition probability matrix and Gaussian distribution parameters are output. For example, the output transition probability matrix may show that the probability of the system transferring from the high power generation state to the medium power generation state is 0.8, and the probability of transferring from the high power generation state to the shutdown state is 0.01. At the same time, the output Gaussian distribution parameters may show that the mean of the system output power under the high power generation state is 900MW and the variance is 10MW.
[0131] In this embodiment, the mixed Gaussian model is used to more finely characterize the distribution of observation values under different states, thereby improving the accuracy of state prediction. The hidden Markov model can capture the dynamic changes of the system state, and the variational expectation maximization algorithm can effectively handle situations with hidden variables, so that the model can better adapt to the complex power system operating environment. By normalizing the state transition probability matrix and using a large amount of data for model training, the robustness of the model is enhanced, enabling it to cope with data noise and abnormal situations.
[0132] S3. Construct a two-layer action evaluation control network according to the timing association control strategy and the probability transfer matrix, generate operation control instructions and obtain a control instruction sequence based on state entropy, optimize the control instruction sequence through the dominant action evaluation algorithm and confidence interval constraints, obtain smooth control instructions, and coordinately control the intelligent body through a distributed reinforcement learning framework to obtain a coordinated control strategy, construct an online learning module based on an adaptive dynamic programming algorithm and approximate the system value function according to a radial basis function network, determine the update step size in combination with Lyapunov stability analysis, obtain optimized controller parameters, extract control scenario features through a meta-learning-based strategy migration mechanism and a task embedding network, optimize the coordinated control strategy according to the rapid gradient descent, and obtain the optimal control strategy.
[0133] The two-layer action evaluation control network is a control strategy network for reinforcement learning, which combines a two-layer structure to evaluate and optimize the action selection process. The state entropy is a measure of the uncertainty or randomness of the state. In information theory, the higher the entropy, the more uncertain the state of the system. The dominant action evaluation algorithm is a method for evaluating the quality of actions in reinforcement learning. It helps the intelligent agent select the optimal action by calculating the advantage of the current action over other actions. The adaptive dynamic programming algorithm is a dynamic optimization method that can adaptively adjust the parameters in the learning and optimization process so that the system can find the optimal solution in a dynamic environment. The radial basis function network is a feedforward neural network that uses radial basis functions as activation functions and is mainly used for pattern recognition and function approximation problems. The Lyapunov stability analysis is a mathematical method for analyzing the stability of a dynamic system. It is based on the Lyapunov function to determine whether the system will return to a state of equilibrium under disturbance.
[0134] In an optional embodiment,
[0135] A two-layer action evaluation control network is constructed according to the timing association control strategy and the probability transfer matrix, and an operation control instruction is generated and a control instruction sequence is obtained based on the state entropy. The control instruction sequence is optimized by the dominant action evaluation algorithm and the confidence interval constraint to obtain a stable control instruction and to coordinately control the intelligent agent through a distributed reinforcement learning framework to obtain a coordinated control strategy. An online learning module is constructed based on an adaptive dynamic programming algorithm and the system value function is approximated according to a radial basis function network. The update step size is determined in combination with Lyapunov stability analysis to obtain optimized controller parameters. The control scene features are extracted through a meta-learning-based strategy migration mechanism and a task embedding network. The coordinated control strategy is optimized according to the rapid descent of the gradient to obtain the optimal control strategy including:
[0136] A double-layer action evaluation control network is constructed according to a timing association control strategy and a probability transfer matrix, and the temperature information, pressure information, and flow information are normalized to obtain a normalized state quantity, and the valve opening information and the motor speed information are normalized to obtain a normalized action quantity. The bottom network of the double-layer action evaluation control network receives the normalized state quantity and the normalized action quantity, extracts features through multiple hidden layers and rectified linear activation functions, and the output layer uses a linear activation function to generate an operation control instruction. The top network of the double-layer action evaluation control network uses a long short-term memory network structure, receives the output sequence of the bottom network at multiple consecutive moments through a memory unit, extracts the timing correlation features of the action sequence, divides the system state space into a temperature sub-interval and a pressure sub-interval, and counts the state distribution frequency of the temperature sub-interval and the pressure sub-interval within a preset time period, calculates the state entropy value based on the state distribution frequency, and screens the operation control instructions according to the state entropy value to obtain a control instruction sequence;
[0137] The cumulative reward value of the current action is calculated by the dominant action evaluation algorithm in the dynamic evaluation window, the dominant value is obtained by performing a difference operation between the cumulative reward value and the average reward value of all actions in the dynamic evaluation window, and the upper and lower limits of the dominant value are set by a confidence interval constraint mechanism to obtain a stable control instruction;
[0138] Construct a distributed reinforcement learning framework, divide the system into a reactor subsystem, a heat exchanger subsystem, and a separator subsystem, configure an intelligent agent for each subsystem to perform state observation and control, and the intelligent agent interactively shares observation data and control experience through a communication network, and collaboratively executes the stable control instructions to obtain a coordinated control strategy;
[0139] An online learning module is constructed based on an adaptive dynamic programming algorithm. A radial basis function network is used to approximate the system value function. An experience replay pool is constructed to store the interactive data of the state vector, action vector, reward value, next state vector, and termination flag. The state samples are clustered by a clustering method to obtain the center point of the basis function. A value function approximation network is constructed based on the center point of the basis function. The system state change rate is obtained in real time and a Lyapunov stability analysis is performed. If the state change rate is greater than a preset change rate threshold, the network parameter update step size is increased, otherwise the network parameter update step size is reduced to obtain the optimized controller parameters.
[0140] A meta-learning-based strategy migration mechanism is constructed. The control scenario features are extracted through a task embedding network. The task embedding network extracts features of the input state through multiple convolutional layers, compresses features through pooling layers, and encodes the compressed features into task embedding vectors through fully connected layers. The coordinated control strategy is optimized using the fast gradient descent method. During the optimization process, the strategy network parameters are adjusted in combination with the task embedding vector, and the momentum term is superimposed to accelerate convergence. The optimal parameter configuration is saved by periodically evaluating the strategy performance to obtain the optimal control strategy.
[0141] The basis function center point is an important parameter in the radial basis function network, which represents the center position of each radial basis function of the network. By adjusting the center point, the network can better fit the input data and improve the prediction accuracy. The momentum term is a concept in the gradient descent method, which means that in the optimization process, the gradient update direction at the previous moment will affect the update at the current moment.
[0142] A two-layer action evaluation control network is constructed to generate control instructions, such as adjusting valve opening and motor speed. The bottom layer of the network receives normalized state information such as temperature, pressure, and flow, as well as normalized action information such as valve opening and motor speed, and outputs operation control instructions after extracting features through a multi-layer network. The top network adopts a long short-term memory network structure, receives the output sequence of the bottom network at multiple consecutive moments, considers the temporal correlation of the action sequence, and screens the operation control instructions according to the entropy value of the system state to form a control instruction sequence. For example, the temperature is divided into sub-intervals such as 0-50℃ and 50-100℃, and the pressure is divided into sub-intervals such as 0-1MPa and 1-2MPa. The frequency of occurrence of each sub-interval within a preset time period (for example, 1 minute) is counted, and the state entropy value is calculated based on the frequency. If a control instruction causes the system state entropy value to decrease, the instruction is retained.
[0143] Optimize the control instruction sequence. Use the dominant action evaluation algorithm to calculate the cumulative reward value of each action within a time window (for example, the past 10 minutes). Compare the cumulative reward value of each action with the average reward value of all actions in the time window to get the dominance value. Set upper and lower limits for the dominance value and limit the out-of-range dominance values to get smooth control instructions. For example, if the cumulative reward value of an action is much higher than the average, limit its dominance value to avoid overly aggressive control.
[0144] A distributed reinforcement learning framework is constructed to divide the system into reactor subsystem, heat exchanger subsystem and separator subsystem. Each subsystem is equipped with an agent responsible for observing the subsystem status and performing control. Agents share data and experience through network communication, and cooperate to execute smooth control instructions to form a coordinated control strategy. For example, the reactor agent shares temperature information with the heat exchanger agent, and the heat exchanger agent adjusts the heat exchange efficiency according to the temperature information.
[0145] Construct an online learning module and use a radial basis function network to approximate the system value function. Store interaction data such as state vector, action vector, reward value, next state vector, and termination flag in the experience replay pool. Use a clustering method to cluster state samples, obtain the center point of the basis function, and build a value function approximation network based on the center point. Monitor the system state change rate in real time and perform Lyapunov stability analysis. If the state change rate exceeds the preset threshold, increase the network parameter update step size; otherwise, reduce the update step size. For example, if the temperature change rate exceeds 10℃ / minute, increase the update step size to speed up learning.
[0146] Construct a meta-learning-based policy transfer mechanism. Use a task embedding network to extract control scenario features. The network extracts state features through multiple convolutional layers, compresses features through pooling layers, and finally encodes features into task embedding vectors through fully connected layers. Use the gradient fast descent method to optimize the coordinated control strategy. During the optimization process, adjust the policy network parameters in combination with the task embedding vector, and superimpose the momentum term to accelerate convergence. Periodically evaluate the policy performance and save the optimal parameter configuration to obtain the optimal control policy. For example, if the current control scenario is similar to the historical scenario, use the experience of the historical scenario to accelerate learning.
[0147] In this embodiment, through the two-layer action evaluation network and state entropy control, more accurate control instructions can be generated, and through the dominant action evaluation algorithm and confidence interval constraints, the control can be made smoother, avoiding overreaction and oscillation of the control, thereby improving the control accuracy. Through the distributed reinforcement learning framework and the online learning module, the system can adapt to different working conditions and disturbances, and through the Lyapunov stability analysis, the stability of the system is guaranteed and the risk of system loss of control is avoided. Through the meta-learning strategy migration mechanism, the existing control experience can be used to quickly adapt to new control scenarios, reducing learning time and cost, thereby improving learning efficiency.
[0148] In an optional embodiment,
[0149] The state samples are clustered by clustering method to obtain the center point of basis function, and the value function approximation network is constructed based on the center point of basis function, and the system state change rate is obtained in real time and Lyapunov stability analysis is performed, including:
[0150] A hybrid clustering strategy combining density peak clustering and adaptive K-means is used to process state samples. By calculating the local density and distance factor of the points corresponding to the state samples, state sample points whose number of adjacent sample points exceeds 20% of the total samples and whose mutual distance is greater than a preset distance threshold are selected as candidate center points, and the candidate center points are used as the initial clustering centers.
[0151] A dynamic category number adjustment mechanism is introduced into the hybrid clustering strategy. Based on the intra-category sample variance threshold, for each category, if the intra-category sample variance exceeds the intra-category sample variance threshold, the current category is split. If the center distance of the adjacent category corresponding to the current category is less than the preset distance, the adjacent category is merged with the current category to obtain the final cluster center point.
[0152] A multi-scale radial basis function network is constructed based on the final cluster center points, and a kernel width parameter, a small kernel width parameter, and a large kernel width parameter that are positively correlated with the distribution range of the current category samples are configured for each of the final cluster center points to obtain an initial value function network;
[0153] The state samples are predicted by the initial value function network, and the error sequence between the predicted value and the actual value is calculated and added to the sliding time window to obtain the system state change rate sequence, and the statistical characteristics of the system state change rate sequence are calculated according to the sliding time window, wherein the statistical characteristics include the change rate mean and the change rate variance;
[0154] A Lyapunov function is constructed based on the statistical characteristics, the change rate mean and the change rate variance are used as input variables of the Lyapunov function, and the ratio of the difference between the Lyapunov function values at two adjacent sampling moments and the sampling period is calculated to obtain the Lyapunov function change rate.
[0155] The density peak clustering is a density-based clustering method that groups data points by identifying density peaks around the data points. The adaptive K-means is an improved version of the classic K-means clustering algorithm that can automatically adjust the number of clusters or the shape of clusters according to changes in data distribution. The distance factor is a parameter in the clustering algorithm that is used to measure the similarity between data points and is often used as a metric to determine which data points should be classified into the same cluster during the clustering process. The clustering results will also be different based on different choices of the distance factor. The kernel width parameter is an important parameter in the kernel function (such as the Gaussian kernel), which controls the width of the kernel function, that is, it affects the calculation of the similarity of data points to other data points.
[0156] Collect state sample data during the system operation. For example, collect the position, velocity, acceleration and other data of a robot during its motion, and store these data as a sample data set. Assume that the collected sample data contains 1000 sample points, and each sample point contains three-dimensional data, namely position, velocity and acceleration.
[0157] A hybrid clustering strategy combining density peak clustering and adaptive K-means is used to cluster state samples. The distance between each sample point and its surrounding sample points is calculated, and the local density of each sample point is determined based on the distance. At the same time, the minimum distance from each sample point to a sample point with a greater density than it is calculated as the distance factor of the sample point. Sample points with a large local density and a large distance factor are selected as candidate center points. For example, sample points whose number of adjacent sample points exceeds 20% of the total number of samples and whose mutual distance is greater than the preset distance threshold of 0.5 are set as candidate center points.
[0158] A dynamic category number adjustment mechanism is introduced into the hybrid clustering strategy. According to the variance of samples within each category, it is determined whether the category needs to be split or merged. For example, the intra-category sample variance threshold is set to 0.2. If the sample variance of a category is greater than 0.2, the category is split into two subcategories. If the center distance between a category and its adjacent category is less than the preset distance 0.3, the two categories are merged into one category. By continuously adjusting the number of categories, a set of cluster center points is finally obtained.
[0159] A multi-scale radial basis function network is constructed based on the final cluster center points. A kernel width parameter that is positively correlated with the distribution range of the current category samples is configured for each cluster center point. For example, different kernel width parameters are set according to the average distance from the sample points to the center point in each category. The farther the sample points are from the center point, the larger the kernel width parameter they correspond to. At the same time, small kernel width parameters and large kernel width parameters are configured for each category to adapt to sample distributions of different scales.
[0160] The constructed initial value function network is used to predict the state samples, and the error sequence between the predicted value and the actual value is added to a sliding time window. For example, the length of the sliding time window is set to 10. According to the error sequence in the sliding time window, the statistical characteristics of the system state change rate sequence are calculated, including the change rate mean and change rate variance.
[0161] The Lyapunov function is constructed based on the calculated statistical characteristics. The mean and variance of the rate of change are used as input variables of the Lyapunov function. The ratio of the difference between the Lyapunov function values at two adjacent sampling moments and the sampling period is calculated to obtain the rate of change of the Lyapunov function. According to the Lyapunov stability theory, if the rate of change of the Lyapunov function is always less than zero, the system is stable.
[0162] In this embodiment, by combining the clustering method and Lyapunov stability analysis, the stability of the system can be evaluated more accurately, avoiding the limitations of traditional methods. By constructing a value function approximation network, the system state change rate can be obtained in real time, which improves the efficiency of stability evaluation. It is suitable for various complex dynamic systems and has strong universality.
[0163] Figure 2 FIG. 1 is a schematic diagram of a structure of a permanent magnet motor energy storage and optimized operation system based on electricity price according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0164] The first unit is used to collect the operation data of the permanent magnet motor and decompose the operation data of the permanent magnet motor into multiple intrinsic mode functions through variational mode decomposition, determine the optimal number of modes based on signal mutual information and iteratively update the center frequency of each mode in combination with particle swarm optimization, obtain the operation feature component and add it to the deep belief neural network, pre-train the restricted Boltzmann machine layer by layer through the contrast divergence algorithm and add the self-attention module to extract the time series feature weight, output the power prediction result, construct a spatiotemporal graph convolutional neural network for electricity price prediction, construct a dynamic adjacency matrix in combination with the spatiotemporal correlation of the historical electricity price data obtained in advance, extract multi-scale features through parallel convolutional layers and determine long-term and short-term dependencies to obtain the electricity price prediction sequence, quantify the system error through the Gaussian mixture model, estimate the model parameters in combination with variational Bayesian inference, and adaptively adjust the number of Gaussian distributions through the marginal likelihood function criterion to obtain a multidimensional probability distribution model;
[0165] The second unit is used to obtain the power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model, perform global optimization through a quantum genetic optimization algorithm, represent optimization variables through quantum bit encoding, update the population in combination with a quantum revolving gate evolution operator and introduce a quantum entanglement mechanism to obtain an initial optimization strategy set, add the initial optimization strategy set to a meta-heuristic hybrid optimization model, perform strategy search according to adaptive temperature regulation and a dynamic taboo table, generate experimental solutions in combination with a differential strategy, construct a timing optimization network based on an attention mechanism, extract control sequence features and establish timing constraint relationships based on a causal attention mechanism, set a discrete action layer based on Gobel-Soft maximum sampling, obtain a timing-related control strategy, perform system modeling through a hybrid Gaussian-Hidden Markov model, estimate model parameters through a variational expectation maximization algorithm, and obtain a probability transfer matrix based on a particle filter prediction state transition process;
[0166] The third unit is used to construct a two-layer action evaluation control network according to the timing association control strategy and the probability transfer matrix, generate operation control instructions and obtain a control instruction sequence based on state entropy, optimize the control instruction sequence through the dominant action evaluation algorithm and the confidence interval constraint, obtain smooth control instructions, and coordinate the intelligent body through a distributed reinforcement learning framework to obtain a coordinated control strategy, build an online learning module based on an adaptive dynamic programming algorithm and approximate the system value function according to a radial basis function network, determine the update step size in combination with Lyapunov stability analysis, obtain optimized controller parameters, extract control scenario features through a meta-learning-based strategy migration mechanism and a task embedding network, optimize the coordinated control strategy according to the rapid descent of the gradient, and obtain the optimal control strategy.
[0167] According to a third aspect of the embodiments of the present invention,
[0168] An electronic device is provided, comprising:
[0169] processor;
[0170] a memory for storing processor-executable instructions;
[0171] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0172] A fourth aspect of the embodiments of the present invention is:
[0173] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0174] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for energy storage and optimized operation of a permanent magnet motor based on electricity price, characterized in that: include: The operation data of the permanent magnet motor is collected and decomposed into multiple intrinsic mode functions through variational mode decomposition. The optimal number of modes is determined based on the signal mutual information and the center frequency of each mode is iteratively updated in combination with particle swarm optimization to obtain the operation feature component and add it to the deep belief neural network. The restricted Boltzmann machine is pre-trained layer by layer through the contrast divergence algorithm and a self-attention module is added to extract the time series feature weights, the power prediction result is output, and a spatiotemporal graph convolutional neural network is constructed to predict the electricity price. A dynamic adjacency matrix is constructed in combination with the spatiotemporal correlation of the historical electricity price data obtained in advance. Multi-scale features are extracted through parallel convolutional layers and long-term and short-term dependencies are determined to obtain the electricity price prediction sequence. The system error is quantified through a Gaussian mixture model, and the model parameters are estimated in combination with variational Bayesian inference. The number of Gaussian distributions is adaptively adjusted through the marginal likelihood function criterion to obtain a multidimensional probability distribution model. The power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model are obtained, global optimization is performed through a quantum genetic optimization algorithm, optimization variables are represented by quantum bit encoding, the population is updated in combination with a quantum revolving gate evolution operator, and a quantum entanglement mechanism is introduced to obtain an initial optimization strategy set, the initial optimization strategy set is added to a meta-heuristic hybrid optimization model, strategy search is performed according to adaptive temperature regulation and a dynamic taboo table, experimental solutions are generated in combination with a differential strategy, a timing optimization network based on an attention mechanism is constructed, control sequence features are extracted and timing constraint relationships are established based on a causal attention mechanism, a discrete action layer based on Gobel-Soft maximum sampling is set, a timing association control strategy is obtained, system modeling is performed through a hybrid Gaussian-Hidden Markov model, model parameters are estimated through a variational expectation maximization algorithm, and a probability transfer matrix is obtained according to a particle filter prediction state transition process; A two-layer action evaluation control network is constructed according to the timing association control strategy and the probability transfer matrix, an operation control instruction is generated, and a control instruction sequence is obtained based on the state entropy. The control instruction sequence is optimized by the dominant action evaluation algorithm and the confidence interval constraint, a stable control instruction is obtained, and the intelligent agent is collaboratively controlled by a distributed reinforcement learning framework to obtain a coordinated control strategy. An online learning module is constructed based on an adaptive dynamic programming algorithm, and the system value function is approximated according to a radial basis function network. The update step size is determined in combination with Lyapunov stability analysis to obtain optimized controller parameters. The control scene features are extracted by a meta-learning-based strategy migration mechanism and a task embedding network. The coordinated control strategy is optimized according to the rapid descent of the gradient to obtain the optimal control strategy, including: A double-layer action evaluation control network is constructed according to a timing association control strategy and a probability transfer matrix, and the temperature information, pressure information, and flow information are normalized to obtain a normalized state quantity, and the valve opening information and the motor speed information are normalized to obtain a normalized action quantity. The bottom network of the double-layer action evaluation control network receives the normalized state quantity and the normalized action quantity, extracts features through multiple hidden layers and rectified linear activation functions, and the output layer uses a linear activation function to generate an operation control instruction. The top network of the double-layer action evaluation control network uses a long short-term memory network structure, receives the output sequence of the bottom network at multiple consecutive moments through a memory unit, extracts the timing correlation features of the action sequence, divides the system state space into a temperature sub-interval and a pressure sub-interval, and counts the state distribution frequency of the temperature sub-interval and the pressure sub-interval within a preset time period, calculates the state entropy value based on the state distribution frequency, and screens the operation control instructions according to the state entropy value to obtain a control instruction sequence; The cumulative reward value of the current action is calculated by the dominant action evaluation algorithm in the dynamic evaluation window, the dominant value is obtained by performing a difference operation between the cumulative reward value and the average reward value of all actions in the dynamic evaluation window, and the upper and lower limits of the dominant value are set by a confidence interval constraint mechanism to obtain a stable control instruction; Construct a distributed reinforcement learning framework, divide the system into a reactor subsystem, a heat exchanger subsystem, and a separator subsystem, configure an intelligent agent for each subsystem to perform state observation and control, and the intelligent agent interactively shares observation data and control experience through a communication network, and collaboratively executes the stable control instructions to obtain a coordinated control strategy; An online learning module is constructed based on an adaptive dynamic programming algorithm. A radial basis function network is used to approximate the system value function. An experience replay pool is constructed to store the interactive data of the state vector, action vector, reward value, next state vector, and termination flag. The state samples are clustered by a clustering method to obtain the center point of the basis function. A value function approximation network is constructed based on the center point of the basis function. The system state change rate is obtained in real time and a Lyapunov stability analysis is performed. If the state change rate is greater than a preset change rate threshold, the network parameter update step size is increased, otherwise the network parameter update step size is reduced to obtain the optimized controller parameters. A meta-learning-based strategy migration mechanism is constructed. The control scenario features are extracted through a task embedding network. The task embedding network extracts features of the input state through multiple convolutional layers, compresses features through pooling layers, and encodes the compressed features into task embedding vectors through fully connected layers. The coordinated control strategy is optimized using the fast gradient descent method. During the optimization process, the strategy network parameters are adjusted in combination with the task embedding vector, and the momentum term is superimposed to accelerate convergence. The optimal parameter configuration is saved by periodically evaluating the strategy performance to obtain the optimal control strategy.
2. The method according to claim 1, characterized in that: The operation data of the permanent magnet motor is collected and decomposed into multiple intrinsic mode functions through variational mode decomposition. The optimal number of modes is determined based on the signal mutual information and the center frequency of each mode is iteratively updated in combination with particle swarm optimization to obtain the operation feature component and add it to the deep belief neural network. The restricted Boltzmann machine is pre-trained layer by layer through the contrast divergence algorithm and a self-attention module is added to extract the time series feature weights, the power prediction result is output, and a spatiotemporal graph convolutional neural network is constructed to predict the electricity price. A dynamic adjacency matrix is constructed by combining the spatiotemporal correlation of the historical electricity price data obtained in advance. Multi-scale features are extracted through parallel convolutional layers and long-term and short-term dependencies are determined to obtain the electricity price prediction sequence. The system error is quantified through a Gaussian mixture model, and the model parameters are estimated in combination with variational Bayesian inference. The number of Gaussian distributions is adaptively adjusted through the marginal likelihood function criterion to obtain a multidimensional probability distribution model including: Collecting permanent magnet motor operation data, the permanent magnet motor operation data includes a stator current signal, a stator voltage signal and a rotor speed signal, inputting the permanent magnet motor operation data into a variational mode decomposition module, and decomposing the permanent magnet motor operation data into a plurality of intrinsic mode functions through the variational mode decomposition module, wherein the variational mode decomposition module converts signal decomposition into an optimization problem by constructing an augmented Lagrangian function, iteratively updates each mode function by an alternating direction multiplier method, calculates the signal mutual information value between adjacent intrinsic mode functions, constructs an evaluation index based on the signal mutual information value, and if the difference of the signal mutual information value is less than a preset threshold, determines the corresponding mode number as the optimal mode number; The center frequency of each intrinsic mode function is iteratively updated through the particle swarm optimization algorithm. The position and speed of the particle swarm are initialized, and the fitness function is constructed based on the mean square error between the reconstructed signal and the original signal. The position of each particle is expressed as a combination of the center frequencies. The individual optimal solution and the global optimal solution are calculated according to the fitness function. The particle position and speed are dynamically updated according to the inertia weight, the individual learning factor and the social learning factor to obtain the corresponding operating characteristic components of the permanent magnet motor. Input the operation feature component into a deep belief neural network, wherein the deep belief neural network is composed of a plurality of stacked restricted Boltzmann machines, the restricted Boltzmann machines are pre-trained by a contrastive divergence algorithm, a network reconstruction sample is obtained by Gibbs sampling, forward propagation and back propagation are alternately performed based on the conditional probability distribution between the visible layer and the hidden layer, the gradients of the weight parameters and the bias parameters are calculated and the network parameters are updated, a self-attention module is constructed in the deep belief neural network, the self-attention module maps the operation feature component into a query matrix, a key matrix and a value matrix, an attention score is obtained by performing a dot product operation of the query matrix and the key matrix and performing softmax normalization, the attention score is multiplied by the value matrix to obtain a weighted feature representation, and a power prediction result is obtained; Constructing a spatiotemporal graph convolutional neural network, calculating the correlation coefficient between historical electricity price data at different time points, constructing a dynamic adjacency matrix based on the correlation coefficient, setting a multi-scale parallel convolution layer in the spatiotemporal graph convolutional neural network, the multi-scale parallel convolution layer uses convolution kernels with different receptive fields to extract features, and fuses multi-scale information through feature splicing, setting a graph convolution layer to aggregate spatial features, the graph convolution layer calculates the Laplace matrix based on the dynamic adjacency matrix, and implements graph convolution operations in combination with Chebyshev polynomial approximation, constructing a long short-term memory unit, the long short-term memory unit controls the importance of current input information through an input gate, uses a forget gate to forget historical information, and uses an output gate to adjust the unit state output to achieve long-term dependency modeling, and inputs the historical electricity price data and the dynamic adjacency matrix into the spatiotemporal graph convolutional neural network to generate an electricity price prediction sequence; The prediction error of the power prediction result and the prediction error of the electricity price prediction sequence are input into a Gaussian mixture model, and the distribution parameters of the Gaussian mixture model are estimated according to the variational Bayesian inference algorithm. The variational distribution is introduced to approximate the posterior distribution, and a variational lower bound is constructed as the optimization target. The coordinate ascent method is used to iteratively optimize the mean, covariance and mixing weight of the Gaussian components. The model complexity is calculated based on the marginal likelihood function. If the improvement of the marginal likelihood function after adding a new Gaussian component is less than a preset improvement threshold, the Gaussian component is stopped from being added, and a multidimensional probability distribution model of the system prediction error is established.
3. The method according to claim 2, characterized in that The distribution parameters of the Gaussian mixture model are estimated according to the variational Bayesian inference algorithm. The variational distribution is introduced to approximate the posterior distribution, and the variational lower bound is constructed as the optimization target. The mean, covariance and mixture weight of the Gaussian components are iteratively optimized using the coordinate ascent method. The model complexity is calculated based on the marginal likelihood function, including: Constructing a Gaussian mixture model, taking power prediction error data and electricity price prediction error data as observation data input, normalizing the observation data to obtain standardized observation data, initializing the number of Gaussian components of the Gaussian mixture model, and setting an initial mean parameter, an initial covariance matrix parameter, and an initial mixing weight parameter for each Gaussian component, wherein the initial mean parameter is obtained by adding the mean of the standardized observation data to a random disturbance value, the initial covariance matrix parameter is obtained by multiplying a unit matrix by a preset positive number, and the initial mixing weight parameter is set using a uniform distribution; Constructing a variational distribution model, the variational distribution model includes a latent variable Dirichlet distribution module, a mean Gaussian distribution module and a covariance Wishart distribution module, and constructing an optimization objective function based on the divergence between the variational distribution model and the true posterior distribution, the optimization objective function includes an observation data likelihood expectation term and a distribution relative entropy term; The parameters of the Gaussian mixture model are iteratively optimized by using the coordinate ascent method, and the parameters are randomly selected in each iterative optimization to be fixed and the distribution parameters of the latent variable Dirichlet distribution module are updated, the posterior probability of the Gaussian component corresponding to each of the standardized observation data is calculated, and the distribution parameters of the mean Gaussian distribution module and the distribution parameters of the covariance Wishart distribution module are updated based on the posterior probability, wherein the distribution parameters of the mean Gaussian distribution module are updated by calculating sufficient statistics and a posterior precision matrix, and the distribution parameters of the covariance Wishart distribution module are updated by accumulating second-order statistics; During the parameter updating process, the marginal likelihood function value is calculated, wherein the marginal likelihood function value is used to characterize the fitting degree and complexity of the Gaussian mixture model to obtain the model complexity.
4. The method according to claim 1, characterized in that: The power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model are obtained, and global optimization is performed through a quantum genetic optimization algorithm. The optimization variables are represented by quantum bit encoding, and the population is updated by combining the quantum revolving gate evolution operator and the quantum entanglement mechanism is introduced to obtain an initial optimization strategy set. The initial optimization strategy set is added to the meta-heuristic hybrid optimization model, and strategy search is performed according to adaptive temperature regulation and dynamic taboo table. The experimental solution is generated by combining the differential strategy, and a timing optimization network based on the attention mechanism is constructed. The control sequence features are extracted and the timing constraint relationship is established based on the causal attention mechanism. A discrete action layer based on the Gobel-Soft maximum sampling is set to obtain a timing association control strategy. The system is modeled by a mixed Gaussian-hidden Markov model, and the model parameters are estimated by a variational expectation maximization algorithm. The probability transfer matrix obtained according to the particle filter prediction state transfer process includes: Obtain a power prediction result data set, an electricity price prediction sequence, and a multidimensional probability distribution model, generate quantum bit encoding to represent optimization variables based on a quantum genetic optimization algorithm and initialize the population, describe the quantum state of each quantum bit by setting the amplitude value range to 0 to 1 and the phase angle range to 0 to 2π, and combine multiple quantum bits to construct a chromosome encoding structure; Performing a quantum revolving door evolution operation to update the population, calculating the fitness value of each individual in the population and determining the rotation angle of the quantum revolving door based on the fitness value, sorting the individuals in the population according to the fitness value and selecting individuals with fitness values in the top 20% as dominant individuals, selecting local gene fragments of the dominant individuals to perform entangled state superposition operations with other individuals to form a new quantum state, performing a quantum measurement operation on the quantum state to collapse it into a classical solution to construct an initial optimization strategy set; Input the initial optimization strategy set into the meta-heuristic hybrid optimization model, establish a temperature adaptive adjustment mechanism to dynamically adjust the temperature parameters, construct a dynamic taboo table to record and update the characteristic values of the visited solution space area, and at the same time screen feasible solutions according to the dynamic taboo table, perform a differential strategy generation operation based on population evolution, select the basis vector in the current population and extract the historical optimal solution, perform a differential operation on the basis vector and the historical optimal solution to generate a differential vector, apply the differential vector to the current solution to generate a test solution and perform boundary constraint correction; Construct a multi-layer attention structured timing optimization network and perform sequence feature extraction. Perform embedded coding on the input original control sequence. Extract sequence features of different time scales through multiple attention heads set in parallel. Calculate the similarity scores between the query vector, key vector, and value vector and generate attention weights. Perform mask operations of the causal attention mechanism to filter information after the current moment. Construct an attention weight matrix to represent the temporal dependencies at different moments. Perform concatenation operations on the features output by each attention head and input them to the next network layer through linear transformation. Based on the Gubel-Soft maximum sampling, a discretization mapping is performed, the score value of each possible action is calculated, and probability sampling is performed according to the score value to generate the final discrete control action sequence, a mixed Gaussian-hidden Markov model is constructed and multiple discrete hidden states are set, and a corresponding Gaussian distribution model is configured for each hidden state to describe the distribution of observation values, and a Markov transfer relationship between the hidden states is established. The variational expectation maximization algorithm is executed to iteratively optimize the parameters of the mixed Gaussian-hidden Markov model, and the posterior distribution calculation and the likelihood function maximization operation are alternately performed, and the state transition probability matrix and the Gaussian distribution parameters corresponding to each hidden state are output; Execute the particle filter to predict the state transfer process, initialize the particle swarm and predict the motion trajectory of each particle based on the state transfer probability, calculate the weight value of each particle according to the observed data, perform the importance resampling operation to screen high-weight particles and update the particle swarm distribution, and summarize the particle swarm distribution to obtain the probability matrix of system state transfer.
5. The method according to claim 4, characterized in that Establish a Markov transfer relationship between the hidden states, execute the variational expectation maximization algorithm to iteratively optimize the parameters of the mixed Gaussian-hidden Markov model, alternately perform posterior distribution calculation and likelihood function maximization operations, and output the state transfer probability matrix and the Gaussian distribution parameters corresponding to each hidden state, including: Divide multiple discrete hidden states according to system operation characteristics and establish a hidden state set, set the discrete hidden states to high power generation state, medium power generation state, low power generation state, extremely low power generation state and shutdown state, and configure a Gaussian distribution model for each discrete hidden state to describe the probability distribution characteristics of the observation value; The data of the acquisition system operation is used to construct a sampling sequence, wherein the sampling interval of the sampling sequence is five minutes, and the sampling sequence includes 288 sampling points. Based on the sampling sequence, the transition frequencies between the hidden states are counted, and a state transition probability matrix is constructed according to the transition frequencies. The dimension of the state transition probability matrix is five times five, and the probability of the system transferring from the current hidden state to the next hidden state is represented by the state transition probability matrix. The state transition probability matrix is normalized so that the sum of the state transition probabilities is one. Classify the sampling sequence according to the hidden state category to obtain state classification data, calculate the Gaussian distribution parameters corresponding to each hidden state based on the state classification data, the Gaussian distribution parameters include mean parameters and variance parameters, execute the variational expectation maximization algorithm to optimize the state transition probability matrix and the Gaussian distribution parameters, calculate the posterior probability of each hidden state corresponding to the observed value based on the current parameters, and calculate the marginal probability distribution of the state sequence through the forward-backward recursive algorithm; The state transition probability matrix is updated according to the posterior probability, the number of transitions between each pair of hidden states is counted and normalized to obtain the updated transition probability, the Gaussian distribution parameters corresponding to each hidden state are updated based on the posterior probability and the observed value, the posterior probability calculation and parameter update operations are performed alternately until the model parameters converge, and the optimized state transition probability matrix and Gaussian distribution parameters are output.
6. The method according to claim 1, characterized in that The state samples are clustered by clustering method to obtain the center point of basis function, and the value function approximation network is constructed based on the center point of basis function, and the system state change rate is obtained in real time and Lyapunov stability analysis is performed, including: A hybrid clustering strategy combining density peak clustering and adaptive K-means is used to process state samples. By calculating the local density and distance factor of the points corresponding to the state samples, state sample points whose number of adjacent sample points exceeds 20% of the total samples and whose mutual distance is greater than a preset distance threshold are selected as candidate center points, and the candidate center points are used as the initial clustering centers. A dynamic category number adjustment mechanism is introduced into the hybrid clustering strategy. Based on the intra-category sample variance threshold, for each category, if the intra-category sample variance exceeds the intra-category sample variance threshold, the current category is split. If the center distance of the adjacent category corresponding to the current category is less than the preset distance, the adjacent category is merged with the current category to obtain the final cluster center point. A multi-scale radial basis function network is constructed based on the final cluster center points, and a kernel width parameter, a small kernel width parameter, and a large kernel width parameter that are positively correlated with the distribution range of the current category samples are configured for each of the final cluster center points to obtain an initial value function network; The state samples are predicted by the initial value function network, and the error sequence between the predicted value and the actual value is calculated and added to the sliding time window to obtain the system state change rate sequence, and the statistical characteristics of the system state change rate sequence are calculated according to the sliding time window, wherein the statistical characteristics include the change rate mean and the change rate variance; A Lyapunov function is constructed based on the statistical characteristics, the change rate mean and the change rate variance are used as input variables of the Lyapunov function, and the ratio of the difference between the Lyapunov function values at two adjacent sampling moments and the sampling period is calculated to obtain the Lyapunov function change rate.
7. A permanent magnet motor energy storage and optimized operation system based on electricity price, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to collect the operation data of the permanent magnet motor and decompose the operation data of the permanent magnet motor into multiple intrinsic mode functions through variational mode decomposition, determine the optimal number of modes based on signal mutual information and iteratively update the center frequency of each mode in combination with particle swarm optimization, obtain the operation feature component and add it to the deep belief neural network, pre-train the restricted Boltzmann machine layer by layer through the contrast divergence algorithm and add the self-attention module to extract the time series feature weight, output the power prediction result, construct a spatiotemporal graph convolutional neural network for electricity price prediction, construct a dynamic adjacency matrix in combination with the spatiotemporal correlation of the historical electricity price data obtained in advance, extract multi-scale features through parallel convolutional layers and determine long-term and short-term dependencies to obtain the electricity price prediction sequence, quantify the system error through the Gaussian mixture model, estimate the model parameters in combination with variational Bayesian inference, and adaptively adjust the number of Gaussian distributions through the marginal likelihood function criterion to obtain a multidimensional probability distribution model; The second unit is used to obtain the power prediction result, the electricity price prediction sequence and the multidimensional probability distribution model, perform global optimization through a quantum genetic optimization algorithm, represent optimization variables through quantum bit encoding, update the population in combination with a quantum revolving gate evolution operator and introduce a quantum entanglement mechanism to obtain an initial optimization strategy set, add the initial optimization strategy set to a meta-heuristic hybrid optimization model, perform strategy search according to adaptive temperature regulation and a dynamic taboo table, generate experimental solutions in combination with a differential strategy, construct a timing optimization network based on an attention mechanism, extract control sequence features and establish timing constraint relationships based on a causal attention mechanism, set a discrete action layer based on Gobel-Soft maximum sampling, obtain a timing-related control strategy, perform system modeling through a hybrid Gaussian-Hidden Markov model, estimate model parameters through a variational expectation maximization algorithm, and obtain a probability transfer matrix based on a particle filter prediction state transition process; The third unit is used to construct a two-layer action evaluation control network according to the timing association control strategy and the probability transfer matrix, generate operation control instructions and obtain a control instruction sequence based on state entropy, optimize the control instruction sequence through the dominant action evaluation algorithm and the confidence interval constraint, obtain smooth control instructions, and coordinate the intelligent body through a distributed reinforcement learning framework to obtain a coordinated control strategy, build an online learning module based on an adaptive dynamic programming algorithm and approximate the system value function according to a radial basis function network, determine the update step size in combination with Lyapunov stability analysis, obtain optimized controller parameters, extract control scenario features through a meta-learning-based strategy migration mechanism and a task embedding network, optimize the coordinated control strategy according to the rapid descent of the gradient, and obtain the optimal control strategy.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Robust optimization scheduling method and device for comprehensive energy system
CN111401664A
Park system access power distribution network optimization method based on load prediction uncertainty
CN117559566A
Cited By
Direct current brushless motor working condition self-adaptive control method and system based on multi-source data fusion
CN122419278A