An abnormal flow recognition system at the gas meter end based on reinforcement learning

Through the gas meter terminal abnormal flow recognition system based on reinforcement learning, the identification model is automatically updated, and the problem of inefficient identification in the existing technology is solved, and efficient abnormal flow detection without manual labeling is realized, ensuring the safe and economical operation of the gas meter.

CN116561679BActive Publication Date: 2025-07-08BEIJING GAS GRP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310514720.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-09
Publication Date
2025-07-08
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

Existing gas meter abnormal flow detection technology cannot be automatically updated with changes in user habits and the increase in gas meter service time, and requires a large number of manual marking of training samples, resulting in inefficient recognition and high cost.

Method used

The abnormal flow recognition system at the gas meter terminal based on reinforcement learning is adopted, and the decision-making module, feature extraction module and value generation module of the reinforcement learning framework are used to extract and make decisions through the status data sequence at the gas meter terminal, and the identification model is automatically updated to avoid manual data screening and labeling.

Benefits of technology

It realizes automated abnormal traffic recognition without manual data labeling, improves identification efficiency and accuracy, reduces operating costs, and ensures the safety and economic benefits of gas companies and users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116561679B_ABST
    Figure CN116561679B_ABST
Patent Text Reader

Abstract

The present invention provides a system for identifying abnormal flow at a gas meter end based on reinforcement learning. The system adopts a reinforcement learning framework, including a decision module, a feature extraction module and a value generation module as an intelligent agent; the feature extraction module extracts feature vectors based on the state data sequence output by the gas meter end, and the decision module outputs a decision value indicating whether the gas flow is normal, and each decision value is used as an action performed by the intelligent agent; the value generation module outputs the value created by the decision value; the value is used as the expected benefit, and the maximum expected benefit is obtained through reinforcement learning. By adopting a reinforcement learning framework, the present invention does not need to manually screen and label the data, and only needs to call the gas meter end's own data and the gas business entity's own historical data for model training to identify whether abnormal flow occurs at the gas meter end.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gas consumption monitoring, and particularly relates to an abnormal flow recognition system at the gas meter end based on reinforcement learning. Background Art

[0002] User gas meters are used to measure the gas consumption of users and are one of the most important measuring instruments in people's livelihood measurement. At present, phenomena such as gas pipeline breakage and overpressure / undervoltage, hose aging and detachment causing gas leakage, gas appliance failure / damage, and its own measurement inaccuracy are all directly reflected as abnormal flow display readings at the gas meter end. Obviously, with the increasing number of gas users year by year and the impact of events such as the epidemic on the production and life of the majority of users, the above situations not only increase the operating costs of gas operators, cause property losses to gas users, but also bring non-negligible safety hazards to both parties. Therefore, only by effectively identifying the abnormal flow at the user gas meter end can measures be taken in a timely manner to ensure the vital interests of gas enterprises and residential gas users.

[0003] Existing abnormal flow detection technologies mainly include the following: One is the rule / threshold-based method, that is, for the instantaneous flow of gas at the meter end or the cumulative flow within a certain period of time, an overlimit value or overlimit range is preset according to industry experience. When the measured data at the meter end exceeds this preset value or range, it is identified as abnormal flow. This method has simple and clear calculations and strong interpretability for the recognition results; however, the recognition rules are relatively single and rigid, and cannot automatically update the recognition rules with the evolution of user habits and the increase in the service time of gas meters. Another is the multi-variable classification-based method, that is, the identification of the normal / abnormal state of the flow is equivalent to a normal / abnormal binary classification problem. Multivariate variables such as the flow rate, pressure loss read at the meter end, and even the concentration of combustible gas and ambient temperature detected by the meter end sensors are used as input features to train a binary classification model (such as a support vector machine, multi-layer perceptron, etc.) to identify normal / abnormal flow. This method makes full use of the existing data, but it is necessary to repeatedly observe, retrieve, check, and label a large number of training samples for the model to learn, and even requires a large number of personnel to enter households for operations. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides an abnormal flow recognition system at the gas meter end based on reinforcement learning.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions.

[0006] An abnormal flow recognition system for gas meter terminals based on reinforcement learning. The system adopts a reinforcement learning framework, including a decision-making module, a feature extraction module, and a value generation module as agents. The inputs of the decision-making module and the feature extraction module are the state data sequences output by the gas meter terminals. The feature extraction module extracts feature vectors based on the state data sequences. The decision-making module outputs a decision value indicating whether the gas flow is normal, and each decision value is used as an action executed by the agent. The decision value output by the decision-making module and the feature vector output by the feature extraction module are fused and then input into the value generation module, and the value generation module outputs the value created by the decision value. Using the value as the expected return, the maximum expected return is obtained through reinforcement learning.

[0007] Furthermore, the state data sequence output by the gas meter terminal is:

[0008] S T ={s1, s2, …, s t , …, s T}

[0009] In the formula, S T is the state data sequence, s t is the state data of the t-th data acquisition period, t = 1, 2, …, T, and T is the length of the state data sequence.

[0010] Even further, the s t at least includes instantaneous flow rate, average flow rate within a fixed time, gas pressure, ambient combustible gas concentration, instantaneous ambient temperature, and average ambient temperature within a fixed time.

[0011] Even further, the decision-making module is composed of a first recurrent neural network, and the calculation method of its output decision value includes:

[0012] Set the initial hidden variable h0 of the first recurrent neural network as a vector of all 0s;

[0013] Calculate the t-th hidden variable h t according to the following formula:

[0014] h t =f1(U A s t-1 +W A h t-1 +b A )

[0015] In the formula, U A , W A , b A are trainable parameters, f1() is an activation function, t = 1, 2, …, T;

[0016] Calculate the decision value output by the decision-making module according to the following formula:

[0017] α = f1(V A h T + c A )

[0018] where α is the decision value, V A , c A are trainable parameters.

[0019] Furthermore, the feature extraction module consists of a second recurrent neural network, and the calculation method of its output feature vector includes:

[0020] Set the initial hidden variable e0 of the second recurrent neural network to a vector of all zeros;

[0021] Calculate the t-th hidden variable e according to the following formula t :

[0022] e t = f2(U F s t-1 + W F e t-1 + b F )

[0023] where U F , W F , b F are trainable parameters, f2() is the activation function, and t = 1, 2,..., T;

[0024] Calculate the feature vector output by the feature extraction module according to the following formula:

[0025]

[0026] where f T is the feature vector output by the feature extraction module.

[0027] Furthermore, the first recurrent neural network and the second recurrent neural network are long short-term memory networks.

[0028] Furthermore, the activation functions f1(), f2() are sigmoid or ReLU functions.

[0029] Further, the decision value α output by the decision module is inserted at the first position of the feature vector f T output by the feature extraction module, and the fusion vector [α, f T is obtained.

[0030] Furthermore, the output Q T of the value generation module is:

[0031] Q T = WQ [α,f T ]+b Q

[0032] Where W Q , b Q is a trainable parameter.

[0033] Furthermore, the recognition system is optimized using the stochastic gradient descent method, and the optimization objective function is:

[0034]

[0035] In the formula, G T The target true value of the value generation module is calculated by:

[0036] If S T If the daily operation and maintenance records within the covered time period are consistent with the output of the decision module, then:

[0037]

[0038] If S T If the daily operation and maintenance records within the covered time period are inconsistent with the output of the decision module, then:

[0039]

[0040] Where, L T is the service time of the gas meter at the current moment, Y is the maximum service life, and R is the accumulated recharge income of the gas meter at the current moment.

[0041] Compared with the prior art, the present invention has the following beneficial effects.

[0042] The system of the present invention adopts a reinforcement learning framework, including a decision module as an intelligent agent, a feature extraction module and a value generation module. The decision module outputs a decision value indicating whether the gas flow is normal, which is used as an action for the intelligent agent to perform. The feature extraction module outputs an extracted feature vector. The value generation module takes the fusion vector of the decision value and the feature vector as input, outputs the value created by the decision value, takes the value as the expected benefit, and obtains the maximum expected benefit through reinforcement learning. By adopting a reinforcement learning framework, the system of the present invention does not need to manually screen and label the data. It only needs to call the gas meter end's own data and the gas business entity's own historical data for model training to identify whether abnormal flow occurs at the gas meter end. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1It is a block diagram of a system for identifying abnormal gas flow at the gas meter end based on reinforcement learning according to an embodiment of the present invention. In the figure, 1 is a decision-making module, 2 is a feature extraction module, and 3 is a value generation module.

[0044] Figure 2 It is a schematic diagram of the composition of the decision-making module.

[0045] Figure 3 It is a schematic diagram of the composition of the feature extraction module. Detailed implementation manners

[0046] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0047] Figure 1 It is a block diagram of a system for identifying abnormal gas flow at the gas meter end based on reinforcement learning according to an embodiment of the present invention. The system adopts a reinforcement learning framework, including a decision-making module 1, a feature extraction module 2 and a value generation module 3 as agents; the inputs of the decision-making module 1 and the feature extraction module 2 are the state data sequences output from the gas meter end, the feature extraction module 2 extracts feature vectors based on the state data sequences, the decision-making module 1 outputs a decision value indicating whether the gas flow is normal, and each decision value is used as an action executed by the agent; the decision value output by the decision-making module 1 and the feature vector output by the feature extraction module 2 are fused and then input into the value generation module 3, and the value generation module 3 outputs the value created by the decision value; using the value as the expected return, the maximum expected return is obtained through reinforcement learning.

[0048] This embodiment provides an abnormal flow recognition system based on reinforcement learning at the gas meter end. Reinforcement learning (RL), also known as enhanced learning, is one of the paradigms and methodologies of machine learning, which is used to describe and solve the problem that an intelligent agent maximizes the reward or achieves a specific goal by learning strategies during the interaction with the environment. Reinforcement learning belongs to unsupervised machine learning and includes five core components: Environment, Agent, State, Action, and Reward. Reinforcement learning regards learning as a trial and evaluation process. The intelligent agent selects an action to act on the environment. After the environment accepts the action, the state changes, and at the same time, a reinforcement signal (reward value) is generated and fed back to the intelligent agent. The intelligent agent then selects the next action based on the reinforcement signal and the current state of the environment. The selection principle is to maximize the cumulative reward value or expected return obtained. The above is the principle of classical reinforcement learning. However, the characteristic of the application scenario of this embodiment is that the normal data and abnormal data are significantly unbalanced, and the sample size of normal data is much larger than that of abnormal data. In this case, the branch DQN (Deep Q Learning) of reinforcement learning is more effective. Therefore, the reinforcement learning adopted in this embodiment is actually DQN.

[0049] In this embodiment, the system follows the typical implementation method of the reinforcement learning framework and mainly consists of three independent neural network modules, namely, the decision-making module 1, the feature extraction module 2, and the value generation module 3. The decision-making module 1 is the intelligent agent of the reinforcement learning framework (i.e., the model for identifying whether the gas flow is normal or abnormal). Its input is the sequence of state data collected in real time from the gas meter end according to the set collection period, and the output is the decision value. The decision value is the recognition result indicating whether the current gas flow is normal or abnormal. The recognition result (normal or abnormal) of this embodiment can be regarded as the action executed by the intelligent agent. If the recognition result is correct, no economic loss will be caused; if the recognition result is incorrect, economic loss will be caused. Whether economic loss is caused can be regarded as the external feedback. The feature extraction module 2 is used to extract features based on the sequence of state data at the gas meter end. The extracted feature vector is fused with the decision value and then input into the value generation module 3. The value generation module 3 calculates the value generated by the decision value based on the vector after the fusion of the feature vector and the decision value. Taking the value as the expected return, the maximum expected return is obtained through reinforcement learning.

[0050] In this embodiment, the recognition model for whether the gas flow is normal is trained by combining the service status of the gas meter itself, the remotely transmitted data of the gas meter, the verification results of the abnormal flow report data of the gas meter, the daily maintenance data of the gas meter, and the user payment data in the existing information management system of the gas industry. This process can run through the entire service process of the gas meter and keep the recognition model automatically updated all the time. This embodiment neither needs to introduce additional operations into the existing operation and maintenance process of the gas enterprise, nor does it require the technical personnel, service personnel, and residential users in the gas industry to report any additional data, and at the same time, it does not need to manually collect and label the training samples of the model.

[0051] As an alternative embodiment, the state data sequence output by the gas meter end is:

[0052] S T ={s1, s2, …, s t , …, s T}

[0053] In the formula, S T is the state data sequence, and s t is the state data of the t-th data acquisition cycle, where t = 1, 2, …, T, and T is the length of the state data sequence.

[0054] This embodiment gives the expression of the state data sequence S T output by the gas meter end. In this embodiment, S T is composed of T state data, and the state data of the t-th data acquisition cycle is s t . The acquisition cycle of the state data is the acquisition time interval between two adjacent state data s t-1 and s t . The smaller the acquisition cycle, the higher the data processing accuracy, but the greater the computational amount of data processing, which is not conducive to realizing real-time processing. Therefore, the size of the acquisition cycle should be selected in a trade-off manner. The state data includes characteristic data related to the identification of abnormal gas flow, such as gas flow, pressure, etc. A specific state data will be given in the following embodiments.

[0055] As an alternative embodiment, the s t at least includes instantaneous flow rate, average flow rate within a fixed time, gas pressure, ambient gas concentration, instantaneous ambient temperature, and average ambient temperature within a fixed time.

[0056] This embodiment provides a specific state data. As described above, the state data includes feature data related to the identification of abnormal gas flow. To identify whether the gas flow is normal, at least the magnitude of the gas flow needs to be known. Therefore, the state data should at least include the gas flow. Additionally, auxiliary identification can be performed based on gas pressure, ambient gas concentration (related to leakage), and ambient temperature. Therefore, the state data provided in this embodiment includes gas flow, gas pressure, and ambient gas concentration. The gas flow includes both instantaneous flow and average flow within a set fixed time. The ambient temperature includes both instantaneous ambient temperature and average ambient temperature within a set fixed time.

[0057] As an optional embodiment, the decision-making module 1 is composed of a first recurrent neural network, and the calculation method of its output decision value includes:

[0058] Set the initial hidden variable h0 of the first recurrent neural network as a vector of all 0s;

[0059] Calculate the t-th hidden variable h according to the following formula t :

[0060] h t = f1(U A s t-1 + W A h t-1 + b A )

[0061] In the formula, U A , W A , b A are trainable parameters, f1() is an activation function, t = 1, 2,..., T;

[0062] Calculate the decision value output by the decision-making module 1 according to the following formula:

[0063] α = f1(V A h T + c A )

[0064] In the formula, α is the decision value, V A , c A are trainable parameters.

[0065] This embodiment provides a network structure of the decision-making module 1 and a calculation method for decision values. Since the decision-making module 1 in this embodiment performs abnormal flow recognition based on the state data sequence output from the gas meter end, the decision-making module 1 in this embodiment mainly consists of a recurrent neural network capable of processing time series or sequence data. A recursive neural network (RNN) is also known as a generalization of a recurrent neural network (RNN) (when each parent node of the recursive neural network is only connected to one child node, its structure is equivalent to a fully connected recurrent neural network), and it is an artificial neural network with a tree-like hierarchical structure and the network nodes recursively process the input information according to their connection order. In a feedforward neural network, the information transmission is unidirectional, and each input is independent, that is, the output of the network only depends on the current input. However, in many real-world tasks, the output of the network not only depends on the current input but also is related to the output in the past period. An RNN is a type of neural network with short-term memory function, and its neurons can not only receive information from other neurons but also receive their own information, forming a network structure with loops. Therefore, an RNN can be used to process time series data. The structure of the first recurrent neural network adopted by the decision-making module 1 in this embodiment is as Figure 2 shown, mainly consisting of T cascaded neural network modules, namely hidden layers, and the output of each hidden layer corresponds to a hidden variable. The calculation formulas for each hidden variable and the decision value output by the decision-making module 1 are as above, and no detailed description will be given here.

[0066] As an optional embodiment, the feature extraction module 2 consists of a second recurrent neural network, and the calculation method for its output feature vector includes:

[0067] Set the initial hidden variable e0 of the second recurrent neural network to a vector of all 0s;

[0068] Calculate the t-th hidden variable e according to the following formula t :

[0069] e t = f2(U F s t-1 + W F e t-1 + b F )

[0070] In the formula, U F , W F , b F are trainable parameters, f2() is an activation function, and t = 1, 2,..., T;

[0071] Calculate the feature vector output by the feature extraction module 2 according to the following formula:

[0072]

[0073] In the formula, f T is the feature vector output by the feature extraction module 2.

[0074] This embodiment provides a network structure of the feature extraction module 2 and a calculation method for the output feature vector. Since the processing object of the feature extraction module 2 is still the state data sequence output from the gas meter end, the feature extraction module 2 is mainly composed of a recurrent neural network, namely the second recurrent neural network. The structure of the second recurrent neural network is as Figure 3 shown, different from Figure 2 which, the output feature vector is equal to the mean of T hidden variables.

[0075] As an optional embodiment, the first recurrent neural network and the second recurrent neural network are long short-term memory networks.

[0076] This embodiment provides a specific recurrent neural network. In this embodiment, both the first recurrent neural network and the second recurrent neural network adopt the long short-term memory network LSTM (Long Short-Term Memory). LSTM is a time recurrent neural network, which is specifically designed to solve the problem of long-term dependence existing in general RNNs. All RNNs have a chain form of repeating neural network modules. However, RNNs have problems of gradient vanishing or explosion during the learning process, and it is difficult to model the dependencies between states with long time intervals. To improve the long-range dependence problem of RNNs, a good solution is to introduce a gating mechanism to control the accumulation speed of information, including selectively adding new information and selectively forgetting the previously accumulated information. LSTM is such an RNN that introduces an input gate, a forget gate, and an output gate. The input gate is used to control how much information of the candidate state at the current moment needs to be saved; the forget gate is used to control how much information of the internal state at the previous moment needs to be forgotten; the output gate is used to control how much information of the internal state at the current moment needs to be output to the external state.

[0077] As an optional embodiment, the activation functions f1() and f2() are sigmoid or ReLU functions.

[0078] This embodiment provides specific functions that the activation functions f1() and f2() can adopt. The activation function is a function that runs on the neurons of an artificial neural network and is used to map the input of a neuron to the output end. The introduction of the activation function is to increase the non-linearity of the neural network model. Without an activation function, the output of each layer is a linear function of the upper-layer input. No matter how many layers the neural network has, the output is a linear combination of the inputs. The activation function introduces non-linear factors into the neurons, enabling the neural network to approximate any non-linear function arbitrarily, so that the neural network can be applied to many non-linear models.

[0079] There are many available activation functions, and the commonly used ones include the Sigmoid function, Tanh function, and ReLU function, etc. The activation functions f1() and f2() in this embodiment are the sigmoid or ReLU function.

[0080] The expression of the sigmoid function is:

[0081] f(x) = (1 + e -x ) -1 ;

[0082] The expression of the ReLU function is:

[0083] f(x) = max(0, x).

[0084] As an alternative embodiment, the decision value α output by the decision module 1 is inserted at the first position of the feature vector f output by the feature extraction module 2 to obtain the fusion vector [α, f T . T

[0085] This embodiment provides a fusion scheme for the decision value and the feature vector. The fusion scheme given in this embodiment can be regarded as splicing. Specifically, it is to insert the decision value α at the first position of the feature vector f T to obtain the fusion vector [α, f T .

[0086] As an alternative embodiment, the output Q T of the value generation module 3 is:

[0087] Q T = W Q [α, f T + b Q

[0088] In the formula, W Q , b Q are trainable parameters.

[0089] This embodiment provides the output Q of the value generation module 3T The calculation formula. The input of the value generation module 3 is the fusion vector [α, f T , and its output is the weight matrix W Q multiplied by [α, f T and then adding a bias b Q .

[0090] As an optional embodiment, the recognition system is optimized by the stochastic gradient descent method, and the optimization objective function is:

[0091]

[0092] In the formula, G T is the target true value of the value generation module 3, and the calculation method includes:

[0093] If the daily operation and maintenance records within the time period covered by S T are consistent with the output of the decision module, then:

[0094]

[0095] If the daily operation and maintenance records within the time period covered by S T are inconsistent with the output of the decision module, then:

[0096]

[0097] In the formula, L T is the service time at the gas meter end at the current moment, Y is the longest service life, and R is the cumulative recharge income at the gas meter end at the current moment.

[0098] This embodiment gives a specific optimization objective function for optimizing and training the recognition model. This embodiment uses the stochastic gradient descent method to optimize the recognition model, and the optimization objective function or loss function is half of the square of the difference between the output Q T of the value generation module 3 and the target true value G T . This embodiment also gives the calculation method of the target true value G T . The size of the target true value G T is determined by the service time L T of the gas meter end, the longest service life Y, and the cumulative recharge income R (in yuan) of the gas meter end. These parameters can be directly obtained from the gas meter body and the background of the gas operator. The expression in the small parentheses in the above formula represents the percentage of the remaining service time of the gas meter in its available life, and this percentage is directly regarded as the revenue attenuation coefficient of the gas meter itself. R / L T represents the revenue created per unit time by the gas meter during its service, and the product of the two gives G T, indicating the revenue amount per unit time of the gas meter during its future service. G T The positive or negative of T is related to whether the daily operation and maintenance records during the covered time period are consistent with the output of the decision-making module 1 (consistency includes: the output of the decision-making module 1 is abnormal and the daily operation and maintenance records are also abnormal; the output of the decision-making module 1 is normal and the daily operation and maintenance records are also normal). If they are consistent, G T is positive; if they are inconsistent, G T is negative. The daily operation and maintenance is a periodic inspection, and each inspection has a record, which records the inspection results including whether the gas flow is abnormal.

[0099] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An abnormal flow recognition system at the gas meter end based on reinforcement learning, characterized in that, The system adopts a reinforcement learning framework, including a decision-making module, a feature extraction module, and a value generation module as agents; the inputs of the decision-making module and the feature extraction module are the state data sequences output from the gas meter end. The feature extraction module extracts feature vectors based on the state data sequences. The decision-making module outputs a decision value indicating whether the gas flow is normal, and each decision value is used as an action executed by the agent. The decision value output by the decision-making module and the feature vector output by the feature extraction module are fused and then input into the value generation module, and the value generation module outputs the value created by the decision value; taking the value as the expected return, the maximum expected return is obtained through reinforcement learning. The state data sequence output from the gas meter end is: S T = {s1, s2, …, s t , …, s T} where S T is the state data sequence, s t is the state data of the t-th data acquisition period, t = 1, 2, …, T, and T is the length of the state data sequence; The decision-making module is composed of a first recurrent neural network, and the calculation method of its output decision value includes: Set the initial hidden variable h0 of the first recurrent neural network as a vector of all 0s; Calculate the t-th hidden variable h according to the following formula t :[[-END]] h t = f1(U A s t-1 + W A h t-1 + b A ) where U A , W A , b A are trainable parameters, f1() is an activation function, and t = 1, 2, …, T; Calculate the decision value output by the decision-making module according to the following formula: α = f1(V A h T + c A ) where α is the decision value, V A , c A are trainable parameters; The feature extraction module is composed of a second recurrent neural network, and the calculation method of its output feature vector includes: Set the initial hidden variable e0 of the second recurrent neural network as a vector of all 0s; Calculate the t-th latent variable e according to the following formula t :[[]]END]] e t = f2(U F s t-1 + W F e t-1 + b F ) where U F , W F , b F are trainable parameters, f2() is an activation function, and t = 1, 2, …, T; Calculate the feature vector output by the feature extraction module according to the following formula: where f T is the feature vector output by the feature extraction module; The decision value α output by the decision module is inserted into the first position of the feature vector f output by the feature extraction module T to obtain a fusion vector [α, f T ; The output Q of the value generation module T is as follows: Q T = W Q [α, f T + b Q Where, W Q , b Q are trainable parameters.

2. The abnormal flow recognition system at the gas meter end based on reinforcement learning according to claim 1, characterized in that, The said s t includes at least instantaneous flow rate, average flow rate within a fixed time, gas pressure, ambient combustible gas concentration, instantaneous ambient temperature, and average ambient temperature within a fixed time.

3. The abnormal flow recognition system at the gas meter end based on reinforcement learning according to claim 1, characterized in that The first recurrent neural network and the second recurrent neural network are long short-term memory networks.

4. The abnormal flow recognition system at the gas meter end based on reinforcement learning according to claim 1, wherein The activation functions f1() and f2() are sigmoid or ReLU functions.

5. The abnormal flow recognition system at the gas meter end based on reinforcement learning according to claim 1, characterized in that, The recognition system uses the stochastic gradient descent method for optimization, and the optimization objective function is: where G T is the target true value of the value generation module, and the calculation method includes: If S T the daily operation and maintenance records during the covered time period are consistent with the output of the decision-making module, then: If S T the daily operation and maintenance records during the covered time period are inconsistent with the output of the decision-making module, then: where L T is the service time at the gas meter end at the current moment, Y is the longest service life, and R is the cumulative recharge income at the gas meter end at the current moment.

Citation Information

Patent Citations

  • Malicious software behavior detection and classification system based on deep learning

    CN113961922A