A reinforcement learning-based operation optimization system for waste gas treatment equipment

CN122568945APending Publication Date: 2026-08-14JIANGSU BAMA NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

传统废气处理控制方式主要依据单一时刻的污染物浓度或设备运行参数进行控制,缺乏对废气处理过程中多处理段之间关联关系的动态建模能力,难以准确反映污染物在不同处理段之间的传播过程与演化规律,导致复杂工况下容易出现控制滞后与处理效率波动的问题;已有时序分析方法大多采用普通循环神经网络、长短期记忆网络或传统Transformer结构对运行数据进行建模,无法有效识别废气处理过程中存在的工况相变特征以及异常传播链结构,对于长时间尺度与短时间尺度运行状态之间的关联表达能力不足,限制了复杂工况下的时序特征提取能力;此外,现有强化学习控制方法通常直接根据状态数据输出控制动作,缺少对污染传播路径、处理段干预顺序以及工况演化过程的关联分析机制,难以实现多处理段协同控制场景下的动态优化调节,导致设备能耗与污染物处理效率之间难以保持稳定平衡

Benefits of technology

本发明通过构建包含处理段关联拓扑结构、多变量运行时间窗口序列以及工况传播链序列的废气处理运行表征体系,结合改进型TS2Vec模型与强化学习决策网络的协同设计,针对传统废气处理控制过程中处理段关联关系难以动态表达、复杂工况传播过程难以建模以及多时间尺度运行特征难以统一表征的问题,提出基于工况相变检测、传播链生成以及演化分支编码的工况演化建模策略,显著提高复杂废气处理场景下污染传播路径与运行状态变化过程的表征能力;在时序编码阶段引入处理段拓扑嵌入层、工况传播链生成层以及多尺度折叠编码层,通过处理段连接关系建模、传播节点连接关系构建以及不同时间尺度特征映射,实现不同处理段之间传播关系与长短周期运行特征之间的一致关联表达;在强化学习决策阶段构建包含工况传播链策略编码层、处理段干预顺序生成层、动作原语组合层以及贡献反馈评价层的强化学习决策网络,结合异常传播链中的处理段连接顺序与处理段贡献变化量生成候选动作序列,并依据动作评分值动态调整设备控制动作,提高复杂工况下不同处理段之间的协同控制能力;最终利用运行反馈模块对控制后的运行状态偏移序列、工况传播链序列以及工况演化轨迹特征执行动态迭代更新,实现废气处理设备运行过程中的工况传播感知、多尺度演化建模以及智能动态优化控制。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122568945A_ABST
    Figure CN122568945A_ABST
Patent Text Reader

Abstract

This invention discloses a reinforcement learning-based system for optimizing the operation of waste gas treatment equipment, comprising: a data acquisition module for collecting operational data of the waste gas treatment equipment and generating a state sequence of the treatment section; a time series construction module for generating a multivariate operating time window sequence and a topological structure associated with the treatment section; a condition characterization module for generating multi-scale operating condition evolution features using an improved TS2Vec model; a contribution calculation module for generating the contribution coefficient of the treatment section and reinforcement learning reward data; a reinforcement learning decision module for generating a sequence of equipment control actions using a reinforcement learning decision network; an equipment control module for controlling the operation of the waste gas treatment equipment; and an operation feedback module for updating the multivariate operating time window sequence and the multi-scale operating condition evolution features. This invention improves the stability and dynamic optimization capability of waste gas treatment control under complex operating conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and waste gas treatment technology, and in particular to a waste gas treatment equipment operation optimization system based on reinforcement learning. Background Technology

[0002] With the increasing demands for industrial waste gas treatment and the continuous strengthening of volatile organic compound (VOC) emission regulations, waste gas treatment equipment has been widely used in industrial scenarios such as chemical engineering, electronics manufacturing, spray painting, and new energy material production. Regarding parameter adjustment and operational optimization during the operation of waste gas treatment equipment, existing technologies mainly employ fixed rule control, empirical threshold control, and parameter optimization methods based on traditional machine learning. However, the following problems commonly exist in actual industrial operation scenarios: Traditional waste gas treatment control methods primarily rely on pollutant concentrations or equipment operating parameters at a single moment, lacking the ability to dynamically model the relationships between multiple treatment stages during the waste gas treatment process. This makes it difficult to accurately reflect the propagation process and evolution of pollutants between different treatment stages, leading to control lag and fluctuations in treatment efficiency under complex operating conditions. Existing time series analysis methods mostly use ordinary recurrent neural networks, long short-term memory networks, or traditional Transformer structures to model operating data, failing to effectively identify the phase transition characteristics and abnormal propagation chain structures present in the waste gas treatment process. They also lack the ability to express the correlation between long-term and short-term operating states, limiting the ability to extract time series features under complex operating conditions. Furthermore, existing reinforcement learning control methods typically output control actions directly based on state data, lacking a correlation analysis mechanism for pollution propagation paths, treatment stage intervention sequences, and operating condition evolution processes. This makes it difficult to achieve dynamic optimization and adjustment in multi-treatment stage collaborative control scenarios, resulting in an inability to maintain a stable balance between equipment energy consumption and pollutant treatment efficiency.

[0003] Therefore, how to provide a system for optimizing the operation of waste gas treatment equipment based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] One objective of this invention is to propose an operation optimization system for waste gas treatment equipment based on reinforcement learning. This invention constructs a topological structure of the treatment section association, a sequence of operating condition propagation chains, and multi-scale operating condition evolution characteristics. Combining an improved TS2Vec model with a reinforcement learning decision network, it proposes an operation optimization method for waste gas treatment based on operating condition phase change detection, propagation chain generation, and evolutionary branch coding. This method enables modeling of pollution propagation processes under complex operating conditions, collaborative control of treatment sections, and dynamic optimization of equipment operating status, thereby improving the stability of waste gas treatment and the accuracy of operation control.

[0005] An embodiment of the present invention provides a reinforcement learning-based operation optimization system for waste gas treatment equipment, comprising: The data acquisition module is used to collect the operating data of each treatment section of the waste gas treatment equipment and generate the corresponding treatment section status sequence. The time series construction module is used to splice the state sequences of processing segments at multiple consecutive time points to generate a multivariable runtime window sequence and construct the processing segment association topology. The operating condition characterization module is used to input the multivariate running time window sequence and the processing segment associated topology into the improved TS2Vec model, perform operating condition phase transition detection on the multivariate running time window sequence, generate operating condition phase transition features, and generate multi-scale operating condition evolution features based on the operating condition phase transition features. The contribution calculation module is used to calculate the contribution coefficient of the processing segment based on the multi-scale working condition evolution characteristics, and generate reinforcement learning reward data based on the contribution coefficient of the processing segment. The reinforcement learning decision module is used to input the multi-scale operating condition evolution characteristics and reinforcement learning reward data into the reinforcement learning decision network to generate a sequence of equipment control actions. The equipment control module is used to control the operation of the waste gas treatment equipment according to the equipment control action sequence; The operation feedback module is used to update the multivariable operation time window sequence and multi-scale operating condition evolution characteristics based on the controlled operation data.

[0006] Optionally, the data acquisition module includes: Waste gas flow rate data and spray flow rate data are collected at the locations of the intake pipe, spray pipe and discharge pipe along the direction of waste gas transportation, and flow time series data are generated according to the collection time. Volatile organic compound concentration data, particulate matter concentration data, and gas component concentration data were collected at the inlet and outlet of each treatment section, and concentration time series data were established according to the treatment section number. Temperature, humidity, differential pressure, fan frequency, valve opening, and heating power data are collected at the locations of the fan, reactor, spray tower, and adsorption device, and corresponding state time series data are established. The flow rate time series data, the concentration time series data, and the state time series data are aligned on the time axis. A unified sampling sequence is established according to a unified sampling time interval, and linear interpolation is performed on positions where the time interval exceeds the sampling interval threshold. The parameters of the unified sampling sequence are combined according to the processing section number to establish a processing section status sequence that includes flow rate parameters, concentration parameters, temperature and humidity parameters, differential pressure parameters, and control parameters. Fluctuation amplitude detection is performed on the state sequence of the processing segment, the absolute value of the data difference between adjacent sampling times is calculated, and data segments whose absolute value of the data difference exceeds the fluctuation threshold are marked as abnormal segments; The number of consecutive abnormal sampling points is counted, and data segments with a number of consecutive abnormal sampling points exceeding the abnormality duration threshold are deleted, while the corresponding time position markers are retained.

[0007] Optionally, the timing construction module includes: The processing segment state sequences are sorted according to the acquisition time to form a single-segment time series matrix corresponding to each processing segment; Using the current sampling time as the end point of the window, select single-segment time series matrices of multiple consecutive sampling times forward, and splice them together to form the local time window matrix corresponding to each processing segment; The local time window matrices corresponding to each treatment section are spliced ​​together according to the exhaust gas flow direction and the connection order of the treatment sections to generate a multivariate running time window sequence; Calculate the consistency rate of pollutant concentration change direction, pressure difference change direction, and energy consumption change direction for any two treatment segments within the same time window. The consistency rates of the pollutant concentration change direction, the pressure difference change direction, and the energy consumption change direction are weighted and fused to calculate the correlation strength of the processing segment; Each processing segment is used as a topology node, and the association strength of the processing segment is used as the weight of the topology edge to construct the processing segment association topology structure.

[0008] Optionally, the operating condition characterization module includes: The improved TS2Vec model includes a processing segment topology embedding layer, a shared temporal coding network, a load condition phase transition detection layer, a load condition propagation chain generation layer, a load condition evolution coding layer, and a multi-scale folding coding layer. The processing segment topology embedding layer performs node embedding processing on each processing segment node in the processing segment associated topology structure, and establishes a topology adjacency sequence according to the processing segment connection relationship; The multivariate runtime window sequence and the topological adjacency sequence are input into a shared temporal coding network, and linear mapping processing is performed on the time window features corresponding to different processing segments to generate temporal embedding features of the processing segments. The operating condition phase change detection layer performs differential calculation on the temporal embedding features of the processing segment between adjacent sampling times to generate a gradient sequence of operating condition changes; Sliding aggregation is performed on the gradient sequence of the operating condition change according to the sampling time order, the number of gradient changes exceeding the gradient threshold within the local time window is counted, and the sampling points that continuously exceed the gradient threshold are formed into phase transition candidate regions. Phase transition candidate regions where the number of gradient changes exceeds the phase transition threshold are marked as operating condition phase transition regions, and the processing segment temporal embedding features corresponding to the operating condition phase transition regions are extracted as operating condition phase transition features. The operating condition propagation chain generation layer performs correlation matching processing on the direction of pollutant concentration change, pressure difference change, and energy consumption change in each processing segment during continuous sampling time. It establishes the connection relationship of the propagation nodes of the processing segments according to the consistency of the change direction, and constructs the operating condition propagation chain sequence based on the connection relationship of the propagation nodes of the processing segments. The propagation length and the number of changes in the connection direction information of the propagation nodes in the working condition propagation chain sequence are statistically analyzed at continuous sampling time, and the working condition propagation chain sequence whose propagation length exceeds the propagation threshold and whose number of changes in the connection direction information of the propagation nodes exceeds the direction change threshold is marked as an abnormal propagation chain. The operating condition evolution coding layer uses the operating condition phase transition features and the anomaly propagation chain as the initial state of evolution, and gradually splices the processing segment temporal embedding features corresponding to the continuous sampling time according to the connection order of the propagation nodes to generate multiple operating condition evolution branch sequences. Evolution direction encoding is performed on each working condition evolution branch sequence, the embedding feature difference between adjacent sampling times is calculated, and the embedding feature difference is accumulated in the order of sampling time to generate working condition evolution trajectory features; The multi-scale folded coding layer performs hierarchical sampling processing on the working condition evolution trajectory features according to short time scale, medium time scale and long time scale respectively, maps the sampled features under different time scales to the same embedding dimension, and splices them in the order of time scale to generate multi-scale working condition evolution features.

[0009] Optionally, the contribution calculation module includes: Extract the pollutant concentration evolution characteristics, pressure difference evolution characteristics, energy consumption evolution characteristics, and propagation chain correlation characteristics corresponding to each treatment segment from the multi-scale operating condition evolution characteristics; The pollutant reduction amount of each treatment section is calculated based on the difference between the inlet and outlet pollutant concentrations of each treatment section and the corresponding exhaust gas flow rate of the treatment section. The operating energy consumption of each treatment section is calculated based on the fan power, spray pump power, and heating power of each treatment section. The basic contribution coefficient is generated by dividing the pollutant reduction amount by the sum of the operating energy consumption and the positive energy consumption smoothing parameter. Adjust the weight values ​​corresponding to the basic contribution coefficients according to the propagation chain association features, and generate the processing segment contribution coefficients corresponding to each processing segment. Calculate the difference in the processing segment contribution coefficient between the current control cycle and the previous control cycle, and generate the change in processing segment contribution. The changes in the contribution of the processing section, the evolution characteristics of the pollutant concentration, the evolution characteristics of the pressure difference, and the evolution characteristics of the energy consumption are fused and calculated to generate reinforcement learning reward data.

[0010] Optionally, the reinforcement learning decision module includes: The multi-scale working condition evolution features are arranged according to the execution order of the processing segment number to generate a working condition state input sequence; The reinforcement learning reward data is used to establish a reward time sequence according to the execution time of the control cycle, and then concatenated with the working condition input sequence at the corresponding positions to generate a propagation chain policy state vector. The propagation chain strategy state vector is input into the reinforcement learning decision network; The reinforcement learning decision network includes a working condition propagation chain policy encoding layer, a processing segment intervention sequence generation layer, an action primitive combination layer, and a contribution feedback evaluation layer. The working condition propagation chain strategy encoding layer performs propagation chain state encoding processing on the propagation chain strategy state vector to generate a processing segment propagation state sequence. The processing segment intervention order generation layer performs segmented scoring processing on the processing segment propagation state sequence according to the connection order of processing segments in the abnormal propagation chain, and generates the processing segment intervention order; The action primitive combination layer selects action primitives from the action primitive sets corresponding to fan frequency adjustment action, valve opening adjustment action, spray flow rate adjustment action, and heating power adjustment action according to the intervention order of the processing segment, and performs action splicing processing according to the intervention order of the processing segment to generate multiple candidate action sequences. The contribution feedback evaluation layer calculates the action score value based on the change in the contribution of the processing segment, the change in operating energy consumption, and the change in the evolution of pollutant concentration corresponding to each candidate action sequence, and sorts each candidate action sequence according to the size of the action score value. The candidate action sequence with the highest action score is selected as the device control action sequence.

[0011] Optionally, the device control module includes: The primitives of each action in the device control action sequence are parsed and processed according to the intervention order of the processing segment to generate the corresponding control parameter adjustment sequence; Based on the control parameter adjustment sequence, extract the fan frequency adjustment parameter, valve opening adjustment parameter, spray flow rate adjustment parameter, and heating power adjustment parameter respectively; The control parameter adjustment sequence for each treatment section is arranged according to the direction of waste gas flow, and a treatment section action execution queue is established. According to the action execution queue of the processing segment, the fan operating frequency, valve opening, spray flow rate and heating power in the corresponding processing segment are adjusted sequentially. After the control parameters of the current treatment section are adjusted, collect data on pollutant concentration changes, pressure difference changes, and energy consumption changes for the corresponding treatment section. Update the control parameter adjustment sequence for the remaining treatment sections based on the pollutant concentration change data, pressure difference change data, and energy consumption change data of the current treatment section; Continue executing the control actions corresponding to the remaining processing segments according to the updated control parameter adjustment sequence until all control actions corresponding to the device control action sequence are completed.

[0012] Optionally, the runtime feedback module includes: Collect pollutant concentration data, differential pressure data, temperature data, humidity data, fan frequency data, valve opening data, spray flow rate data, and heating power data after executing the equipment control action sequence; The controlled running data is time-aligned according to the acquisition time sequence and written into the current multivariate running time window sequence according to the acquisition time sequence. Delete historical sampling data that exceed the sampling window length threshold in the multivariate runtime window sequence, and generate an updated multivariate runtime window sequence; Calculate the changes in pollutant concentration, pressure difference, and energy consumption between the corresponding sampling times before and after control, and generate an operating state offset sequence; Update the propagation node connection relationship and propagation node connection direction information in the operating condition propagation chain sequence according to the operating state offset sequence; The updated operating condition propagation chain sequence is subjected to time-scale hierarchical processing, and the updated multivariate running time window sequence is combined with the updated operating condition evolution trajectory features to generate the features iteratively again. The corresponding multi-scale operating condition evolution features are updated based on the regenerated operating condition evolution trajectory features.

[0013] The beneficial effects of this invention are: This invention constructs a waste gas treatment operation representation system that includes a treatment segment association topology, a multivariate operating time window sequence, and an operating condition propagation chain sequence. Combining an improved TS2Vec model with a reinforcement learning decision network, it addresses the challenges of dynamically representing treatment segment relationships, modeling complex operating condition propagation processes, and uniformly representing multi-timescale operating characteristics in traditional waste gas treatment control processes. It proposes an operating condition evolution modeling strategy based on operating condition phase change detection, propagation chain generation, and evolutionary branch coding, significantly improving the representation capability of pollution propagation paths and operating state changes in complex waste gas treatment scenarios. In the time-series coding stage, a treatment segment topology embedding layer, an operating condition propagation chain generation layer, and a multi-scale folding coding layer are introduced. This is achieved through processing segment connection relationship modeling, propagation node connection relationship construction, and... Feature mapping at different time scales enables consistent representation of the propagation relationship between different processing segments and the operational characteristics of long and short cycles. In the reinforcement learning decision-making stage, a reinforcement learning decision network is constructed, comprising a working condition propagation chain strategy encoding layer, a processing segment intervention sequence generation layer, an action primitive combination layer, and a contribution feedback evaluation layer. Candidate action sequences are generated by combining the processing segment connection order and the change in processing segment contribution in the anomaly propagation chain. Based on the action score, the equipment control actions are dynamically adjusted to improve the collaborative control capability between different processing segments under complex operating conditions. Finally, the operation feedback module is used to dynamically iteratively update the post-control operating state offset sequence, the working condition propagation chain sequence, and the working condition evolution trajectory features, achieving working condition propagation perception, multi-scale evolution modeling, and intelligent dynamic optimization control during the operation of the waste gas treatment equipment. Attached Figure Description

[0014] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the structure of a waste gas treatment equipment operation optimization system based on reinforcement learning proposed in this invention; Figure 2 This is a schematic diagram of the structure of an improved TS2Vec model in a reinforcement learning-based waste gas treatment equipment operation optimization system proposed in this invention. Detailed Implementation

[0015] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0016] refer to Figures 1-2 A system for optimizing the operation of waste gas treatment equipment based on reinforcement learning, comprising: The data acquisition module is used to collect the operating data of each treatment section of the waste gas treatment equipment and generate the corresponding treatment section status sequence. The time series construction module is used to splice the state sequences of processing segments at multiple consecutive time points to generate a multivariable runtime window sequence and construct the processing segment association topology. The operating condition characterization module is used to input the multivariate running time window sequence and the associated topology of the processing segment into the improved TS2Vec model, perform operating condition phase transition detection on the multivariate running time window sequence, generate operating condition phase transition features, and generate multi-scale operating condition evolution features based on the operating condition phase transition features. The contribution calculation module is used to calculate the contribution coefficient of the processing segment based on the multi-scale working condition evolution characteristics, and to generate reinforcement learning reward data based on the contribution coefficient of the processing segment. The reinforcement learning decision module is used to input multi-scale operating condition evolution characteristics and reinforcement learning reward data into the reinforcement learning decision network to generate equipment control action sequences; The equipment control module is used to control the operation of the waste gas treatment equipment according to the sequence of equipment control actions; The operation feedback module is used to update the multivariable operating time window sequence and multi-scale operating condition evolution characteristics based on the controlled operating data.

[0017] In this embodiment, the data acquisition module includes: Waste gas flow rate data and spray flow rate data are collected at the locations of the intake pipe, spray pipe and discharge pipe along the direction of waste gas transportation, and flow time series data are generated according to the collection time. Volatile organic compound concentration data, particulate matter concentration data, and gas component concentration data were collected at the inlet and outlet of each treatment section, and concentration time series data were established according to the treatment section number. Temperature, humidity, differential pressure, fan frequency, valve opening, and heating power data are collected at the locations of the fan, reactor, spray tower, and adsorption device, and corresponding state time series data are established. Time axis alignment processing is performed on flow time series data, concentration time series data and state time series data, a unified sampling sequence is established according to a unified sampling time interval, and linear interpolation processing is performed on positions where the time interval exceeds the sampling interval threshold; The parameters of the unified sampling sequence are combined according to the processing section number to establish a processing section status sequence that includes flow rate parameters, concentration parameters, temperature and humidity parameters, differential pressure parameters, and control parameters. Fluctuation amplitude detection is performed on the state sequence of the processing segment, the absolute value of the data difference between adjacent sampling times is calculated, and data segments whose absolute value of the data difference exceeds the fluctuation threshold are marked as abnormal segments; The number of consecutive abnormal sampling points is counted, and data segments with a number of consecutive abnormal sampling points exceeding the abnormality duration threshold are deleted, while the corresponding time position markers are retained.

[0018] In this implementation, electromagnetic flowmeters are installed at the inlet, spray, and outlet pipelines, with a sampling interval of 2 seconds for exhaust gas flow and 5 seconds for spray flow. Volatile organic compound (VOC) and particulate matter (PM) concentration sensors are installed at the inlet and outlet of each treatment section, with a sampling range of 0 ppm to 5000 ppm for VOC concentration and 0 mg / m³ to 300 mg / m³ for PM concentration. Vibration and frequency sensors are installed at the fan location, temperature and humidity sensors at the spray tower location, differential pressure sensors at the adsorption device location, and temperature and power sensors at the reactor location. A uniform sampling interval of 5 seconds is used; linear interpolation is performed when the time difference between adjacent sampling points exceeds 8 seconds. Fluctuation thresholds are dynamically generated based on the mean and standard deviation of 30 consecutive sampling points; data differences exceeding 2.5 times the corresponding standard deviation are marked as abnormal segments. When the number of consecutive abnormal sampling points reaches 6, the corresponding abnormal data segment is deleted, but the abnormal time and location markers are retained.

[0019] In this embodiment, the timing construction module includes: The processing segment state sequences are sorted according to the acquisition time to form a single-segment time series matrix corresponding to each processing segment; Using the current sampling time as the end point of the window, select single-segment time series matrices of multiple consecutive sampling times forward, and splice them together to form the local time window matrix corresponding to each processing segment; The local time window matrices corresponding to each treatment section are spliced ​​together according to the exhaust gas flow direction and the connection order of the treatment sections to generate a multivariate running time window sequence; Calculate the consistency rate of pollutant concentration change direction, pressure difference change direction, and energy consumption change direction for any two treatment segments within the same time window. The consistency rates of pollutant concentration change direction, pressure difference change direction, and energy consumption change direction are weighted and fused to calculate the correlation strength of the processed segments. Each processing segment is used as a topology node, and the association strength of the processing segments is used as the weight of the topology edges to construct the association topology structure of the processing segments.

[0020] In this implementation, the single-segment time series matrix is ​​arranged according to a 5-second sampling time interval, and each single-segment time series matrix contains 60 consecutive sampling moments, corresponding to 300 seconds of running data; the local time window matrix is ​​constructed using a sliding window method, with 30 overlapping sampling moments retained between adjacent windows; the multivariate running time window sequence is connected sequentially to the local time window matrices corresponding to the pretreatment section, spray section, adsorption section, and emission section according to the exhaust gas flow direction; the consistency rate of pollutant concentration change direction is calculated by the proportion of the number of sampling points with the same concentration change direction within the same time window to the total number of sampling points, and the consistency rate of pressure difference change direction and energy consumption change direction are calculated in the same way; when the consistency rate of change direction exceeds 0.75, an association edge is established between the corresponding treatment sections; the weight of the consistency rate of pollutant concentration change direction is set to 0.5, the weight of the consistency rate of pressure difference change direction is set to 0.3, and the weight of the consistency rate of energy consumption change direction is set to 0.2; when the association strength of the treatment section is lower than 0.35, the corresponding topological edge is deleted, and when it is higher than 0.65, the weight level of the corresponding topological edge is increased.

[0021] In this embodiment, the working condition characterization module includes: The improved TS2Vec model includes a processing segment topology embedding layer, a shared temporal coding network, a load condition phase transition detection layer, a load condition propagation chain generation layer, a load condition evolution coding layer, and a multi-scale folding coding layer. The processing segment topology embedding layer performs node embedding processing on each processing segment node in the processing segment associated topology structure, and establishes a topology adjacency sequence according to the processing segment connection relationship; The multivariate runtime window sequence and the topological adjacency sequence are input into a shared temporal coding network. Linear mapping is performed on the time window features corresponding to different processing segments to generate temporal embedding features of the processing segments. The operating condition phase change detection layer performs differential calculation on the temporal embedding features of the processing segment between adjacent sampling times to generate a gradient sequence of operating condition changes; Sliding aggregation is performed on the gradient sequence of operating conditions according to the sampling time order. The number of gradient changes exceeding the gradient threshold within the local time window is counted, and the sampling points that continuously exceed the gradient threshold are formed into phase transition candidate regions. Phase transition candidate regions where the number of gradient changes exceeds the phase transition threshold are marked as operating condition phase transition regions, and the processing segment temporal embedding features corresponding to the operating condition phase transition regions are extracted as operating condition phase transition features. The working condition propagation chain generation layer performs correlation matching processing on the direction of pollutant concentration change, pressure difference change, and energy consumption change in each treatment segment during continuous sampling time. It establishes the connection relationship of propagation nodes in the treatment segment according to the consistency of the change direction, and constructs the working condition propagation chain sequence based on the connection relationship of propagation nodes in the treatment segment. The propagation length and the number of changes in the connection direction information of the propagation nodes in the propagation chain sequence at consecutive sampling times are statistically analyzed, and the propagation chain sequence with a propagation length exceeding the propagation threshold and the number of changes in the connection direction information of the propagation nodes exceeding the direction change threshold are marked as abnormal propagation chains. The working condition evolution coding layer uses the working condition phase transition features and the anomaly propagation chain as the initial state of evolution, and gradually splices the processing segment temporal embedding features corresponding to the continuous sampling time according to the connection order of the propagation nodes to generate multiple working condition evolution branch sequences. Evolution direction encoding is performed on the evolution branch sequence of each working condition, the embedding feature difference between adjacent sampling times is calculated, and the embedding feature difference is accumulated in the order of sampling time to generate the working condition evolution trajectory features; The multi-scale folded coding layer performs hierarchical sampling processing on the working condition evolution trajectory features according to short time scale, medium time scale and long time scale respectively. The sampled features at different time scales are mapped to the same embedding dimension and spliced ​​in the order of time scale to generate multi-scale working condition evolution features.

[0022] In this implementation, the improved TS2Vec model adopts a four-layer temporal coding structure, with each layer having a coding dimension of 128. The processing segment topology embedding layer establishes a topological adjacency sequence according to the connection order of the processing segments. The initial topological edge weight between adjacent processing segments is set to 1.0, and the initial topological edge weight between non-adjacent processing segments is set to 0.2. The operating condition phase change detection layer uses a sliding time window with a length of 20 sampling points to perform gradient statistical processing. The gradient threshold is set to twice the average difference of the embedded features at adjacent sampling times, and the phase change threshold is set to the number of sampling points that continuously exceed the gradient threshold reaching 8. The operating condition propagation chain generation layer establishes propagation node connection relationships according to the consistency rate of pollutant concentration change direction, the consistency rate of pressure difference change direction, and the consistency rate of energy consumption change direction. A propagation connection is established when the consistency rate of change direction exceeds 0.7. The propagation length threshold is set to 4, and the direction change threshold is set to 3. The multi-scale folding coding layer performs hierarchical sampling processing using time scales of 30s, 300s, and 1800s, and then uniformly maps them to a 256-dimensional embedding space before performing temporal scale sequential splicing. Both the improved TS2Vec model and the TS2Vec model adopt a time series self-supervised representation learning structure, perform time series encoding processing on multivariate running data in continuous time windows, and use a shared time series encoding network to extract local and global change features in the time series. The improved TS2Vec model retains the time series embedding mechanism and multi-layer encoding structure in the TS2Vec model, while also retaining the time window sliding sampling method and the unified embedding space mapping method. The improved TS2Vec model adds a treatment segment topology embedding layer, a condition phase change detection layer, a condition propagation chain generation layer, a condition evolution coding layer, and a multi-scale folding coding layer to the TS2Vec model. The treatment segment topology embedding layer performs topology embedding processing on the connection relationship between each treatment segment of the waste gas treatment equipment. The condition phase change detection layer performs gradient aggregation processing on the embedding feature difference between adjacent sampling times. The condition propagation chain generation layer establishes a propagation chain sequence based on the direction of pollutant concentration change, pressure difference change, and energy consumption change. The condition evolution coding layer performs evolutionary branch coding processing on the propagation chain sequence. The multi-scale folding coding layer performs hierarchical folding processing on the condition evolution trajectory features at different time scales. The improved TS2Vec model incorporates the correlation between treatment sections, the phase change process of operating conditions, and the pollution propagation process during the operation of the waste gas treatment equipment into the time series representation process, reducing the dependence of traditional time series models on single time-varying features. The operating condition propagation chain generation structure improves the ability to identify propagation paths between different treatment sections, and the multi-scale folding coding structure improves the ability to express the correlation between long-cycle operating conditions and short-cycle fluctuating operating conditions, thereby enhancing the ability to represent the evolutionary features under complex waste gas treatment conditions.

[0023] In this embodiment, the contribution calculation module includes: Extract pollutant concentration evolution characteristics, pressure difference evolution characteristics, energy consumption evolution characteristics, and propagation chain correlation characteristics corresponding to each treatment segment from the multi-scale operating condition evolution characteristics; The pollutant reduction amount of each treatment section is calculated based on the difference between the inlet and outlet pollutant concentrations of each treatment section and the corresponding exhaust gas flow rate of the treatment section. The operating energy consumption of each treatment section is calculated based on the fan power, spray pump power, and heating power of each treatment section. The basic contribution coefficient is generated by dividing the pollutant reduction amount by the sum of the operating energy consumption and the positive energy consumption smoothing parameter. Adjust the weight values ​​corresponding to the basic contribution coefficients based on the propagation chain association characteristics to generate the processing segment contribution coefficients for each processing segment. Calculate the difference in the processing segment contribution coefficient between the current control cycle and the previous control cycle, and generate the change in processing segment contribution. The changes in the contribution of the treatment section, the evolution characteristics of pollutant concentration, the evolution characteristics of pressure difference, and the evolution characteristics of energy consumption are fused and calculated to generate reinforcement learning reward data.

[0024] In this implementation, the pollutant reduction is calculated by multiplying the difference between the pollutant concentration at the inlet and outlet of the treatment section by the corresponding exhaust gas flow rate. The fan power, spray pump power, and heating power are accumulated according to the average power value within the same control cycle to generate the operating energy consumption. The positive energy consumption smoothing parameter is set to 0.01 to reduce the sudden change in the contribution coefficient under low energy consumption conditions. The propagation chain association feature is jointly calculated by the propagation length corresponding to the propagation chain sequence under the operating condition and the number of changes in the propagation node connection direction information. The weight of the propagation length is set to 0.6, and the weight of the number of changes in the propagation node connection direction information is set to 0.4. The treatment section contribution coefficient is corrected by multiplying the basic contribution coefficient by the corresponding weight value of the propagation chain association feature. The change in treatment section contribution is calculated by subtracting the treatment section contribution coefficient of the previous control cycle from the current control cycle. The reinforcement learning reward data is calculated by performing a weighted fusion calculation according to the weight values ​​corresponding to the change in treatment section contribution, pollutant concentration evolution feature, pressure difference evolution feature, and energy consumption evolution feature. The weight of the pollutant concentration evolution feature is set to 0.5, the weight of the pressure difference evolution feature is set to 0.2, and the weight of the energy consumption evolution feature is set to 0.3.

[0025] In this embodiment, the reinforcement learning decision-making module includes: The multi-scale operating condition evolution features are arranged according to the execution order of the processing segment number to generate the operating condition state input sequence; The reinforcement learning reward data is used to establish a reward time sequence according to the execution time of the control cycle, and then concatenated with the working condition input sequence at the corresponding position to generate the propagation chain policy state vector. Input the propagation chain policy state vector into the reinforcement learning decision network; The reinforcement learning decision network includes a working condition propagation chain policy encoding layer, a processing segment intervention sequence generation layer, an action primitive combination layer, and a contribution feedback evaluation layer; The working condition propagation chain strategy encoding layer performs propagation chain state encoding processing on the propagation chain strategy state vector to generate a processing segment propagation state sequence. The intervention sequence generation layer performs segmented scoring processing on the propagation state sequence of the processing segments according to the connection order of the processing segments in the anomaly propagation chain, and generates the intervention sequence of the processing segments. The action primitive combination layer selects action primitives from the action primitive sets corresponding to fan frequency adjustment action, valve opening adjustment action, spray flow rate adjustment action, and heating power adjustment action according to the intervention order of the processing segment, and performs action splicing processing according to the intervention order of the processing segment to generate multiple candidate action sequences. The contribution feedback evaluation layer calculates the action score based on the change in the contribution of the processing segment, the change in operating energy consumption, and the change in the evolution of pollutant concentration corresponding to each candidate action sequence, and sorts each candidate action sequence according to the size of the action score. The candidate action sequence with the highest action score is selected as the device control action sequence.

[0026] In this implementation, the propagation chain strategy state vector is represented by a 256-dimensional feature vector, and the control period is set to 30 seconds. The working condition propagation chain strategy encoding layer adopts a three-layer temporal encoding structure, with each layer having a 128-dimensional encoding dimension. The processing segment intervention sequence generation layer performs segmented scoring processing based on the propagation length, the number of changes in the propagation node connection direction information, and the change in the processing segment contribution in the abnormal propagation chain. The weight of propagation length is set to 0.4, the weight of the number of changes in the propagation node connection direction information is set to 0.3, and the weight of the change in the processing segment contribution is set to 0.3. The action primitive set includes increasing the fan frequency by 5 Hz, decreasing the fan frequency by 5 Hz, and increasing the valve opening by 10 Hz. The action sequence can be generated by combining up to four action primitives according to the intervention order of the treatment section to generate candidate action sequences. The contribution feedback evaluation layer performs weighted scoring processing based on the change in the contribution of the treatment section, the change in operating energy consumption, and the change in pollutant concentration. The weight of the change in the contribution of the treatment section is set to 0.5, the weight of the change in operating energy consumption is set to 0.2, and the weight of the change in pollutant concentration is set to 0.3. Candidate action sequences with an action score value lower than 0.35 are deleted from the ranking results.

[0027] In this embodiment, the device control module includes: The primitives of each action in the equipment control action sequence are parsed and processed according to the intervention order of the processing segment to generate the corresponding control parameter adjustment sequence; Based on the control parameter adjustment sequence, extract the fan frequency adjustment parameters, valve opening adjustment parameters, spray flow rate adjustment parameters, and heating power adjustment parameters respectively; The control parameter adjustment sequence for each treatment section is arranged according to the direction of waste gas flow, and a treatment section action execution queue is established. According to the action execution queue of the processing section, the fan operating frequency, valve opening degree, spray flow rate and heating power in the corresponding processing section are adjusted sequentially; After the control parameters of the current treatment section are adjusted, collect data on pollutant concentration changes, pressure difference changes, and energy consumption changes for the corresponding treatment section. Update the control parameter adjustment sequence for the remaining treatment sections based on the pollutant concentration change data, pressure difference change data, and energy consumption change data of the current treatment section; The control actions corresponding to the remaining processing segments will continue to be executed according to the updated control parameter adjustment sequence until all control actions corresponding to the equipment control action sequence are completed.

[0028] In this implementation, the equipment control action sequence adopts a segmented progressive execution method. Within the same control cycle, the control actions corresponding to each treatment segment are executed sequentially according to the direction of waste gas flow. Each action primitive in the control parameter adjustment sequence corresponds to an independent control register address. The fan operating frequency, valve opening, spray flow rate, and heating power are written as execution parameters through analog output interface and industrial communication bus, respectively. The treatment segment action execution queue is dynamically adjusted according to the connection order of propagation nodes in the treatment segment association topology. When the pollutant concentration change data corresponding to the current treatment segment exceeds the concentration fluctuation threshold, the execution priority of the corresponding treatment segment in the action execution queue is increased. When the differential pressure change data exceeds the differential pressure fluctuation threshold, the fan frequency adjustment amplitude of the corresponding treatment segment is reduced. When the energy consumption change data exceeds the energy consumption fluctuation threshold, the heating power adjustment amplitude is limited. The control parameter adjustment sequence update cycle is set to 30s, the treatment segment action execution interval is set to 5s, and a stable operating time interval is reserved between two adjacent control parameter adjustments.

[0029] In this embodiment, the operation feedback module includes: Collect pollutant concentration data, differential pressure data, temperature data, humidity data, fan frequency data, valve opening data, spray flow rate data, and heating power data after the execution of the equipment control action sequence; The controlled running data is time-aligned according to the acquisition time sequence and written into the current multivariate running time window sequence according to the acquisition time sequence. Delete historical sampling data that exceed the sampling window length threshold in the multivariate runtime window sequence, and generate an updated multivariate runtime window sequence; Calculate the changes in pollutant concentration, pressure difference, and energy consumption between the corresponding sampling times before and after control, and generate an operating state offset sequence; Update the connection relationships and connection directions of propagation nodes in the operating condition propagation chain sequence according to the operating status offset sequence; The updated operating condition propagation chain sequence is subjected to time-scale hierarchical processing, and the updated multivariate running time window sequence is combined with the updated operating condition evolution trajectory features to generate the features iteratively again. The corresponding multi-scale operating condition evolution features are updated based on the regenerated operating condition evolution trajectory features.

[0030] In this implementation, the sampling period for the controlled operational data is set to 5 seconds, and the threshold for the sampling window length corresponding to the multivariate operational time window sequence is set to 120 sampling points. When the number of sampling points in the multivariate operational time window sequence exceeds the sampling window length threshold, the earliest sampling data is deleted. The operational state offset sequence is generated by combining the changes in pollutant concentration, pressure difference, and energy consumption between the corresponding sampling times before and after control. When the change in pollutant concentration exceeds 15 ppm, the connection weight of the corresponding propagation node is increased; when the change in pressure difference exceeds 120 Pa, the connection strength of the corresponding propagation node is decreased; and when the change in energy consumption exceeds 8 kW, the connection direction information of the corresponding propagation node is adjusted. The operational condition propagation chain sequence is processed hierarchically using three time scales: 30 seconds, 300 seconds, and 1800 seconds. When re-iterating and generating the operational condition evolution trajectory features, a unified dimension mapping process is performed on the trajectory features corresponding to the same propagation node at different time scales, and the mapping dimension is set to 256 dimensions. The update period for the multi-scale operational condition evolution features is set to 30 seconds.

[0031] Example 1: To verify the feasibility of this invention in practice, it was applied to the volatile organic compound (VOC) waste gas treatment system of a new energy material production enterprise. The waste gas treatment system includes a pretreatment section, a spray treatment section, an activated carbon adsorption section, and an RTO thermal oxidation section. The waste gas sources include coating processes, solvent recovery processes, and high-temperature reaction processes. The waste gas components mainly include toluene, xylene, ethyl acetate, and a small amount of particulate matter. Due to the large fluctuations in waste gas concentration during different production periods, traditional control methods based on fixed thresholds cannot accurately reflect the propagation relationship between the treatment sections. When the waste gas concentration rises rapidly in a short period, problems such as lag in the spray section, pressure fluctuations in the adsorption section, and abnormal increases in the RTO furnace temperature easily occur, leading to increased system energy consumption. Furthermore, the outlet pollutant concentration approaches the emission limit during certain time periods.

[0032] In actual operation, this invention first collects real-time data on flow rate, temperature, humidity, pressure difference, fan frequency, spray flow rate, and pollutant concentration in the pretreatment section, spray treatment section, activated carbon adsorption section, and RTO thermal oxidation treatment section, and establishes a multivariate operating time window sequence according to a uniform time interval. Then, it uses the treatment section topology embedding layer in the improved TS2Vec model to establish the correlation topology between different treatment sections, and simultaneously identifies concentration abrupt change regions during operation through the operating condition phase change detection layer. When the volatile organic compound concentration at the spray section outlet increases rapidly, the operating condition propagation chain generation layer... The consistency of pollutant concentration change direction and pressure difference change direction constructs the operating condition propagation chain sequence and identifies abnormal propagation chains. The operating condition evolution coding layer generates multiple operating condition evolution branch sequences based on abnormal propagation chains, predicting the operational change trends of different treatment sections under different adjustment actions. The reinforcement learning decision network combines the treatment section contribution coefficient, propagation chain state, and changes in operating energy consumption to generate equipment control action sequences, dynamically adjusting the fan frequency, spray flow rate, and RTO heating power. The operation feedback module continuously updates the operating condition propagation chain sequence and multi-scale operating condition evolution characteristics, realizing dynamic iterative control under complex operating conditions.

[0033] During a 30-day continuous industrial operation test, this invention was compared with traditional PID control and conventional reinforcement learning control. The average concentration of exhaust gas at the inlet was maintained within the range of 850 ppm to 2100 ppm during the test, and high-concentration fluctuation conditions and sudden load increases in the spray section were randomly simulated. Test results show that this invention can identify pollution propagation trends in advance and complete coordinated adjustment of treatment sections before a significant increase in pollutant concentration occurs. During high-concentration fluctuations, the pressure difference fluctuation amplitude in the spray section controlled by this invention is significantly reduced, the outlet concentration change in the activated carbon adsorption section is more stable, and the temperature adjustment frequency in the RTO thermal oxidation section is reduced. Compared with traditional PID control, this invention shows significant improvements in average outlet pollutant concentration, unit energy consumption, and system stability. Compared with conventional reinforcement learning control, this invention has higher stability in terms of complex operating condition propagation identification and multi-treatment section coordinated control capabilities. The experimental results are shown in Table 1. Table 1 Comparison of Operational Performance of Waste Gas Treatment Systems

[0034] As can be seen from the comparison results in Table 1 above, there are significant differences in the stability of waste gas treatment, pollutant control capability, and system energy consumption among different control methods. Traditional PID control mainly adjusts parameters based on fixed thresholds. When the inlet concentration of waste gas fluctuates rapidly, the control process exhibits significant lag, resulting in an average outlet VOC concentration of 78 ppm and a peak VOC concentration of 132 ppm. Simultaneously, the fan adjustment frequency reaches 26 times / hour, indicating that frequent adjustments to operating parameters easily cause pressure fluctuations and equipment instability. While ordinary reinforcement learning control can dynamically adjust control parameters according to the operating status, it lacks the ability to model the propagation process of operating conditions and the correlation between treatment sections. Under complex propagation conditions, it still exhibits certain control deviations, resulting in pressure fluctuations reaching 176 Pa and exceeding limits 3 times / month. This invention dynamically models the pollution propagation and operational evolution processes by constructing a topological structure associated with treatment sections, a sequence of operational propagation chains, and multi-scale operational evolution characteristics. This results in a reduction of the average VOC concentration at the outlet to 38 ppm, a peak VOC concentration to 66 ppm, and an increase in system stability to 98.7%. Simultaneously, the average unit energy consumption under this invention's control is reduced to 351 kWh, and the number of fan adjustments is reduced to 11 times / h. This demonstrates that the invention can reduce ineffective control actions and improve the coordinated adjustment capability between treatment sections under complex operating conditions.

[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A system for optimizing the operation of waste gas treatment equipment based on reinforcement learning, characterized in that, include: The data acquisition module is used to collect the operating data of each treatment section of the waste gas treatment equipment and generate the corresponding treatment section status sequence. The time series construction module is used to splice the state sequences of processing segments at multiple consecutive time points to generate a multivariable runtime window sequence and construct the processing segment association topology. The operating condition characterization module is used to input the multivariate running time window sequence and the processing segment associated topology into the improved TS2Vec model, perform operating condition phase transition detection on the multivariate running time window sequence, generate operating condition phase transition features, and generate multi-scale operating condition evolution features based on the operating condition phase transition features. The contribution calculation module is used to calculate the contribution coefficient of the processing segment based on the multi-scale working condition evolution characteristics, and generate reinforcement learning reward data based on the contribution coefficient of the processing segment. The reinforcement learning decision module is used to input the multi-scale operating condition evolution characteristics and reinforcement learning reward data into the reinforcement learning decision network to generate a sequence of equipment control actions. The equipment control module is used to control the operation of the waste gas treatment equipment according to the equipment control action sequence; The operation feedback module is used to update the multivariable operation time window sequence and multi-scale operating condition evolution characteristics based on the controlled operation data.

2. The waste gas treatment equipment operation optimization system based on reinforcement learning according to claim 1, characterized in that, The data acquisition module includes: Waste gas flow rate data and spray flow rate data are collected at the locations of the intake pipe, spray pipe and discharge pipe along the direction of waste gas transportation, and flow time series data are generated according to the collection time. Volatile organic compound concentration data, particulate matter concentration data, and gas component concentration data were collected at the inlet and outlet of each treatment section, and concentration time series data were established according to the treatment section number. Temperature, humidity, differential pressure, fan frequency, valve opening, and heating power data are collected at the locations of the fan, reactor, spray tower, and adsorption device, and corresponding state time series data are established. The flow rate time series data, the concentration time series data, and the state time series data are aligned on the time axis. A unified sampling sequence is established according to a unified sampling time interval, and linear interpolation is performed on positions where the time interval exceeds the sampling interval threshold. The parameters of the unified sampling sequence are combined according to the processing section number to establish a processing section status sequence that includes flow rate parameters, concentration parameters, temperature and humidity parameters, differential pressure parameters, and control parameters. Fluctuation amplitude detection is performed on the state sequence of the processing segment, the absolute value of the data difference between adjacent sampling times is calculated, and data segments whose absolute value of the data difference exceeds the fluctuation threshold are marked as abnormal segments; The number of consecutive abnormal sampling points is counted, and data segments with a number of consecutive abnormal sampling points exceeding the abnormality duration threshold are deleted, while the corresponding time position markers are retained.

3. The waste gas treatment equipment operation optimization system based on reinforcement learning according to claim 1, characterized in that, The timing construction module includes: The processing segment state sequences are sorted according to the acquisition time to form a single-segment time series matrix corresponding to each processing segment; Using the current sampling time as the end point of the window, select single-segment time series matrices of multiple consecutive sampling times forward, and splice them together to form the local time window matrix corresponding to each processing segment; The local time window matrices corresponding to each treatment section are spliced ​​together according to the exhaust gas flow direction and the connection order of the treatment sections to generate a multivariate running time window sequence; Calculate the consistency rate of pollutant concentration change direction, pressure difference change direction, and energy consumption change direction for any two treatment segments within the same time window. The consistency rates of the pollutant concentration change direction, the pressure difference change direction, and the energy consumption change direction are weighted and fused to calculate the correlation strength of the processing segment; Each processing segment is used as a topology node, and the association strength of the processing segment is used as the weight of the topology edge to construct the processing segment association topology structure.

4. The waste gas treatment equipment operation optimization system based on reinforcement learning according to claim 1, characterized in that, The operating condition characterization module includes: The improved TS2Vec model includes a processing segment topology embedding layer, a shared temporal coding network, a load condition phase transition detection layer, a load condition propagation chain generation layer, a load condition evolution coding layer, and a multi-scale folding coding layer. The processing segment topology embedding layer performs node embedding processing on each processing segment node in the processing segment associated topology structure, and establishes a topology adjacency sequence according to the processing segment connection relationship; The multivariate runtime window sequence and the topological adjacency sequence are input into a shared temporal coding network, and linear mapping processing is performed on the time window features corresponding to different processing segments to generate temporal embedding features of the processing segments. The operating condition phase change detection layer performs differential calculation on the temporal embedding features of the processing segment between adjacent sampling times to generate a gradient sequence of operating condition changes; Sliding aggregation is performed on the gradient sequence of the operating condition change according to the sampling time order, the number of gradient changes exceeding the gradient threshold within the local time window is counted, and the sampling points that continuously exceed the gradient threshold are formed into phase transition candidate regions. Phase transition candidate regions where the number of gradient changes exceeds the phase transition threshold are marked as operating condition phase transition regions, and the processing segment temporal embedding features corresponding to the operating condition phase transition regions are extracted as operating condition phase transition features. The operating condition propagation chain generation layer performs correlation matching processing on the direction of pollutant concentration change, pressure difference change, and energy consumption change in each processing segment during continuous sampling time. It establishes the connection relationship of the propagation nodes of the processing segments according to the consistency of the change direction, and constructs the operating condition propagation chain sequence based on the connection relationship of the propagation nodes of the processing segments. The propagation length and the number of changes in the connection direction information of the propagation nodes in the working condition propagation chain sequence are statistically analyzed at continuous sampling time, and the working condition propagation chain sequence whose propagation length exceeds the propagation threshold and whose number of changes in the connection direction information of the propagation nodes exceeds the direction change threshold is marked as an abnormal propagation chain. The operating condition evolution coding layer uses the operating condition phase transition features and the anomaly propagation chain as the initial state of evolution, and gradually splices the processing segment temporal embedding features corresponding to the continuous sampling time according to the connection order of the propagation nodes to generate multiple operating condition evolution branch sequences. Evolution direction encoding is performed on each working condition evolution branch sequence, the embedding feature difference between adjacent sampling times is calculated, and the embedding feature difference is accumulated in the order of sampling time to generate working condition evolution trajectory features; The multi-scale folded coding layer performs hierarchical sampling processing on the working condition evolution trajectory features according to short time scale, medium time scale and long time scale respectively, maps the sampled features under different time scales to the same embedding dimension, and splices them in the order of time scale to generate multi-scale working condition evolution features.

5. The waste gas treatment equipment operation optimization system based on reinforcement learning according to claim 1, characterized in that, The contribution calculation module includes: Extract the pollutant concentration evolution characteristics, pressure difference evolution characteristics, energy consumption evolution characteristics, and propagation chain correlation characteristics corresponding to each treatment segment from the multi-scale operating condition evolution characteristics; The pollutant reduction amount of each treatment section is calculated based on the difference between the inlet and outlet pollutant concentrations of each treatment section and the corresponding exhaust gas flow rate of the treatment section. The operating energy consumption of each treatment section is calculated based on the fan power, spray pump power, and heating power of each treatment section. The basic contribution coefficient is generated by dividing the pollutant reduction amount by the sum of the operating energy consumption and the positive energy consumption smoothing parameter. Adjust the weight values ​​corresponding to the basic contribution coefficients according to the propagation chain association features, and generate the processing segment contribution coefficients corresponding to each processing segment. Calculate the difference in the processing segment contribution coefficient between the current control cycle and the previous control cycle, and generate the change in processing segment contribution. The changes in the contribution of the processing section, the evolution characteristics of the pollutant concentration, the evolution characteristics of the pressure difference, and the evolution characteristics of the energy consumption are fused and calculated to generate reinforcement learning reward data.

6. The waste gas treatment equipment operation optimization system based on reinforcement learning according to claim 1, characterized in that, The reinforcement learning decision module includes: The multi-scale working condition evolution features are arranged according to the execution order of the processing segment number to generate a working condition state input sequence; The reinforcement learning reward data is used to establish a reward time sequence according to the execution time of the control cycle, and then concatenated with the working condition input sequence at the corresponding positions to generate a propagation chain policy state vector. The propagation chain strategy state vector is input into the reinforcement learning decision network; The reinforcement learning decision network includes a working condition propagation chain policy encoding layer, a processing segment intervention sequence generation layer, an action primitive combination layer, and a contribution feedback evaluation layer. The working condition propagation chain strategy encoding layer performs propagation chain state encoding processing on the propagation chain strategy state vector to generate a processing segment propagation state sequence. The processing segment intervention order generation layer performs segmented scoring processing on the processing segment propagation state sequence according to the connection order of processing segments in the abnormal propagation chain, and generates the processing segment intervention order. The action primitive combination layer selects action primitives from the action primitive sets corresponding to fan frequency adjustment action, valve opening adjustment action, spray flow rate adjustment action, and heating power adjustment action according to the intervention order of the processing segment, and performs action splicing processing according to the intervention order of the processing segment to generate multiple candidate action sequences. The contribution feedback evaluation layer calculates the action score value based on the change in the contribution of the processing segment, the change in operating energy consumption, and the change in the evolution of pollutant concentration corresponding to each candidate action sequence, and sorts each candidate action sequence according to the size of the action score value. The candidate action sequence with the highest action score is selected as the device control action sequence.

7. The waste gas treatment equipment operation optimization system based on reinforcement learning according to claim 1, characterized in that, The device control module includes: The primitives of each action in the device control action sequence are parsed and processed according to the intervention order of the processing segment to generate the corresponding control parameter adjustment sequence; Based on the control parameter adjustment sequence, extract the fan frequency adjustment parameter, valve opening adjustment parameter, spray flow rate adjustment parameter, and heating power adjustment parameter respectively; The control parameter adjustment sequence for each treatment section is arranged according to the direction of waste gas flow, and a treatment section action execution queue is established. According to the action execution queue of the processing segment, the fan operating frequency, valve opening, spray flow rate and heating power in the corresponding processing segment are adjusted sequentially. After the control parameters of the current treatment section are adjusted, collect data on pollutant concentration changes, pressure difference changes, and energy consumption changes in the corresponding treatment section. Update the control parameter adjustment sequence for the remaining treatment sections based on the pollutant concentration change data, pressure difference change data, and energy consumption change data of the current treatment section; Continue executing the control actions corresponding to the remaining processing segments according to the updated control parameter adjustment sequence until all control actions corresponding to the device control action sequence are completed.

8. The waste gas treatment equipment operation optimization system based on reinforcement learning according to claim 1, characterized in that, The operation feedback module includes: Collect pollutant concentration data, differential pressure data, temperature data, humidity data, fan frequency data, valve opening data, spray flow rate data, and heating power data after executing the equipment control action sequence; The controlled running data is time-aligned according to the acquisition time sequence and written into the current multivariate running time window sequence according to the acquisition time sequence. Delete historical sampling data that exceed the sampling window length threshold in the multivariate runtime window sequence, and generate an updated multivariate runtime window sequence; Calculate the changes in pollutant concentration, pressure difference, and energy consumption between the corresponding sampling times before and after control, and generate an operating state offset sequence; Update the propagation node connection relationship and propagation node connection direction information in the operating condition propagation chain sequence according to the operating state offset sequence; The updated operating condition propagation chain sequence is subjected to time-scale hierarchical processing, and the updated multivariate running time window sequence is combined with the updated operating condition evolution trajectory features to generate the features iteratively again. The corresponding multi-scale operating condition evolution features are updated based on the regenerated operating condition evolution trajectory features.