Carbon dioxide concentration control method and device, terminal and medium

By using reinforcement learning modules and preset optimal equations, combined with action value functions, the problem of signal drift of traditional sensors in dynamic environments is solved, enabling precise control of carbon dioxide concentration and improving the system's adaptability and accuracy.

CN121879446AInactive Publication Date: 2026-04-17SOUTHWEST PETROLEUM UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST PETROLEUM UNIV
Filing Date
2025-12-30
Publication Date
2026-04-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional sensors are prone to signal drift and accuracy degradation in dynamic environments, resulting in a significant reduction in the accuracy of carbon dioxide concentration measurement. Existing static compensation models cannot adapt to complex and ever-changing real-time environmental conditions, limiting the robustness and practicality of the system.

Method used

By employing a reinforcement learning module combined with a preset optimal equation and action value function, the system obtains the current carbon dioxide concentration and environmental parameters, calculates the optimal state function value, selects and executes target actions, adjusts the correction coefficient to control the carbon dioxide concentration, and utilizes a preset hardware interface to achieve precise calibration.

Benefits of technology

It enables precise and rapid control of carbon dioxide concentration in dynamic environments, reduces measurement errors, and improves the robustness and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121879446A_ABST
    Figure CN121879446A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon dioxide concentration control method and device, a terminal and a medium. The method comprises the steps that the current carbon dioxide concentration corresponding to a target environment and environment parameters corresponding to the target environment are obtained; performing data processing on the current carbon dioxide concentration and the environmental parameters to obtain a state vector corresponding to a target environment; calculating an optimal state function value corresponding to the state vector based on a preset optimal equation calculation formula and the state vector; based on the state vector, the scene complexity corresponding to the target environment and the optimal state function value, determining an action value function corresponding to the state vector; and based on a preset screening strategy, screening out a target action from the action space corresponding to the action value function, and executing the target action to control the carbon dioxide concentration corresponding to the target environment. The invention aims to realize accurate and rapid control of the carbon dioxide concentration of the target environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, terminal and medium for controlling carbon dioxide concentration. Background Technology

[0002] Accurate real-time monitoring of carbon dioxide (CO2) plays a crucial role in fields such as indoor air quality control, industrial safety early warning, agricultural greenhouse regulation, and environmental science research.

[0003] However, traditional sensors are prone to signal drift and accuracy degradation in dynamic environments (such as drastic fluctuations in temperature and humidity, interference from pollutants, or long-term continuous operation), leading to a significant reduction in measurement accuracy. For example, in dynamic environments, non-dispersive infrared (NDIR) sensors can generate errors as high as 43% due to fluctuations in temperature, humidity, and air pressure. To address these challenges, existing methods typically rely on static compensation models based on fixed parameters. While these models can reduce some fundamental errors, they cannot adapt to complex and changing real-time environmental conditions. Their preset parameters have limited generalization ability in uncalibrated states and cannot achieve adaptive optimization during long-term operation, which limits the robustness and practicality of the system. Linear or polynomial correction methods ignore nonlinear interaction effects such as humidity-air pressure, which can lead to errors exceeding 10% in nonlinear cases. Summary of the Invention

[0004] The main objective of this application is to provide a method, device, terminal, and medium for controlling carbon dioxide concentration, which can achieve precise and rapid control of carbon dioxide concentration in a target environment.

[0005] To achieve the above objectives, this application provides a method for controlling carbon dioxide concentration, the method comprising: Obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment; The current carbon dioxide concentration and the environmental parameters are processed to obtain the state vector corresponding to the target environment. Based on the preset optimal equation calculation formula and the state vector, the optimal state function value corresponding to the state vector is calculated, wherein the optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment. Based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, the action value function corresponding to the state vector is determined, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment. Based on a preset screening strategy, target actions are selected from the action space corresponding to the action value function, and the target actions are executed to control the carbon dioxide concentration corresponding to the target environment.

[0006] Specifically, determining the action value function corresponding to the state vector based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value includes: Based on the aforementioned scenario complexity, the processing method of the preset reinforcement learning module for the state vector is determined; Based on the optimal state function value, the state vector is processed by the preset reinforcement learning module to obtain the action value function.

[0007] Specifically, determining the processing method of the preset reinforcement learning module for the state vector based on the scene complexity includes: If the scene complexity is less than the preset complexity threshold, then the state vector is processed by calling the preset function of the preset reinforcement learning module; If the scene complexity is greater than or equal to a preset complexity threshold, then the state vector is determined to be processed by a preset deep network of a preset reinforcement learning module.

[0008] Specifically, the step of filtering target actions from the action space corresponding to the action value function based on a preset filtering strategy includes: Determine the screening strategy value and the probability value respectively, wherein the screening strategy value is greater than 0 and less than 1, the probability value is greater than or equal to 0, and the probability value is less than 1; If the probability value is less than the filtering strategy value, then any action in the action space is determined as the target action; If the probability value is greater than or equal to the filtering strategy value, then the action corresponding to the maximum action value function value in the action space is determined as the target action.

[0009] Specifically, the target action includes adjusting the correction coefficient; The execution of the target action includes: Adjust the correction coefficients to obtain the adjusted correction coefficients; Substitute the adjusted correction coefficients into the preset calibration model to calculate the calibrated carbon dioxide concentration. The preset hardware interface controls the preset action hardware to adjust the current carbon dioxide concentration in the target environment to the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration.

[0010] Specifically, after performing the target action, the method further includes: Based on the target action, the preset carbon dioxide reference concentration, the current carbon dioxide concentration, the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration, and the calibrated carbon dioxide concentration, a reward function value is calculated using a preset reward function. Based on the reward function value and the state vector corresponding to the calibrated carbon dioxide concentration, the parameters of the preset reinforcement learning module are updated to obtain the updated preset reinforcement learning module.

[0011] Specifically, the target action includes a first type of action and a second type of action, and the preset reward function includes a first preset reward function and a second preset reward function; The step of calculating a reward function value based on the target action, a preset carbon dioxide reference concentration, the current carbon dioxide concentration, the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration, and the calibrated carbon dioxide concentration, through a preset reward function, includes: If the target action is the first type of action, then the preset carbon dioxide reference concentration and the calibrated carbon dioxide concentration are substituted into the first preset reward function to calculate the reward function value; If the target action is the second type of action, then the current carbon dioxide concentration and the target carbon dioxide concentration are substituted into the second preset reward function to calculate the reward function value.

[0012] To achieve the above objectives, this application also provides a carbon dioxide concentration control device, the device comprising: The first unit is used to obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment. The second unit is used to process the current concentration of carbon dioxide and the environmental parameters to obtain the state vector corresponding to the target environment. The third unit is used to calculate the optimal state function value corresponding to the state vector based on the preset optimal equation calculation formula and the state vector. The optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment. The fourth unit is used to determine the action value function corresponding to the state vector based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment. The fifth unit is used to select target actions from the action space corresponding to the action value function based on a preset screening strategy, and execute the target actions to control the carbon dioxide concentration corresponding to the target environment.

[0013] To achieve the above objectives, this application also provides a terminal, including a memory storing multiple instructions; the processor loads instructions from the memory to execute the steps in any of the methods provided in this application.

[0014] To achieve the above objectives, this application also provides a medium storing a plurality of instructions adapted for loading by a processor to execute the steps in any of the methods provided in this application.

[0015] This application provides a method, device, terminal, and medium for controlling carbon dioxide concentration. The method first acquires the current carbon dioxide concentration and environmental parameters corresponding to the target environment; processes the current carbon dioxide concentration and environmental parameters to obtain a state vector corresponding to the target environment; calculates the optimal state function value corresponding to the state vector based on a preset optimal equation and the state vector; determines the action value function corresponding to the state vector based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value; and selects a target action from the action space corresponding to the action value function based on a preset filtering strategy, and executes the target action to accurately and quickly control the carbon dioxide concentration corresponding to the target environment. Attached Figure Description

[0016] Figure 1 A flowchart illustrating the method provided in the embodiments of this application; Figure 2 A schematic diagram comparing the convergence of different reward formulas provided in the embodiments of this application; Figure 3 A schematic diagram of the Pareto front corresponding to the reward formula provided in the embodiments of this application; Figure 4 This is a schematic diagram of the sensitivity curve for α changes provided in the embodiments of this application; Figure 5 A schematic diagram illustrating the change of carbon dioxide concentration over time in an indoor test provided in this application embodiment; Figure 6 This is a schematic diagram of the device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the terminal structure provided in an embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] Traditional sensors are prone to signal drift and accuracy degradation in dynamic environments (such as those with drastic temperature and humidity fluctuations, pollutant interference, or long-term continuous operation), leading to a significant reduction in measurement accuracy. For example, in dynamic environments, non-dispersive infrared (NDIR) sensors can experience errors as high as 43% due to fluctuations in temperature, humidity, and air pressure. To address these challenges, existing methods typically rely on static compensation models based on fixed parameters. While these models can reduce some fundamental errors, they cannot adapt to complex and changing real-time environmental conditions. Their preset parameters have limited generalization ability in uncalibrated states and cannot achieve adaptive optimization during long-term operation, which limits the robustness and practicality of the system. Linear or polynomial correction methods neglect nonlinear interaction effects such as humidity-air pressure, which can lead to errors exceeding 10% in nonlinear cases.

[0019] Therefore, embodiments of this application provide a method, apparatus, terminal, and medium for controlling carbon dioxide concentration to solve practical technical problems.

[0020] In some embodiments, the device may be integrated into an electronic device, such as a terminal or server.

[0021] In some embodiments, the server may also be implemented as a terminal.

[0022] The server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0023] The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited herein.

[0024] The following sections provide detailed descriptions of each example. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0025] This application provides a method for controlling carbon dioxide concentration, which can achieve precise and rapid control of carbon dioxide concentration in a target environment.

[0026] In some embodiments, a 50㎡ office houses 10 employees who work daily. The CO2 concentration needs to be stabilized at 700-800ppm (target concentration Ctarget=750ppm), balancing monitoring accuracy with air conditioning and ventilation energy consumption. System configuration: CO2 sensor (MH-Z19B), environmental sensors (DHT22 temperature and humidity sensor + BMP280 barometric pressure sensor), processing unit (ARM Cortex-M4), ventilation actuator (1000W fan), and filtering strategy values. =0.1, preset complexity threshold is 0.6.

[0027] Furthermore, the first preset reward function is defined as being expressible by the following calculation expression:

[0028] in, This indicates the calibrated carbon dioxide concentration. This indicates the preset carbon dioxide reference concentration; A second preset reward function is also defined, which can be represented by the following calculation expression:

[0029] in, This indicates the current concentration of carbon dioxide. This indicates the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration. This indicates the energy consumption generated by controlling carbon dioxide emissions. This represents the nonlinear energy consumption index.

[0030] like Figure 1 The specific process of the method can be as follows: S110. Obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment.

[0031] In some embodiments, the current original concentration of CO2 in the office is collected by a CO2 sensor (MH-Z19B). =850ppm; Environmental parameters were collected synchronously via environmental sensors: temperature =850℃, humidity =55%, pressure =1012hPa; Simultaneously record the CO2 concentration change data for the past 10 minutes (780ppm→850ppm) for subsequent calculation of the concentration change rate.

[0032] S120. Perform data processing on the current carbon dioxide concentration and the environmental parameters to obtain the state vector corresponding to the target environment.

[0033] In some embodiments, a 5-point median window can be used. =850ppm and environmental parameters were subjected to median filtering (to filter sensor noise), and the processed data was kept at the specified value. =850ppm =850℃ =55%, =1012 hPa; In some embodiments, the rate of change of carbon dioxide concentration can be calculated using the following methods: = (850ppm-780ppm) / (10min×60s / min)=0.117ppm / s.

[0034] In some embodiments, the filtered data and the concentration change rate are integrated to obtain a state vector: .

[0035] S130. Based on the preset optimal equation calculation formula and the state vector, calculate the optimal state function value corresponding to the state vector, wherein the optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment.

[0036] In some embodiments, the preset optimal equation calculation formula can be the Bellman optimality equation, i.e.:

[0037] in, The optimal state function value represents the optimal state function in state . The expected cumulative reward that can be obtained by taking the optimal action. :state; :action. In state Next action Instant rewards. Discount factor (0≤γ<1), measures the importance of future rewards. : Next state. For expectation operators.

[0038] Specifically, the state vector Substituting [850, 28, 55, 1012, 0.117] into the equation and combining it with historical "state-action-reward" data, the optimal state function value is obtained. =92, which represents the upper limit of the expected cumulative reward that can be obtained by taking the optimal action under the current office environment conditions.

[0039] S140. Based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, determine the action value function corresponding to the state vector, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment.

[0040] In some embodiments, determining the action value function corresponding to the state vector based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value includes the following steps S141 to S142: S141. Based on the scene complexity, determine the processing method of the preset reinforcement learning module for the state vector.

[0041] In some embodiments, determining the processing method of the preset reinforcement learning module for the state vector based on the scene complexity includes the following specific implementation process: If the scene complexity is less than the preset complexity threshold, then the state vector is processed by calling the preset function of the preset reinforcement learning module; If the scene complexity is greater than or equal to a preset complexity threshold, then the state vector is determined to be processed by a preset deep network of a preset reinforcement learning module.

[0042] Specifically, the complexity of the office scene is calculated to be 0.4 (small fluctuations and gradual changes) using the system's preset complexity evaluation model (combining the fluctuation range of environmental parameters and the rate of change of CO2 concentration). Since 0.4 < preset complexity threshold 0.6, it is determined that the state vector will be processed by calling the preset Q function of the preset reinforcement learning module.

[0043] S142. Based on the optimal state function value, the state vector is processed by the preset reinforcement learning module to obtain the action value function.

[0044] The state vector =[850,28,55,1012,0.117] Input the preset Q function to obtain the optimal state function value. Using 92 as the value benchmark, the relationship between all possible actions and the expected cumulative reward in the current state is quantified to obtain the action value function:

[0045] in, The optimal representation of the action-value function Value, representing the state Next action Expected cumulative rewards that can be obtained Indicates the instant reward value. Indicates the next state Execute the next action The expected cumulative reward that can be obtained.

[0046] S150. Based on a preset screening strategy, select the target action from the action space corresponding to the action value function, and execute the target action to control the carbon dioxide concentration corresponding to the target environment.

[0047] In some embodiments, the step of filtering target actions from the action space corresponding to the action value function based on a preset filtering strategy includes the following steps S151 to S152: S151. Determine the screening strategy value and the probability value respectively, wherein the screening strategy value is greater than 0 and less than 1, the probability value is greater than or equal to 0, and the probability value is less than 1.

[0048] In some embodiments, the filtering strategy value =0.1 (initial value), and the system random generator generates a probability value of 0.08 in the interval [0,1).

[0049] S152. If the probability value is less than the filtering strategy value, then any action in the action space is determined as the target action; if the probability value is greater than or equal to the filtering strategy value, then the action corresponding to the maximum action value function value in the action space is determined as the target action.

[0050] In some embodiments, since the generated probability value 0.08 is less than the filtering strategy value 0.1, an action is randomly selected from the action space as the target action, and the target action is ultimately determined to be a first-class action. =0.01 (adjusting correction factor).

[0051] In some embodiments, the target action includes adjusting the correction coefficient.

[0052] Specifically, performing the target action includes the steps A1 to A3 as shown below: A1. Adjust the correction coefficients to obtain the adjusted correction coefficients.

[0053] In some embodiments, the adjustment correction factor is: the system initial drift correction factor. =1.0, based on the target action =0.01, update to obtain the adjusted correction coefficient. =0.01+1.0=1.01.

[0054] A2. Substitute the adjusted correction coefficients into the preset calibration model to calculate the calibrated carbon dioxide concentration.

[0055] In some embodiments, =0.01+1.0=1.01 Substitute into the preset calibration model: ,in The nonlinear compensation terms for temperature, humidity, and pressure are calculated. =-15ppm, obtained =1.01×850+( 15) = 843.5 ppm.

[0056] A3. Control the preset action hardware through the preset hardware interface to adjust the current carbon dioxide concentration of the target environment to the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration.

[0057] In some embodiments, a control signal can be sent to the CO2 sensor via a pulse width modulation (PWM, 25kHz) hardware interface to update the sensor's measurement calibration parameters. =0.01+1.0=1.01, which synchronizes the CO2 concentration output by the sensor to 843.5ppm, while recording the current ventilation rate as 25% (no ventilation adjustment triggered).

[0058] In some embodiments, after performing the target action, the method further includes steps B1 to B2 as shown below: B1. Based on the target action, the preset carbon dioxide reference concentration, the current carbon dioxide concentration, the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration, and the calibrated carbon dioxide concentration, the reward function value is calculated through the preset reward function.

[0059] In some embodiments, the target action includes a first type of action and a second type of action, and the preset reward function includes a first preset reward function and a second preset reward function.

[0060] Specifically, the step of calculating the reward function value based on the target action, the preset carbon dioxide reference concentration, the current carbon dioxide concentration, the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration, and the calibrated carbon dioxide concentration, through a preset reward function, includes the following specific implementation process: If the target action is the first type of action, then the preset carbon dioxide reference concentration and the calibrated carbon dioxide concentration are substituted into the first preset reward function to calculate the reward function value; if the target action is the second type of action, then the current carbon dioxide concentration and the target carbon dioxide concentration are substituted into the second preset reward function to calculate the reward function value.

[0061] Specifically, the target action is the first type of action (calibration coefficient adjustment), which involves setting the preset carbon dioxide reference concentration. =845ppm (measured by standard reference instrument Picarro G2201-i) and calibrated carbon dioxide concentration Substituting 843.5ppm into the first preset reward function, the reward function value is calculated. =-|843.5-845|=-1.5.

[0062] B2. Based on the reward function value and the state vector corresponding to the calibrated carbon dioxide concentration, update the parameters of the preset reinforcement learning module to obtain the updated preset reinforcement learning module.

[0063] In some embodiments, a new state vector can be acquired after the target action is performed: calibrated sensor output. =843.5ppm, new environmental parameters were collected simultaneously. =28℃ =55%, =1012 hPa, calculate the rate of change of the new concentration. =0.05ppm / s, resulting in a new state vector. =[843.5,28,55,1012,0.05].

[0064] In some embodiments, the result is obtained by calculating a preset Q function. =88, which is the maximum action value function value corresponding to the new state.

[0065] In some embodiments, the Q-learning-based update formula is used:

[0066] Substitute the parameters, =89、 =0.1, =-1.5、 =0.9、 = =88, the updated value is calculated as follows: =89+0.1*(-1.5+0.9*88-89)=87.87.

[0067] The updated preset reinforcement learning module is output for action decisions in the next round of CO2 concentration control.

[0068] In some embodiments, steps S110 to B2 are repeated every 5 minutes. If the CO2 concentration is calculated in the next round... If the concentration is 860ppm (exceeding the target concentration of 750ppm), the target action may trigger a second type of action (adjustment of ventilation volume). At this time, the reward value will be calculated through the second preset reward function, and the parameters of the reinforcement learning module will be continuously optimized to ultimately achieve a stable CO2 concentration in the office within the target range and the lowest energy consumption.

[0069] The method will be illustrated below through another specific embodiment: (1) System architecture: This embodiment presents an adaptive CO2 sensing system based on reinforcement learning (RL). The system achieves high-precision CO2 measurement and control through the deep integration of hardware, data processing, and reinforcement learning. Subsections describe the reinforcement learning module, dynamic calibration, and intelligent control technology. The design principles emphasize three core elements: modular architecture, network scalability, and real-time performance. The modular design supports multi-gas detection capabilities, network scalability supports large-scale deployment, and real-time performance stems from efficient data processing and decision-making loop mechanisms.

[0070] CO2 sensors include two models: NDIR (model MH-Z19B, range 0-5000ppm, accuracy ±1%FS, resolution 1ppm, response time <60 seconds, recovery time <120 seconds) and TDLAS (model LGR-ICOS, range 0-10000ppm, accuracy ±0.1%FS, resolution 0.1ppm, response time <1 second, recovery time <5 seconds). These two sensors were chosen because they have complementary advantages in terms of cost and accuracy. Their response time (<0.5 seconds) is optimized by reinforcement learning inference beyond hardware performance (for example, the baseline response time of MH-Z19B was originally less than 60 seconds, which was reduced to an average of 0.4 seconds after calculation by the built-in Q function). The environmental sensor can measure temperature (-40℃ to 80℃, accuracy ±0.5℃), humidity (0% to 100%, accuracy ±2%) and pressure (300 hPa to 1100 hPa, accuracy ±1 hPa) [8]. The system employs an ARM Cortex-M4 (168 MHz, 512 KB RAM) for low-power tasks and an NVIDIA Jetson Nano (4 GB RAM) for complex computations. The control interface uses pulse width modulation (PWM, 25 kHz) or the Modbus protocol. Raw data and parameters undergo median filtering (using a 5-point median window), and reinforcement learning (RL) outputs actions accordingly, forming a closed-loop control. The system supports network topology modeling via Modbus communication (node ​​model: latency = transmission delay + processing time, error rate less than 5%), thus achieving scalability. In the network security field, AES-256 encryption and OAuth 2.0 technology effectively prevent tampering (threat models: DDoS attacks on Wi-Fi, code injection via 4G networks). Anomaly detection technology can identify 95% of simulated attacks (false alarm rate of 5% in 50 industrial tests, which can be reduced to 2% through threshold adjustment).

[0071] (2) Reinforcement Learning Module: The reinforcement learning module adopts the Markov Decision Process (MDP) framework and solves for the optimal policy through the Bellman optimality equation.

[0072]

[0073] Bellman's optimal equation, meaning: the optimal state-value function, representing the state... The expected cumulative reward that can be obtained by taking the optimal action. :state The optimal value function. :state; :action. In state Next action Instant rewards. Discount factor (0≤ <1), to measure the importance of future rewards. : Next state.

[0074] The state space (st) is defined as a multidimensional vector:

[0075] This represents the original concentration of carbon dioxide. For temperature; Humidity; For pressure; The rate of change of carbon dioxide concentration (ppm / s).

[0076] in, The parameter range is 0-5000ppm, with an accuracy of ±10ppm; The parameter range is -10℃ to 50℃, with an accuracy of ±0.1℃. The parameter range is 0%-100%, with an accuracy of ±1%. The parameter range is 800hPa-1200hPa, with an accuracy of ±1hPa; (ppm / s) is used to capture dynamic trends. This vector integrates multivariate inputs for robust prediction, with variances of less than 1% in sensitivity analyses across different ecosystems.

[0077] The reward function is designed based on the task objective. The first formula below is used during calibration; this negative error reward mechanism effectively reduces measurement error by penalizing deviations from the reference value. The second formula below is used for control.

[0078] : The calibrated CO2 concentration. Reference CO2 concentration (true value).

[0079] Current CO2 concentration. Target CO2 concentration. E: Energy consumption. Energy consumption nonlinearity index (e.g., 1.5). Weighting coefficients (e.g., 0.7 and 0.3).

[0080] Among them, it is set as =0.7、 =0.3、 =1000、 =1.5. These parameters were determined using a grid search method, aiming to optimize the balance between computational accuracy and energy consumption (search range: ∈[0.5–0.9]、 ∈[0.1–0.5], step size = 0.1; convergence is considered when the Pareto front fluctuates by less than 5% in 50 iterations.

[0081] When the value of k takes different nonlinear parameters (e.g.) =1.0 and 2.0), in When the energy-accuracy tradeoff is 1.5, the error is lowest (3% lower than the multi-objective optimization scheme, n=100 experiments, paired t-test p<0.05). This is because the nonlinear penalty mechanism can more accurately reflect the increasing characteristics of energy cost. When adopting a hierarchical optimization strategy (i.e., prioritizing accuracy over energy in a two-level reward mechanism), such as... Figure 2 As shown, although the convergence speed is faster ( Figure 2 The proposed weighted scalar method performs 850 iterations compared to 950 iterations in multi-objective optimization, but its error is 2% higher in the balanced scenario (95% confidence interval [1.5%-2.5%], p<0.05). This verifies that the proposed weighted scalar method has advantages in embedded stability optimization, which is consistent with the research conclusions in the literature.

[0082] This multi-objective optimization scheme was validated through comparative experiments (100 runs, CO2 concentration 400-1200 ppm, temperature 15-35℃). Compared with the single-objective scheme, its trade-off performance was improved by 5% (average 5.2%, 95% confidence interval [4.1%-6.3%], p<0.05). In terms of balance, this scheme outperformed the hierarchical optimization scheme (error reduced by 2%, 95% confidence interval [1.5%-2.5%], p<0.05). The weight parameters were obtained through grid search ( =0.7, =0.3) can minimize the error to 1.8% (95% confidence interval [1.6%-2.0%]). Quantitative comparisons are performed using the Pareto frontier method. For example... Figure 3 As shown, in the accuracy-energy consumption dimension, the multi-objective frontier scheme outperforms the hierarchical scheme by 4% (e.g., with a 2% error, the energy saving rate is 28% compared to 24%, and the Wilcoxon test after 50 iterations shows p<0.05).

[0083] Based on value function learning analysis under piecewise linear control, these principles are implemented by embedding the value function in hardware for reward structure analysis. Unlike simple simulation, the Q-function of the integrated sensor hardware enables real-time piecewise adjustment. The infinite time-domain Q-function is shown below. It ensures long-term drift modeling, and its performance is superior to the finite-time model:

[0084] The optimal action-value function represents the state. Next action The expected cumulative reward that can be obtained. Optimal Q value. r: Immediate reward.

[0085] Regarding learning algorithms, the Q-learning update formula for simple scenarios is the first formula below. This algorithm approximates the infinite temporal Q-function by balancing immediate and future rewards. For complex scenarios, the Deep Q-Network (DQN) loss function can be used, as shown in the second formula below, which effectively handles high-dimensional inputs.

[0086] In the mean squared error loss function of Deep Q Network (DQN) : Parameters of the current network.

[0087] This is a closed-loop feedback mechanism where raw sensor data is directly fed back to the reinforcement learning state, actions adjust calibration parameters in real time, and the reward mechanism continuously improves accuracy and efficiency. This hardware-software collaboration is achieved by embedding the reinforcement learning algorithm into the processing unit (such as an ARM Cortex-M4), bridging the gap through direct data flow: sensor output, after preprocessing (filtering through a median window), serves as the state input; reinforcement learning calculates actions (such as...)... (Update) and executed via a hardware interface (such as sending a PWM signal to the actuator).

[0088] To make the infinite time domain assumption compatible with non-Markovian realities (such as potential drift in humidity hysteresis), the Bellman optimality equation can be extended using a partially observable model (POMDP boundary), thereby limiting the error to a specific range:

[0089] The POMDP error bound is the upper bound of the value function error in a partially observable Markov decision process. Value function error. Discount factor. This represents the maximum residual of the hidden state.

[0090] in The maximum residual of the hidden state is represented by ε=0.02 (data sourced from ARM Cortex-M4 testing, using Gaussian noise σ=10, for a total of 100 runs). Real-world hardware testing (100 noisy runs on an ARM Cortex-M4) shows that the algorithm converges within 1200 iterations (mean error <0.5%, 95% confidence interval [0.3%~0.7%]), while divergence in the finite-time domain occurs only 15%, verifying the robustness of the method for continuous calibration in partially observable environments. This conclusion is further validated by simulation experiments injecting non-Markovian noise (n=200), showing that the algorithm also converges within 1200 iterations (mean error <0.5%, 95% confidence interval [0.3%~0.7%]).

[0091] The derivation process of the control strategy and signals is as follows:

[0092] A deterministic control strategy selects the action that maximizes the Q value. The strategy (action selection) in state s. Action value function.

[0093] By combining the Bellman update algorithm with an ε-greedy exploration strategy (initial ε value 0.1, gradually decaying to 0.01), actions that maximize long-term benefits (such as ventilation regulation) are selected, thus balancing exploration and utilization in a dynamic environment and avoiding getting trapped in local optima due to noise and CO2 drift. Control signals (such as pulsed control signals PWM) are generated from the strategy output, enabling precise execution. Its core advantages are: the look-ahead characteristic of the Q function improves efficiency; deterministic derivation simplifies embedded hardware implementation; and the exploration mechanism handles uncertainty.

[0094] To address noise interference and model drift, the system employs a robust strategy for dynamic environments to ensure stability. By utilizing an empirical replay buffer with 10,000 samples and updating the target network every 100 steps, this approach reduces variance by 8% in Gaussian noise tests (σ=5). Linear interpolation effectively fills in missing data points. For model drift, the system employs a dual mechanism: setting a Q-value deviation threshold (triggers a reset if it exceeds 5% within 10 iterations) and using historical averages to construct redundant paths to replace missing inputs.

[0095] The optimization objective is to maximize the cumulative reward derived from the policy using the Bellman algorithm. A grid search method is employed to optimize parameters α and γ. The specific process is implemented in Python and run on an ARM Cortex-M4 embedded system, satisfying the embedded constraints. After a complete grid search involving 100 sets of experiments, with each set converging after approximately 1000 iterations, the optimal parameters α=0.1 and γ=0.9 were finally determined, minimizing the error (1.8%, 95% confidence interval [1.6%-2.0%]) and controlling the variance below 2%.

[0096] Sensitivity analysis: Increasing α by 10% (to 0.11) reduces the number of iterations by 15% (mean 850 iterations, 95% confidence interval [800, 900]), but increases the Q variance by 5% (mean 0.08, confidence interval [0.06, 0.10]). Decreasing γ to 0.09 stabilizes the variance (<2%, confidence interval [1.5%, 2.5%]), but prolongs the convergence time by 20% (mean 1200 iterations, confidence interval [1100, 1300]). Increasing γ by 10% (to 0.99) in mobile devices improves energy efficiency by 3% (mean 28.5%, confidence interval [27%, 30%]), but increases short-term error by 0.5% (confidence interval [0.3%, 0.7%]); decreasing it by 10% (to 0.81) prioritizes accuracy (error decreases by 0.8%, confidence interval [0.6%, 1.0%]), but increases energy consumption by 4% (confidence interval [3%, 5%]). Cross-ecosystem testing showed error variation of less than 1% (mean 0.7%, confidence interval [0.4%, 1.0%], Wilcoxon test p < 0.01, 100 runs). Regarding ecosystem robustness, the error change after migration was less than 1% (mean 0.7%, confidence interval [0.4%, 1.0%], p < 0.05, t-test for 100 runs) compared to the 5% error without migration (baseline retraining, Wilcoxon test p < 0.01), and transfer learning reduced adaptation time by 40% (mean 600 iterations, confidence interval [500, 700]). Robustness was validated by 100 runs for each variant, confirmed by ±10% parameter changes and MSE metrics. Sensitivity analysis was extended through real-world cross-ecosystem tests (e.g., indoor to mobile device migration). Hardware experiments were conducted, migrating the trained model from an office environment (15℃–35℃, static) to a drone flight environment (0–30℃, 0–10 m / s wind speed, n=50 runs per migration). Adaptation curves ( Figure 4The results show that the initial 5% error in a new environment is reduced to below 1% after transfer learning (average adaptation iterations of 600, 95% confidence interval [550, 650]), and the adaptation time is shortened by 40%. Uncertainty is also categorized and quantified. Specifically, internal uncertainties (such as sensor noise) are mitigated through robust filtering and reinforcement learning updates; external uncertainties (such as unpredictable environmental changes) are adaptively adjusted through state space. For parametric uncertainties (such as model coefficients), estimation using a parametric model with a Bayesian prior (e.g., using a normal prior for α in structural inference with a variance of 0.01) reduces the estimation error by 7% in testing. For nonparametric uncertainties (e.g., unknown parameters in rare events), see [link to relevant documentation]. Figure 4 The sensitivity curve for α variation is shown. Monte Carlo simulation was used for distribution analysis, with each scenario run 1000 times to ensure system stability. There is a trade-off in computational power. According to performance analysis of the STM32F4 Cortex-M4 chip, the Q-learning algorithm typically requires 50 KiB of memory, with inference time of 0.1 to 0.2 seconds (basic operations require 10,000-12,000 cycles, affected by cache and interrupts), and power consumption of 4 to 5 milliwatts. In contrast, the DQN algorithm requires only 0.3 to 0.5 seconds of inference time (a typical model requires approximately 45,000-50,000 cycles) and shares 4GB of memory. While offering a 20% improvement in accuracy, it increases power consumption by 25% to 30%. Therefore, Q-learning is suitable for low-power mobile devices, while DQN is more suitable for high-precision industrial applications. In resource-constrained scenarios, the lower computational requirements of Q-learning make it an ideal choice for battery-powered devices.

[0097] (3) Dynamic calibration: The calibration module performs spectral line fitting on the NDIR / TDLAS absorption spectrum at a wavelength of 4.26 μm and adjusts parameters in real time to compensate for drift through real-time correction (RL) operations. The specific process is as follows: acquire raw spectral data; construct the state according to the strategy and calculate the normalization matrix (Δk); update the normalization matrix (Δk) through the spectral line fitting model.

[0098] The calibration model is:

[0099] in, : The calibrated CO2 concentration; This represents the original concentration of CO2. This is the drift correction factor (initial value is 1.0). For temperature, For humidity, For pressure. It is a polynomial with a nonlinear effect. The error is calculated by comparing it with a standard gas and is used to guide the learning process.

[0100] The control process is as follows: Collect raw data and environmental parameters; generate states and calculate actions Δk according to the policy; update the value of k to k+Δk and apply it. The reward is evaluated based on the reference error, and the Q function is updated accordingly.

[0101] The core of the dynamic calibration method lies in applying time-varying signals, such as sinusoidal CO2 concentration changes with frequencies up to 1 Hz, to both the instrument under test (DUT) and a reference "fast" instrument (Picaro G2201-i, with a response time of less than 1 second). This method effectively analyzes frequency-dependent gain (amplitude ratio) and phase shift (time lag), thereby enabling spectral filtering correction in post-processing. The protocol specifies three frequencies of sinusoidal CO2 signals: 0.1 Hz (for slow drift, amplitude 500 ppm), 0.5 Hz (for moderate changes, amplitude 1000 ppm), and 1 Hz (for rapid changes, amplitude 2000 ppm). These signals are generated in a controlled chamber, co-located with the Picaro G2201-i reference instrument, and data is recorded at 1 Hz for Fourier transform analysis to ensure that the gain after calibration is less than 1.05 and the phase difference is less than 5 degrees.

[0102] This process is enhanced by reinforcement learning (RL), an algorithm that automatically adjusts the calibration coefficient k based on observed deviations, thus overcoming the limitations of traditional mean recalibration and adapting to nonlinear drift through action space optimization. The experimental protocol involves generating a signal (e.g., a 0.1Hz-1Hz sinusoidal CO2 wave with an amplitude of 500ppm-2000ppm using a Picaro reference source) in a controlled chamber, with a reference instrument co-located there. Data is recorded at a 1Hz frequency, and gain and phase are analyzed. For example, in a chamber test where humidity is gradually increased from 40% to 80% over 30 minutes (temperature fixed at 25°C), reinforcement learning dynamically adjusts the k value by 0.005-0.01 per cycle, reducing the phase shift from 20° to 5° while maintaining gain stability within 2% (variance 0.015). Figure 5 As shown, time series data validated these improvements.

[0103] (4) Intelligent control: The control module predicts CO2 trends and optimizes connected devices to achieve a balance between air quality and energy efficiency. The control process is as follows: predict future concentration trends using data from the past 10 minutes; assess the state and determine Vt based on current and predicted values; control actuators such as fans via PWM signals; update the strategy every 5 minutes to ensure real-time performance and continuous learning.

[0104] The prediction function is based on time series analysis, using CO2 concentration data from the past 10 minutes to predict the concentration for the next 5 to 10 minutes:

[0105] The predicted CO2 concentration; This represents the CO2 concentration over the past 10 minutes. It is a three-layer neural network. The model was trained using the UCI air quality dataset and augmented with a Gaussian process to accommodate CO2 (mean 800 ppm, standard deviation 200 ppm, sample size 1000, input stability normalized by z-value). A neural network-based prediction model utilizes the latest time-series data for short-term forecasting, enabling proactive control. This model was chosen for its ability to capture time-series patterns, effectively reducing trend prediction errors. PWM modulation technology is used for fine-tuning of the equipment in the control signal. Tests show a prediction error rate below 5%, and overfitting is effectively avoided through L2 regularization (λ=0.01-0.1) and an early termination strategy (no change for 10 consecutive cycles).

[0106] In summary, this application provides a method for controlling carbon dioxide concentration, which can accurately and quickly control the carbon dioxide concentration in the target environment.

[0107] To better implement the above methods, this application also provides a carbon dioxide concentration control device, which can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer; the server can be a single server or a server cluster composed of multiple servers.

[0108] For example, in this embodiment, the method of this application embodiment will be described in detail by taking the carbon dioxide concentration control device specifically integrated into the terminal as an example.

[0109] For example, such as Figure 6 As shown, the carbon dioxide concentration control device 600 may include a first unit 601, a second unit 602, a third unit 603, a fourth unit 604, and a fifth unit 605. The device includes: The first unit 601 is used to obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment. The second unit 602 is used to process the current concentration of carbon dioxide and the environmental parameters to obtain the state vector corresponding to the target environment. The third unit 603 is used to calculate the optimal state function value corresponding to the state vector based on the preset optimal equation calculation formula and the state vector, wherein the optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment. The fourth unit 604 is used to determine the action value function corresponding to the state vector based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment. The fifth unit 605 is used to filter out target actions from the action space corresponding to the action value function based on a preset filtering strategy, and execute the target actions to control the carbon dioxide concentration corresponding to the target environment.

[0110] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0111] As can be seen from the above, the embodiments of this application can accurately and quickly control the carbon dioxide concentration corresponding to the target environment.

[0112] This application also provides an electronic device, which can be a terminal, a server, or other similar device. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0113] In some embodiments, the product processing device may also be integrated into multiple electronic devices, such as multiple servers, with the multiple servers implementing the carbon dioxide concentration control method of this application.

[0114] In this embodiment, the electronic device will be described in detail as a terminal, for example, such as... Figure 7 As shown, it illustrates a structural schematic diagram of the terminal 700 involved in an embodiment of this application. Specifically: The terminal 700 may include components such as a processor 701 with one or more processing cores, a memory 702 with one or more media, a power supply 703, an input module 704, and a communication module 705. Those skilled in the art will understand that... Figure 7 The terminal 700 structure shown does not constitute a limitation on the terminal 700, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 701 is the control center of the terminal 700. It connects various parts of the terminal 700 via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, thereby providing overall monitoring of the terminal 700. In some embodiments, the processor 701 may include one or more processing cores; in some embodiments, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 701.

[0115] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 includes a program storage area and a data storage area. The program storage area can store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area can store data created according to the use of the terminal 700, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0116] The terminal 700 also includes a power supply 703 that supplies power to the various components. In some embodiments, the power supply 703 can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0117] The terminal 700 may also include an input module 704, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0118] The terminal 700 may also include a communication module 705. In some embodiments, the communication module 705 may include a wireless module. The terminal 700 can perform short-range wireless transmission through the wireless module of the communication module 705, thereby providing users with wireless broadband internet access. For example, the communication module 705 can be used to help users send and receive emails, browse web pages, and access streaming media.

[0119] Although not shown, terminal 700 may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in terminal 700 loads the executable files corresponding to the processes of one or more applications into memory 702 according to the following instructions, and the processor 701 runs the applications stored in memory 702 to realize various functions, as follows: Obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment; The current carbon dioxide concentration and the environmental parameters are processed to obtain the state vector corresponding to the target environment. Based on the preset optimal equation calculation formula and the state vector, the optimal state function value corresponding to the state vector is calculated, wherein the optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment. Based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, the action value function corresponding to the state vector is determined, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment. Based on a preset screening strategy, target actions are selected from the action space corresponding to the action value function, and the target actions are executed to control the carbon dioxide concentration corresponding to the target environment.

[0120] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0121] As can be seen from the above, the embodiments of this application can achieve precise and rapid control of the carbon dioxide concentration in the target environment.

[0122] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be accomplished by instructions, or by instructions controlling related hardware. These instructions can be stored in a medium and loaded and executed by a processor.

[0123] Therefore, embodiments of this application provide a medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the carbon dioxide concentration control methods provided in embodiments of this application. For example, the instructions can execute the following steps: Obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment; The current carbon dioxide concentration and the environmental parameters are processed to obtain the state vector corresponding to the target environment. Based on the preset optimal equation calculation formula and the state vector, the optimal state function value corresponding to the state vector is calculated, wherein the optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment. Based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, the action value function corresponding to the state vector is determined, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment. Based on a preset screening strategy, target actions are selected from the action space corresponding to the action value function, and the target actions are executed to control the carbon dioxide concentration corresponding to the target environment.

[0124] The medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0125] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a medium. A processor of a computer device reads the computer instructions from the medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.

[0126] Since the instructions stored in the medium can execute the steps in any of the carbon dioxide concentration control methods provided in the embodiments of this application, the beneficial effects that any of the carbon dioxide concentration control methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0127] The above provides a detailed description of a carbon dioxide concentration control method, apparatus, terminal, and medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method of controlling carbon dioxide concentration, characterized by, The method includes: Obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment; The current carbon dioxide concentration and the environmental parameters are processed to obtain the state vector corresponding to the target environment. Based on the preset optimal equation calculation formula and the state vector, the optimal state function value corresponding to the state vector is calculated, wherein the optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment. Based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, the action value function corresponding to the state vector is determined, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment. Based on a preset screening strategy, target actions are selected from the action space corresponding to the action value function, and the target actions are executed to control the carbon dioxide concentration corresponding to the target environment.

2. The method of claim 1, wherein, The step of determining the action value function corresponding to the state vector based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value includes: Based on the aforementioned scenario complexity, the processing method of the preset reinforcement learning module for the state vector is determined; Based on the optimal state function value, the state vector is processed by the preset reinforcement learning module to obtain the action value function.

3. The method of claim 2, wherein, The step of determining the processing method of the preset reinforcement learning module for the state vector based on the scene complexity includes: If the scene complexity is less than the preset complexity threshold, then the state vector is processed by calling the preset function of the preset reinforcement learning module; If the scene complexity is greater than or equal to a preset complexity threshold, then the state vector is determined to be processed by a preset deep network of a preset reinforcement learning module.

4. The method of claim 1, wherein, The step of filtering target actions from the action space corresponding to the action value function based on a preset filtering strategy includes: Determine the screening strategy value and the probability value respectively, wherein the screening strategy value is greater than 0 and less than 1, the probability value is greater than or equal to 0, and the probability value is less than 1; If the probability value is less than the filtering strategy value, then any action in the action space is determined as the target action; If the probability value is greater than or equal to the filtering strategy value, then the action corresponding to the maximum action value function value in the action space is determined as the target action.

5. The method of claim 1, wherein, The target action includes adjusting the correction coefficient; The execution of the target action includes: Adjust the correction coefficients to obtain the adjusted correction coefficients; Substitute the adjusted correction coefficients into the preset calibration model to calculate the calibrated carbon dioxide concentration. The preset hardware interface controls the preset action hardware to adjust the current carbon dioxide concentration in the target environment to the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration.

6. The method of claim 5, wherein, After performing the target action, the method further includes: Based on the target action, the preset carbon dioxide reference concentration, the current carbon dioxide concentration, the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration, and the calibrated carbon dioxide concentration, a reward function value is calculated using a preset reward function. Based on the reward function value and the state vector corresponding to the calibrated carbon dioxide concentration, the parameters of the preset reinforcement learning module are updated to obtain the updated preset reinforcement learning module.

7. The method of claim 6, wherein, The target action includes a first type of action and a second type of action, and the preset reward function includes a first preset reward function and a second preset reward function; The step of calculating a reward function value based on the target action, a preset carbon dioxide reference concentration, the current carbon dioxide concentration, the target carbon dioxide concentration corresponding to the calibrated carbon dioxide concentration, and the calibrated carbon dioxide concentration, through a preset reward function, includes: If the target action is the first type of action, then the preset carbon dioxide reference concentration and the calibrated carbon dioxide concentration are substituted into the first preset reward function to calculate the reward function value; If the target action is the second type of action, then the current carbon dioxide concentration and the target carbon dioxide concentration are substituted into the second preset reward function to calculate the reward function value.

8. A device for controlling carbon dioxide concentration, characterized in that, The device includes: The first unit is used to obtain the current carbon dioxide concentration and environmental parameters corresponding to the target environment. The second unit is used to process the current concentration of carbon dioxide and the environmental parameters to obtain the state vector corresponding to the target environment. The third unit is used to calculate the optimal state function value corresponding to the state vector based on the preset optimal equation calculation formula and the state vector. The optimal state function value is used to characterize the expected cumulative reward that can be obtained by taking the optimal action in the current state of the target environment. The fourth unit is used to determine the action value function corresponding to the state vector based on the state vector, the scene complexity corresponding to the target environment, and the optimal state function value, wherein the action value function is used to characterize the relationship between all actions and the expected cumulative reward in the current state of the target environment. The fifth unit is used to select target actions from the action space corresponding to the action value function based on a preset screening strategy, and execute the target actions to control the carbon dioxide concentration corresponding to the target environment.

9. A terminal, characterized in that, The method includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps of the method as described in any one of claims 1 to 7.

10. A medium, characterized in that, The medium stores a plurality of instructions adapted for loading by a processor to execute the steps of the method according to any one of claims 1 to 7.