Small space carbon neutralization strategy generation system based on deep reinforcement learning inference

By combining deep reinforcement learning and fluid simulation engines, a digital carbon space is constructed, which solves the problems of diversity and efficiency in generating carbon neutrality strategies in small spaces, and achieves efficient carbon concentration prediction and strategy optimization.

CN116307387BActive Publication Date: 2026-06-02SHANGHAI JIAOTONG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI JIAOTONG UNIV
Filing Date
2023-03-17
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methods for generating carbon neutralization strategies in small spaces are difficult to adapt to diverse spaces, have low efficiency in carbon concentration prediction, and have low efficiency in adjusting carbon neutralization strategies, making rapid iteration impossible.

Method used

We employ a deep reinforcement learning-based approach to construct a digital carbon space using a fluid simulation engine, generate carbon concentration monitoring sequences using a spatial case library, and combine this with a carbon concentration prediction model to optimize carbon neutrality strategies and generate efficient spatial carbon neutrality strategies.

Benefits of technology

It enables efficient carbon concentration prediction and carbon neutrality strategy generation for any small space, improves automation and execution efficiency, and supports rapid strategy iteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116307387B_ABST
    Figure CN116307387B_ABST
Patent Text Reader

Abstract

A small space carbon neutralization strategy generation system based on deep reinforcement learning reasoning comprises a feature analysis module, a digital carbon space initialization module, a carbon neutralization strategy optimization module, a carbon neutralization strategy derivation module, a model training support module and a carbon concentration prediction module. The present application drives the construction of a digital carbon space based on a fluid simulation engine based on physical formulas, generates a focus concentration monitoring sequence based on a massive space case library, constitutes a training set, extracts rich space-time semantics in the training set through a carbon concentration prediction model, forms an efficient reasoning method for space carbon concentration, uses a deep reinforcement learning model to reason and optimize for any given small space, and thus can directly reason the carbon concentration of a given space-time. The carbon concentration prediction model is used as an interactive environment to construct and train a space carbon neutralization strategy optimization module based on deep reinforcement learning. For the space environment point cloud model, carbon sink point cloud model and carbon source configuration input by the user, an efficient space carbon neutralization strategy is reasoned and generated, and the complete point cloud model and related configuration of the corresponding strategy are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technology in the field of carbon neutrality, specifically a small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference. Background Technology

[0002] Existing methods for generating carbon neutrality strategies for small spaces still have certain problems. First, small spaces are diverse, with varying external shapes and internal layouts. Some existing digital carbon space construction methods rely on software simulation, but the results are completely tied to a given space; others are based on fluid simulation engines, but lack the abstraction of carbon sinks, making them difficult to apply to diverse spaces. Second, the state of small spaces changes dynamically, and existing carbon concentration prediction methods have low execution efficiency, making it difficult to accurately and efficiently predict the carbon concentration at a given time and space. Finally, the elements in small spaces are highly interconnected, and changes in one can have far-reaching consequences. Existing carbon neutrality strategy generation methods rely on manual control of strategy adjustments, making it difficult to support rapid modification and iteration of carbon neutrality strategies. Summary of the Invention

[0003] This invention addresses the shortcomings of existing technologies that fail to consider the impact of carbon sinks (carbon absorption sources) or whose carbon sink locations are fixed and only consider local areas centered on the carbon sink, thus neglecting to quantify the overall net-zero carbon emissions of the space. It proposes a small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference. This system uses a fluid simulation engine based on physical formulas to drive the construction of a digital carbon space. A training set is formed by generating focal concentration monitoring sequences based on a massive spatial case library. A carbon concentration prediction model extracts rich spatiotemporal semantics from the training set, forming an efficient inference method for spatial carbon concentration. For any given small space, a deep reinforcement learning model is used for inference optimization, enabling direct inference of the carbon concentration at a given time and space. The carbon concentration prediction model is used as an interactive environment to construct and train a spatial carbon neutrality strategy optimization module based on deep reinforcement learning. Based on the user-input spatial environment point cloud model, carbon sink point cloud model, and carbon source configuration, an efficient spatial carbon neutrality strategy is generated, and the complete point cloud model and related configurations of the corresponding strategy are output.

[0004] This invention is achieved through the following technical solution:

[0005] This invention relates to a small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference, comprising: a feature parsing module, a digital carbon space initialization module, a carbon neutrality strategy optimization module, a carbon neutrality strategy derivation module, a model training support module, and a carbon concentration prediction module. The model training support module uses a spatial case library as its data source. By constructing and running a data carbon space, it collects carbon concentration monitoring values ​​of key points, processes them together with case data to form a training set, and trains the carbon concentration prediction model. The feature parsing module preprocesses the original environmental point cloud model to filter out invalid information, extracts static spatial features, and parses the carbon source parameter file to obtain position and velocity information for constructing the digital carbon space. The digital carbon space initialization module... The module converts spatial features into relative sizes, then initializes the carbon sink layout based on the total carbon source emission rate, and deploys monitors according to environmental characteristics to form a carbon neutrality strategy to be optimized. The carbon neutrality strategy optimization module changes the internal layout of the digital carbon space by adjusting the carbon neutrality strategy, interacts with the carbon concentration prediction module to obtain the current net carbon emissions of the space, and continuously optimizes the carbon neutrality strategy by adjusting the spatial layout. The carbon neutrality strategy export module registers the input environmental point cloud model and carbon sink point cloud model according to the optimized carbon neutrality strategy to form an intuitive and easy-to-use complete spatial point cloud model, saves the dynamic configuration of carbon sources, carbon sinks and monitors in the strategy in a file, and finally outputs the complete point cloud model and dynamic configuration file as the processing result.

[0006] The feature parsing module imports and recognizes the user-input point cloud model and parameter configuration file, extracts carbon spatial features from the environmental point cloud model and carbon source configuration as valid input for subsequent modules. This feature parsing module includes: a point cloud model preprocessing unit, an environmental feature extraction unit, and a carbon source parameter extraction unit. Specifically: the point cloud model preprocessing unit preprocesses the user-input point cloud model, removing noise caused by device precision and object material through data denoising; it also filters point cloud data while maintaining accuracy through data simplification, extracting effective model information; the environmental feature extraction unit parses the preprocessed point cloud model, extracting the model's external boundaries and the boundaries of internal obstacles and carbon sources, identifying the relative positions of ventilation openings and obstacles, and forming key-value pairs; the carbon source parameter extraction unit takes the carbon source parameter file as input, extracts the attribute configuration related to the carbon source through scene keyword retrieval in text analysis, constructs multiple sets of key-value pairs with the carbon source as the object; and compares and fuses the relative values ​​of environmental features with the absolute values ​​of carbon source parameters to obtain an absolute representation of spatial features, ultimately forming a serialized carbon spatial feature representation.

[0007] The aforementioned attribute configurations include: quantity, location, and emission and absorption rates.

[0008] The digital carbon space initialization module initializes the digital carbon space based on carbon space characteristics, forming a basic spatial carbon neutrality strategy as input for subsequent neural reinforcement learning model inference optimization. This module includes a carbon space feature conversion unit, a carbon sink layout generation unit, and a monitor placement unit. Specifically, the carbon space feature conversion unit normalizes spatial information, including absolute size, during digital carbon space initialization, keeping it numerically within the range of [0,1], thus initially initializing the digital carbon space. The carbon sink layout generation unit sets the carbon sink layout, calculating the maximum total carbon emission rate of the carbon source at all times based on the carbon source spatiotemporal sequence. Based on this, a set of spatiotemporal sequences of carbon sinks are randomly initialized, where the total carbon absorption rate of the carbon sink at any given time is... This ensures that the carbon sink can absorb the total carbon emissions from the carbon source; the monitoring unit places a monitor at each ventilation opening and between each pair of carbon sources and carbon sinks that are close to each other to avoid excessively high local carbon concentrations indoors.

[0009] The carbon neutrality strategy optimization module adjusts the layout of carbon sources, carbon sinks and monitors by calling the carbon neutrality strategy optimization module, evaluates the results of each adjustment and then updates the layout, and iterates until the output value of the carbon neutrality strategy optimization module tends to stabilize.

[0010] The carbon neutrality strategy optimization module incorporates a deep reinforcement learning model. It uses a small-space carbon neutrality strategy as input, a carbon concentration prediction model as the interaction environment, and net-zero carbon emissions from the space as the core metric for training. The inference process uses an unoptimized carbon neutrality strategy as input and an optimized carbon neutrality strategy as output. The quality of each state is determined by the reward. This carbon neutrality strategy optimization module includes a spatial feature extraction unit, a deep convolutional network computation unit, and an action selection unit. Specifically, the spatial feature extraction unit constructs a high-dimensional input vector based on the input carbon neutrality strategy and extracts spatial features using principal component analysis to obtain a low-dimensional abstract vector containing rich features. The deep convolutional network computation unit uses this low-dimensional vector as input to calculate and output the target value Q corresponding to the state. The action selection unit selects the optimal next state based on the target value Q and the corresponding reward and outputs it.

[0011] The aforementioned reward Where: Co is the total cost of the carbon sink device, Rt is the total energy consumption of the carbon sink and monitors, Ct is the total carbon emissions emitted to the outside world, and M represents the changes in the monitoring values ​​of all monitors. , , and As the weight of each part The specific values ​​are selected based on the specific circumstances. The lower the total cost Co, total energy consumption Rt, and total external carbon emissions Ct, the better, while the monitoring values ​​should be as close as possible to... The better, as too high or too low a concentration will result in a lower reward.

[0012] Changes in the monitored values ​​of all the monitors mentioned Where: T is the number of time intervals, and k is the number of monitors. This represents the average atmospheric carbon dioxide concentration. Let t be the monitoring value of the i-th monitor at time t.

[0013] The carbon dioxide concentration That is, the carbon concentration is 400 ppm.

[0014] The total carbon emissions from the window to the outside Where: T is the number of time intervals, and q is the number of windows. Let be the carbon concentration in the i-th window at time t, and let the time interval be fixed. Regardless of how the carbon neutrality strategy sets up the monitor, each prediction requires monitoring the carbon concentration values ​​for each window.

[0015] The carbon neutrality strategy optimization module performs experience replay before training the model and randomly selects states. That is, taking the carbon neutrality strategy as input, calculating each action. That is, the reward corresponding to the strategy adjustment. And the new status This corresponds to the new carbon neutrality strategy, recording the quadruple. To form an experience pool, we repeatedly use these experiences to train the target network. Specifically, we set a batch size (`batch_size`). Within each batch, we use the target network to calculate the evaluation value of the reward, without updating its parameters; only the parameters of the evaluation network are updated. After each batch, we update the parameters of the target network once. When the prediction value given by the evaluation network in each iteration... The target value of the target network is The corresponding loss function Where: r is the reward for the current action, The discount rate decays over time. To evaluate the network parameters, Let E[x] be the parameter of the target network, and E[x] be the expectation of the random variable x.

[0016] The parameters of the target network are updated in any of the following ways:

[0017] a) Directly assign the parameters of the evaluation network to the target network to achieve a hard update;

[0018] b) By introducing a learning rate The target network is then assigned a weighted average of the old target network parameters and the new evaluation network parameters to achieve a soft update. Among them: selecting the soft update method for updating. The model converges when the loss function value stabilizes. In subsequent inference, only the target network is used to calculate the optimal action a for a given state s, i.e., to generate the optimal carbon neutralization strategy.

[0019] The carbon neutrality strategy export module exports the optimized spatial carbon neutrality strategy and generates a corresponding configuration file. This module includes: a parameter configuration export unit, an environment configuration export unit, a point cloud model export unit, and a point cloud model registration unit. Specifically: the point cloud model registration unit registers the carbon sink point cloud model with the environment point cloud model based on the quantity and location information of carbon sinks in the optimized carbon neutrality strategy, forming a complete net-zero carbon spatial point cloud model; the point cloud model export unit exports the complete point cloud model for subsequent modeling and analysis; the environment configuration export unit records the spatial shape and obstacle location information, forming an environment configuration file; and the parameter configuration export unit records the changes in the location and absorption rate of carbon sources and carbon sinks over time, forming carbon source configuration files and carbon sink configuration files, and ultimately creating a complete spatial carbon neutrality strategy output file.

[0020] The model training support module enhances the spatial data of the spatial case library, generating a training set for the carbon concentration prediction model. This module includes: a case data preprocessing unit, a digital carbon space construction unit, a digital carbon space operation unit, and a spatiotemporal sequence post-processing unit. Specifically, the case data preprocessing unit transforms the spatial location information of the data into corresponding locations that the fluid simulation engine can recognize and use. Then, it performs multi-sequence time alignment operations and obtains the location and velocity of the carbon source and carbon sink at each moment by linear interpolating the spatiotemporal sequences of carbon sources and sinks, ultimately yielding the preprocessed case data. The digital carbon space construction unit identifies the configuration of the digital carbon space by importing preprocessed case data, and then initializes the fluid simulation engine to obtain the constructed digital carbon space. The digital carbon space operation unit simulates the diffusion process of carbon dioxide gas in a given space by running the digital carbon space. By randomly selecting multiple monitoring points, it records the gas concentration reading sequence. The resulting carbon concentration monitoring sequence can reflect the overall gas state of the digital carbon space over time. The spatiotemporal sequence post-processing unit generates the model input spatiotemporal sequence after transforming the carbon concentration monitoring sequence data and aligning the time series data.

[0021] The case data includes: fixed static spatial information, i.e., spatial shape, such as cuboid, cylinder, etc.; information on ventilation openings to the outside of the space, i.e., the location, shape and size of windows; and information on obstacles inside the space, i.e., their location, shape and size; and variable dynamic spatial information, including carbon source spatiotemporal sequences and carbon sink spatiotemporal sequences, wherein the carbon source spatiotemporal sequence can be described in detail as the location and carbon emission rate of each carbon source at each moment; and the carbon sink spatiotemporal sequence can be described in detail as the location and carbon absorption rate of each carbon sink at each moment.

[0022] The fluid simulation engine uses a fluid simulation algorithm based on Euler's perspective, which focuses on a fixed set of grid points that constitute the space. By recording changes in data such as gas concentration, velocity, and temperature at the grid points, it reflects the overall changes in the gas.

[0023] The training set generated by the model training support module is further processed by carbon concentration data transformation and time-series data time alignment.

[0024] The aforementioned carbon concentration data transformation refers to the following: Because the simulation engine extrapolates from physical formulas and real-world environmental values, the carbon concentration monitoring sequences generated by the digital carbon space often have generally small values, mostly in the range of 0.0001 to 0.01. Simultaneously, there are also cases with extremely small carbon concentration values, typically tens of negative powers, and both positive and negative values ​​exist. Therefore, based on the characteristics of the data, a data transformation based on polynomial functions is used to process the data, and normalization is performed on a case-by-case basis.

[0025] The aforementioned time-series data time alignment refers to the following: Since the carbon concentration monitoring sequence is generated by running the digital carbon space, its time interval is a fixed value set by the physics engine. However, the time interval between the carbon source spatiotemporal sequence and the carbon sink spatiotemporal sequence is often larger. Therefore, it is necessary to sample the carbon source and carbon sink according to the time interval of carbon concentration to obtain the model input spatiotemporal sequence. When a numerical mutation point is encountered, the value after the mutation is used as the sampled value.

[0026] The model input spatiotemporal sequence includes the model input carbon source sequence, the model input carbon sink sequence, and the model input monitor sequence. All model input spatiotemporal sequences are aggregated to form a training set, which serves as data support for the carbon concentration prediction model.

[0027] The carbon concentration prediction module predicts the carbon concentration at any given time and space. Specifically, it simulates the carbon concentration changes in the digital carbon space corresponding to a carbon neutrality strategy to obtain quantitative indicators reflecting the strategy's effectiveness, thus evaluating the strategy. The training process of the carbon concentration prediction model uses the training set as input, first extracting spatial features through a fully connected network, then extracting temporal features through a temporal feature extraction network, obtaining the predicted carbon concentration value for the monitor at subsequent time points. The inference process uses a given carbon neutrality strategy as input and its score as output, reflecting the strategy's effectiveness through the score. Specifically, it includes:

[0028] 1) Normalize the input information. When the longest side in the space is h, all values ​​x related to the absolute size of the space, such as position and velocity, are normalized to: To avoid the influence of values ​​that are too large or too small, regarding the relationship between the carbon source and the monitor, let the current time be t, and the position of carbon source 1 be... carbon emission direction ,rate Position of monitor m Current monitoring value The directional relationship between the carbon source and the monitor The direction vector can be obtained by combining the carbon emission direction. The carbon source vector can be obtained by combining the results. .

[0029] 2) Regarding the relationship between carbon sinks and monitors, let the current time be t, and the location of carbon sink 2 be... ,rate Position of monitor m Current monitoring value The relationship between carbon sequestration and the direction of the monitor. Since the carbon sink absorption range is spherical and the direction points towards the center of the sphere, the direction vector can be obtained. The carbon sink vector can be obtained by combining the results. .

[0030] 3) Abstract the carbon source and carbon sink as cuboids. For this relationship between the obstacle and the monitor, let the current time be t, and the obstacle position... Base length w, height h, monitor position m Current monitoring value The directional relationship between the obstacle and the monitor , direction vector The obstacle vector can be obtained by combining the results. .

[0031] 4) To , and Spatial feature extraction was performed using different fully connected networks, resulting in a unified 1*C dimensional vector. For example, the fully connected network structure is as follows: It consists of two fully connected layers: a hidden layer and an output layer. The hidden layer produces a 1*H dimensional vector, and a linear rectified activation function is applied before the output layer. Activate, all Directly adding them together yields a 1*C dimensional vector, representing the influence of all carbon sources in space on the monitor m at time t. and The processing is the same as above, thus obtaining three 1*C dimensional vectors, which are concatenated into a 1*3C dimensional vector, serving as the influence of space on monitor m at time t, and also as the input of the temporal feature extraction network at time t.

[0032] 5) A temporal feature extraction network consisting of an input layer, hidden layers, and an output layer extracts the temporal features of the data at time t-1 and outputs the model's predicted carbon concentration for monitor m at time t. The difference is that the hidden layer incorporates the previous hidden layer vector into the current calculation. The specific calculation of the temporal feature extraction network is as follows: , ,in: The input at time t, The result of the hidden layer at time t. Let V be the output of the output layer at time t. V is the weight matrix of the output layer, and g is the activation function. (This is because solving...) The hidden layer takes into account the result of the previous time step, so it is also a recurrent layer, and U is the input. The weight matrix, where W is the value at the previous time step. The weight matrix used as input at this moment, f, is the activation function. Carbon concentration prediction is a regression problem; therefore, the mean squared loss function MSELoss() is used as the loss function for model training, specifically: Where: X is the actual carbon concentration sequence, Y is the model-predicted sequence, and n is the sequence length. and These are the actual value and the predicted value at time t, respectively.

[0033] Technical effect

[0034] Compared to the techniques commonly used for gas simulation, namely the use of physics-based fluid simulation engines, this invention incorporates the abstraction of carbon sinks into the engine, thereby improving the accuracy of gas simulation.

[0035] Compared to techniques commonly used to predict carbon dioxide concentration at a given time and space, this invention trains a carbon concentration prediction model using a rich and complete training set generated by a fluid simulation engine. By combining physics-based simulation methods with machine learning-based feature extraction methods, the carbon concentration prediction method has good accuracy and generalization, enabling the prediction of carbon concentration at any time and space, regardless of the scene.

[0036] Compared to the techniques commonly used to optimize carbon neutrality strategies, which involve fully modeling dynamic scenarios, this invention uses static features and some dynamic features as inputs and a reasoning method based on a deep reinforcement learning model to achieve strategy evaluation and adjustment. This avoids the gas simulation process and greatly improves the automation and execution efficiency of carbon neutrality strategy reasoning. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the system of the present invention;

[0038] Figure 2 This is a schematic diagram illustrating the application of an example. Detailed Implementation

[0039] like Figure 1 As shown in the figure, this embodiment relates to a small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference, including: a feature parsing module, a digital carbon space initialization module, a carbon neutrality strategy optimization module, a carbon neutrality strategy derivation module, a model training support module, and a carbon concentration prediction module.

[0040] like Figure 2 As shown, the embodiment is divided into a web application, a spatial carbon neutrality strategy optimization system, and an infrastructure layer. The spatial carbon neutrality strategy optimization system is the main body of the entire implementation framework and the core of this invention. It responds to calls from the interaction layer and provides carbon neutrality strategy optimization services. It also meets the data processing and model dependency requirements of internal functions by calling data transmission interfaces and model call interfaces. It includes a carbon space builder, a carbon neutrality strategy generator, a model training set generator, and a model trainer. Through the interaction and collaboration between the modules, it supports the reception, processing, construction, optimization, and export of point cloud models and configuration data.

[0041] The carbon space builder includes a point cloud model preprocessing module, a spatial feature extraction module, and a digital carbon space initialization module. The point cloud model preprocessing module performs data denoising and data simplification on the model, filtering out and retaining effective information. The spatial feature extraction module first parses the preprocessed point cloud model, extracting the model's external boundaries and the boundaries of internal obstacles and carbon sources, identifying static information such as the relative positions of ventilation openings and obstacles, forming key-value pairs. Then, it parses the user-input carbon source parameter file, extracting the attribute configurations related to the carbon sources, including quantity, location, emission, and absorption rates, forming multiple key-value pairs based on the carbon source. Finally, it compares and fuses the relative values ​​of environmental features with the absolute values ​​of carbon source parameters to obtain an absolute representation of the spatial features, ultimately forming a serialized carbon space feature representation. The digital carbon space initialization module scales the extracted spatial features to convert them into the digital carbon space, then sets the locations of carbon sinks and monitors, completing the initial construction of the digital carbon space.

[0042] The digital carbon space initialization module includes a carbon space feature conversion unit, a carbon sink layout generation unit, and a monitor placement unit. Specifically, the carbon space feature conversion unit, during digital carbon space initialization, normalizes all spatial information, including absolute values, such as the positions of obstacles and carbon sources, keeping them numerically within the range of [0,1]. It also calculates the maximum total carbon emission rate of the carbon source at all times based on the carbon source spatiotemporal sequence. Based on this, a set of spatiotemporal sequences of carbon sinks are randomly initialized, where the total carbon absorption rate of the carbon sink at any given time is... This ensures that the carbon sink can absorb the total carbon emissions from the carbon source. To reduce the space's net-zero carbon emissions to the outside, it is necessary to monitor the carbon concentration at each ventilation opening. Therefore, a monitor needs to be placed at each ventilation opening. In addition, monitors should be placed between each pair of carbon sources and carbon sinks that are close to each other to avoid excessively high local carbon concentrations indoors.

[0043] The carbon neutrality strategy generator uses a carbon neutrality strategy optimization module to perform inference optimization on an initialized digital carbon space to generate a spatial carbon neutrality strategy with net-zero carbon emissions. The initialized digital carbon space corresponds to a carbon neutrality strategy to be optimized, which does not effectively balance cost and carbon removal. Therefore, the carbon neutrality strategy optimization module uses a carbon concentration prediction model as its interactive environment, and energy consumption, carbon sink costs, and the monitoring values ​​output by the carbon concentration prediction model as rewards. Within a given small space, it solves for the optimal carbon neutrality strategy by adjusting the carbon sink and monitor layout. The optimized carbon neutrality strategy is finally exported to the corresponding configuration file and output to the user.

[0044] The model training set generator includes a digital carbon space construction module, a carbon concentration monitoring sequence generation module, and a spatiotemporal sequence post-processing module. Specifically, it processes the spatial location information of the data into corresponding locations that the fluid simulation engine can recognize and use through data spatial location transformation. Then, it performs multi-sequence time alignment operations. By linearly interpolating the spatiotemporal sequences of carbon sources and sinks, it obtains the location and rate of carbon sources and sinks at each moment, forming preprocessed case data. The preprocessed case data is imported to identify the configuration of the digital carbon space. Then, the fluid simulation engine is initialized to construct the digital carbon space. The fluid simulation engine uses a fluid simulation algorithm based on the Euler perspective, which does not treat gas as particles moving according to physical laws, but instead focuses on a set of fixed grid points constituting the space. By recording changes in gas concentration, rate, and temperature at these grid points, it reflects the overall gas changes. Running the digital carbon space simulates the diffusion process of carbon dioxide gas in a given space. By randomly selecting multiple monitoring points and recording their gas concentration reading sequences, the resulting carbon concentration monitoring sequence reflects the overall gas state within the digital carbon space over time. Taking a single case as a unit, the carbon concentration sequence is normalized by polynomial function transformation, and carbon sources and carbon sinks are sampled according to the time interval of the carbon concentration sequence to obtain the spatiotemporal sequence of the model input, which is then summarized to form a training set.

[0045] The model trainer includes a carbon concentration prediction model training module and a carbon neutrality strategy optimization module training module.

[0046] The carbon concentration prediction model training module trains the carbon concentration prediction model based on the training set, and the model prototype is based on an LSTM neural network. The input at each time step is processed by two fully connected layers to extract spatial features, forming a 1*4 dimensional vector which serves as the input to the LSTM network. For each training iteration, the LSTM performs best when predicting the fifth time step using inputs from four consecutive time steps. The LSTM has two hidden layers, each with a channel size of 4, and the output layer has a channel size of 1, which represents the final prediction result. During training, the Adam optimization function is used, with a base learning rate of 0.005, a weight decay hyperparameter of 0.01, and iterative updates to the model weights. The loss function used is MSELoss.

[0047] The carbon neutrality strategy optimization module training module uses the DQN model as a prototype and a carbon concentration prediction model as the interactive environment. The experience replay pool size is set to 1000. Both the target network and the evaluation network contain two hidden layers. The base learning rate is set to 0.001, the decay factor to 0.95, and the positive and negative reward values ​​to 0.01 and -1, respectively. The rewards are based on the monitored values ​​of energy consumption, carbon sink cost, and the carbon concentration prediction model output. The parameters are set as follows: Given a small space, construct a solution space and update the rewards under different spatial strategies.

[0048] Table 1 shows a comparison of the technical specifications of this embodiment with the technical characteristics of the prior art.

[0049] Technical content This embodiment Existing technology System functional objectives Automatically generate carbon neutrality strategies for arbitrary spaces Manually generate carbon neutrality strategies for a specified space Feature Analysis It supports extracting features from point cloud models and parsing parameter configurations from configuration files. Importing features via files or manually entering relevant configurations results in a high error rate. Digital carbon space construction Based on a fluid dynamics engine, gas simulation is achieved by adding an abstraction of a gas absorption source to the original model, which can construct complex and varied digital carbon spaces. Gas simulation is achieved using a fluid dynamics engine, which only abstracts gas emission sources and obstacles, but does not abstract gas absorption sources. Carbon neutrality strategy generation method Using static spatial features and some dynamic spatial features as input, inference optimization is performed through a deep reinforcement learning model. Complete static and dynamic scenes are modeled, and repeated experiments are conducted by manually adjusting the scene layout. Configuration export In addition to exporting information such as environmental features, carbon source and carbon sink parameters, it also registers the environmental point cloud model and the carbon sink point cloud model, and exports the complete point cloud model for users to use. Relying on manual recording of configuration information such as environmental characteristics, carbon source and carbon sink parameters has a high error rate; and the lack of subsequent usage methods makes it inconvenient. Training set generation Physics-based methods were used to augment the spatial case library, resulting in rich training data. Simulating gas diffusion based on physical methods requires no related technical support. Carbon concentration prediction methods Based on machine learning models that extract spatiotemporal features, the carbon concentration at a given time and space can be predicted directly with a fast prediction rate. It simulates the entire gas diffusion process based on physical formulas and predicts all locations in a given space, but the prediction rate is relatively slow. Generalization Physics-based simulation methods generate accurate and complete datasets, and machine learning models are used to extract rich spatiotemporal features from the datasets, supporting the prediction of carbon concentration in any space. It only supports carbon concentration prediction for a specified space, and the results cannot be transferred or used. Ease of use The output of the preceding module is directly used as the input of the subsequent module. Users only need to upload the scanned point cloud model and related configuration files. During the execution of the method, users need to manually process the point cloud model, adjust the strategy, and start the simulation process. Execution efficiency The carbon concentration prediction method avoids the gas simulation process and only calculates the carbon concentration in a given time and space, resulting in high execution efficiency. Carbon concentration prediction requires repeating the gas simulation process and calculating the carbon concentration at all locations in space, which is inefficient. Maintainability The method has a clear overall structure, simple internal module implementation, and is easy to understand and modify. Physics-based methods require understanding the relevant formulas for simulation, and the implementation process is complex, difficult to understand, and difficult to modify.

[0050] This system fully utilizes real-world case data to construct a digital carbon space based on physical simulation methods, generating a rich and complete dataset. Machine learning models are then used to extract the rich spatiotemporal semantics within this dataset, improving the generalization ability of the method. The carbon neutrality strategy optimization module, based on deep reinforcement learning, directly obtains simulation results by calling a carbon concentration prediction model, thereby evaluating and adjusting the strategy. This avoids manual strategy adjustments and simulations of gas diffusion processes, improving the system's usability and execution efficiency. Furthermore, the overall structure of the method is clear, the internal logic of the modules is simple, making it easy to understand, implement, and modify, and ensuring good maintainability.

[0051] Compared with existing technologies, this invention solves the problems of unpredictable internal carbon concentration, poor accuracy and low efficiency in generating carbon neutrality strategies in small building spaces. Furthermore, it fully utilizes existing spatial case data to participate in model training, improving the system's intelligence level and providing strong data and technical support for generating carbon neutrality strategies.

[0052] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.

Claims

1. A small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference, characterized in that, include: The system comprises a feature parsing module, a digital carbon space initialization module, a carbon neutrality strategy optimization module, a carbon neutrality strategy export module, a model training support module, and a carbon concentration prediction module. Specifically: the model training support module uses a spatial case library as its data source. By constructing and running a data carbon space, it collects carbon concentration monitoring values ​​at key points, processes them together with case data to form a training set, and trains the carbon concentration prediction model; the feature parsing module preprocesses the original environmental point cloud model to filter out invalid information, extracts static spatial features, and parses the carbon source parameter file to obtain location and velocity information for constructing the digital carbon space; the digital carbon space initialization module converts spatial features into relative sizes, and then... The carbon sink layout is initialized based on the total emission rate of carbon sources, and monitors are deployed according to environmental characteristics to form a carbon neutrality strategy to be optimized. The carbon neutrality strategy optimization module changes the internal layout of the digital carbon space by adjusting the carbon neutrality strategy, and interacts with the carbon concentration prediction module to obtain the current net carbon emissions of the space. The carbon neutrality strategy is optimized by continuously adjusting the spatial layout. The carbon neutrality strategy export module registers the input environmental point cloud model and carbon sink point cloud model according to the optimized carbon neutrality strategy to form an intuitive and easy-to-use complete spatial point cloud model. The dynamic configuration of carbon sources, carbon sinks and monitors in the strategy is saved in a file, and finally the complete point cloud model and dynamic configuration file are output as the processing results.

2. The small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference according to claim 1, characterized in that, The feature parsing module imports and recognizes the user-input point cloud model and parameter configuration file, extracts carbon spatial features from the environmental point cloud model and carbon source configuration as valid input for subsequent modules. This feature parsing module includes: a point cloud model preprocessing unit, an environmental feature extraction unit, and a carbon source parameter extraction unit. Specifically: the point cloud model preprocessing unit preprocesses the user-input point cloud model, removing noise caused by device precision and object material through data denoising; it also filters point cloud data while maintaining accuracy through data simplification, extracting effective model information; the environmental feature extraction unit parses the preprocessed point cloud model, extracting the model's external boundaries and the boundaries of internal obstacles and carbon sources, identifying the relative positions of ventilation openings and obstacles, and forming key-value pairs; the carbon source parameter extraction unit takes the carbon source parameter file as input, extracts the attribute configuration related to the carbon source through scene keyword retrieval in text analysis, constructs multiple sets of key-value pairs with the carbon source as the object; and compares and fuses the relative values ​​of environmental features with the absolute values ​​of carbon source parameters to obtain an absolute representation of spatial features, ultimately forming a serialized carbon spatial feature representation.

3. The small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference according to claim 1, characterized in that, The digital carbon space initialization module initializes the digital carbon space based on carbon space characteristics, forming a basic spatial carbon neutrality strategy as input for subsequent neural reinforcement learning model inference optimization. This module includes a carbon space feature conversion unit, a carbon sink layout generation unit, and a monitor placement unit. Specifically, the carbon space feature conversion unit normalizes spatial information, including absolute size, during digital carbon space initialization, keeping it numerically within the range of [0,1], thus initially initializing the digital carbon space. The carbon sink layout generation unit sets the carbon sink layout, calculating the maximum total carbon emission rate of the carbon source at all times based on the carbon source spatiotemporal sequence. Based on this, a set of spatiotemporal sequences of carbon sinks are randomly initialized, where the total carbon absorption rate of the carbon sink at any given time is... This ensures that the carbon sink can absorb the total carbon emissions from the carbon source; the monitoring unit places a monitor at each ventilation opening and between each pair of carbon sources and carbon sinks that are close to each other to avoid excessively high local carbon concentrations indoors.

4. The small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference according to claim 1, characterized in that, The carbon neutrality strategy optimization module adjusts the layout of carbon sources, carbon sinks, and monitors by calling the carbon neutrality strategy optimization module, evaluating the results of each adjustment, and then updating the layout. This process iterates until the output value of the carbon neutrality strategy optimization module stabilizes. This module incorporates a deep reinforcement learning model, using a small-space carbon neutrality strategy as input, a carbon concentration prediction model as the interaction environment, and net-zero carbon emissions from the space as the core metric for training. The inference process uses the unoptimized carbon neutrality strategy as input and the optimized carbon neutrality strategy as output. The carbon neutrality strategy optimization module, which determines the quality of each state based on the level of reward, includes a spatial feature extraction unit, a deep convolutional network computation unit, and an action selection unit. Specifically: the spatial feature extraction unit constructs a high-dimensional input vector based on the input carbon neutrality strategy, extracts spatial features using principal component analysis, and obtains a low-dimensional abstract vector containing rich features; the deep convolutional network computation unit uses this low-dimensional vector as input to calculate and output the target value Q corresponding to the state; and the action selection unit selects the optimal next state based on the target value Q and the corresponding reward and outputs it. The aforementioned reward Where: Co is the total cost of the carbon sink device, Rt is the total energy consumption of the carbon sink and monitors, Ct is the total carbon emissions emitted to the outside world, and M represents the changes in the monitoring values ​​of all monitors. , , and The specific weights for each component are selected based on the specific circumstances. Lower total cost (Co), total energy consumption (Rt), and total external carbon emissions (Ct) are better, while the monitored values ​​should be as close as possible to... The better, as too high or too low a concentration will result in a lower reward, among which: ; Changes in the monitored values ​​of all the monitors mentioned Where: T is the number of time intervals, and k is the number of monitors. This represents the average atmospheric carbon dioxide concentration. Let be the monitoring value of the i-th monitor at time t; The total carbon emissions from the window to the outside Where: T is the number of time intervals, and q is the number of windows. Let be the carbon concentration in the i-th window at time t, and let the time interval be fixed. Regardless of how the carbon neutrality strategy sets up the monitor, each prediction requires monitoring the carbon concentration values ​​for each window.

5. The small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference according to claim 1, characterized in that, The carbon neutrality strategy optimization module performs experience replay before training the model and randomly selects states. That is, taking the carbon neutrality strategy as input, calculating each action. That is, the reward corresponding to the strategy adjustment. And the new status This corresponds to the new carbon neutrality strategy, recording the quadruple. To form an experience pool, we repeatedly use these experiences during the training process. Specifically, we set a batch size (batch_size). Within each batch, we use the target network to calculate the evaluation value of the reward, without updating its parameters; only the parameters of the evaluation network are updated. After each batch, we update the parameters of the target network once. When the prediction value given by the evaluation network in each iteration... The target value of the target network is The corresponding loss function Where: r is the reward for the current action, The discount rate decays over time. To evaluate the network parameters, Let E[x] be the parameter of the target network, and E[x] be the expectation of the random variable x.

6. The small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference according to claim 5, characterized in that, The parameters of the target network are updated in any of the following ways: a) Directly assign the parameters of the evaluation network to the target network to achieve a hard update; b) By introducing a learning rate The target network is then assigned a weighted average of the old target network parameters and the new evaluation network parameters to achieve a soft update. ,in: Select the soft update method to update. When the loss function value stabilizes, the model converges. In subsequent inference, only the target network is used to calculate the optimal action a under a given state s, that is, to generate the optimal carbon neutralization strategy.

7. The small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference according to claim 1, characterized in that, The carbon neutrality strategy export module exports the optimized spatial carbon neutrality strategy and generates a corresponding configuration file. This module includes: a parameter configuration export unit, an environment configuration export unit, a point cloud model export unit, and a point cloud model registration unit. Specifically: the point cloud model registration unit registers the carbon sink point cloud model with the environment point cloud model based on the quantity and location information of carbon sinks in the optimized carbon neutrality strategy, forming a complete net-zero carbon spatial point cloud model; the point cloud model export unit exports the complete point cloud model for subsequent modeling and analysis; the environment configuration export unit records the spatial shape and obstacle location information, forming an environment configuration file; and the parameter configuration export unit records the changes in the location and absorption rate of carbon sources and carbon sinks over time, forming carbon source configuration files and carbon sink configuration files, and ultimately creating a complete spatial carbon neutrality strategy output file.

8. The small-scale spatial carbon neutralization strategy generation system based on deep reinforcement learning inference according to claim 1, characterized in that, The model training support module enhances the spatial data of the spatial case library, generating a training set for the carbon concentration prediction model. This module includes: a case data preprocessing unit, a digital carbon space construction unit, a digital carbon space operation unit, and a spatiotemporal sequence post-processing unit. Specifically, the case data preprocessing unit transforms the spatial location information of the data into corresponding locations that can be recognized and used by the fluid simulation engine through spatial location transformation. Then, it performs multi-sequence time alignment operations and obtains the location and velocity of the carbon source and carbon sink at each moment by linear interpolating the spatiotemporal sequences of carbon sources and sinks, ultimately yielding the preprocessed case data. The digital carbon space construction unit identifies the configuration of the digital carbon space by importing preprocessed case data, and then initializes the fluid simulation engine to obtain the constructed digital carbon space. The digital carbon space operation unit simulates the diffusion process of carbon dioxide gas in a given space by running the digital carbon space. By randomly selecting multiple monitoring points, it records the gas concentration reading sequence. The resulting carbon concentration monitoring sequence reflects the overall gas state of the digital carbon space over time. The spatiotemporal sequence post-processing unit generates the model input spatiotemporal sequence after transforming the carbon concentration monitoring sequence data and aligning the time series data.

9. The small-scale spatial carbon neutrality strategy generation system based on deep reinforcement learning inference according to claim 1, characterized in that, The carbon concentration prediction module predicts the carbon concentration at any given time and space. That is, by simulating the carbon concentration changes in the digital carbon space corresponding to the carbon neutrality strategy, it obtains quantitative indicators that reflect the quality of the strategy and achieves the purpose of evaluating the strategy. The training process of the carbon concentration prediction model takes the training set as input, first extracts spatial features through a fully connected network, and then further inputs it into a temporal feature extraction network to extract temporal features, thereby obtaining the carbon concentration prediction value of the monitor at subsequent time points; The reasoning process takes a given carbon neutralization strategy as input and a score for that strategy as output. The score reflects the effectiveness of the strategy, and specifically includes: 1) Normalize the input information. When the longest side in the space is h, normalize the values ​​x related to position, velocity, and absolute size of the space to: To avoid the influence of values ​​that are too large or too small; regarding the relationship between the carbon source and the monitor, let the current time be t, and the position of carbon source 1 be... carbon emission direction ,rate Position of monitor m Current monitoring value The directional relationship between the carbon source and the monitor The direction vector can be obtained by combining the carbon emission direction. The carbon source vector can be obtained by combining the results. ; 2) Regarding the relationship between carbon sinks and monitors, let the current time be t, and the location of carbon sink 2 be... ,rate Position of monitor m Current monitoring value The relationship between carbon sequestration and the direction of the monitor. Since the carbon sink absorption range is spherical and the direction points towards the center of the sphere, the direction vector can be obtained. The carbon sink vector can be obtained by combining the results. ; 3) Abstract the carbon source and carbon sink as cuboids. For this relationship between the obstacle and the monitor, let the current time be t, and the obstacle position... Base length w, height h, monitor position m Current monitoring value The directional relationship between the obstacle and the monitor , direction vector The obstacle vector can be obtained by combining the results. ; 4) To , and Spatial feature extraction was performed using different fully connected networks, resulting in a unified 1*C dimensional vector; For example, the fully connected network structure is as follows: It consists of two fully connected layers: a hidden layer and an output layer. The hidden layer produces a 1*H dimensional vector, and a linear rectified activation function is applied before the output layer. Activate, all Directly adding them yields a 1*C dimensional vector, representing the influence of all carbon sources in space on the monitor m at time t; for and The processing is the same as above, thus obtaining three 1*C dimensional vectors, which are concatenated into a 1*3C dimensional vector, serving as the influence of space on monitor m at time t, and also as the input of the temporal feature extraction network at time t; 5) A temporal feature extraction network consisting of an input layer, hidden layers, and an output layer extracts the temporal features of the data at time t-1 and outputs the model's predicted carbon concentration for monitor m at time t. The difference is that the hidden layer incorporates the previous hidden layer vector into the current calculation. The specific calculation of the temporal feature extraction network is as follows: , ,in: The input at time t, The result of the hidden layer at time t. Let V be the output of the output layer at time t; V is the weight matrix of the output layer, and g is the activation function; due to solving... The hidden layer takes into account the result of the previous time step, so it is also a recurrent layer, and U is the input. The weight matrix, where W is the value at the previous time step. The weight matrix used as input at this moment, f, is the activation function; carbon concentration prediction is a regression problem, therefore the mean squared loss function MSELoss() is used as the loss function for model training, specifically: Where: X is the actual carbon concentration sequence, Y is the model-predicted sequence, and n is the sequence length. and These are the actual value and the predicted value at time t, respectively.