Indoor environment self-learning regulation system and method based on multi-sensor data fusion

CN122690931APending Publication Date: 2026-09-04GUANGZHOU SAIKE AUTOMATION CONTROL EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610708382.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

然而,这类方法难以应对建筑环境固有的多变量耦合、大惯性、时变性和不确定性等复杂特征

Benefits of technology

[0014] The aforementioned self-learning control method, system, computer equipment, and storage medium for indoor environments based on multi-sensor data fusion divide the target indoor space into multiple control sub-zones and construct graph structure data representing spatial topological relationships. This enables graph convolutional neural networks to accurately extract spatial correlation features between sub-zones, solving the technical problem that traditional methods cannot effectively handle thermal coupling between regions. By augmenting the initial node embedding features with graph data and constructing positive and negative sample pairs, the graph convolutional neural network is optimized using a contrastive loss function, improving the robustness and discriminativeness of feature representation and making subsequent control decisions more accurate. By inputting the optimized node embedding features into the policy network to generate collaborative control parameters, unified control of multiple end devices is achieved. By acquiring the actual environmental response data and total system energy consumption data after control, a multi-objective reward function is constructed. The policy network is updated using a reinforcement learning algorithm with the goal of maximizing the cumulative reward value, achieving adaptive optimization of the control strategy and effectively reducing system energy consumption while meeting comfort requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122690931A_ABST
    Figure CN122690931A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent environment control, and discloses an indoor environment self-learning regulation and control method and system based on multi-sensor data fusion, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-sensor data to generate state vectors of each control sub-area; constructing graph structure data; inputting the state vectors and the graph structure data into a graph convolutional neural network to extract initial node embedding features; performing graph data augmentation on the initial node embedding features and constructing positive and negative sample pairs, optimizing the graph convolutional neural network by using a contrast loss function, and obtaining optimized node embedding features; inputting the optimized node embedding features into a strategy network to generate collaborative control parameters and issuing the parameters for execution; acquiring actual environment response data and system total energy consumption data after regulation and control, constructing a multi-objective reward function and calculating a reward value; and updating network parameters of the strategy network by using a reinforcement learning algorithm. The application realizes multi-area collaboration and self-learning optimization control of an indoor environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent environmental control technology, specifically to an indoor environment self-learning control method, system, computer equipment, and storage medium based on multi-sensor data fusion. Background Technology

[0002] With the development of building intelligence technology, indoor environmental control systems are playing an increasingly important role in improving comfort and reducing energy consumption. Traditional indoor environmental control methods mainly employ logic control based on fixed rules or classic PID control algorithms, starting and stopping equipment through preset temperature thresholds or schedules. However, these methods struggle to cope with the complex characteristics inherent in the building environment, such as multivariate coupling, large inertia, time-varying nature, and uncertainty. For example, different areas have thermophysical relationships; temperature regulation in one area often triggers temperature fluctuations in adjacent areas. Traditional control strategies lack the dynamic perception capability of such spatial coupling relationships, resulting in poor control performance and high energy consumption. Furthermore, although model predictive control addresses these issues to some extent through feedforward optimization, its performance is highly dependent on accurate physical models, making widespread deployment in practical engineering applications difficult. Summary of the Invention

[0003] Therefore, it is necessary to provide an indoor environment self-learning control method, system, computer equipment, and storage medium based on multi-sensor data fusion that can achieve adaptive optimization control based on multi-regional collaborative control, addressing the aforementioned technical problems.

[0004] Firstly, a self-learning control method for indoor environment based on multi-sensor data fusion is provided, the method comprising: Acquire multi-sensor data of the target indoor space and generate state vectors for each of the control sub-zones; the target indoor space is pre-divided into multiple control sub-zones. A graph structure data is constructed using the control sub-regions as nodes and the degree of thermophysical correlation between the sub-regions as edge weights; The state vector and the graph structure data are input into a graph convolutional neural network to extract the initial node embedding features of each control sub-region; The initial node embedding features are augmented with graph data to generate at least two augmented views. The graph convolutional neural network is optimized using a contrastive loss function, with features of the same control subregion under different augmented views as positive sample pairs and features of different control subregions under the same augmented view as negative sample pairs. The optimized node embedding features are obtained. The optimized node embedding features are input into the policy network to generate the cooperative control parameters for each of the control sub-regions; The coordinated control parameters are sent to the end device controller to drive the end device to perform coordinated control. Acquire the actual environmental response data and total system energy consumption data after regulation, and construct a multi-objective reward function based on the deviation between the actual environmental response data and the preset comfort threshold and the total system energy consumption data to calculate the reward value of the current regulation action; with the goal of maximizing the cumulative reward value, use a reinforcement learning algorithm to update the network parameters of the policy network.

[0005] In one embodiment, optimizing the graph convolutional neural network using a contrastive loss function to obtain optimized node embedding features includes: The two feature vectors in the positive sample pair are mapped to the low-dimensional feature space respectively to obtain the first positive sample mapping vector and the second positive sample mapping vector. Each feature vector in the negative sample pair is mapped to the same low-dimensional feature space to obtain multiple negative sample mapping vectors; The cosine similarity between the first positive sample mapping vector and the second positive sample mapping vector is calculated as the positive sample similarity, and the cosine similarity between the first positive sample mapping vector and each of the negative sample mapping vectors is calculated as the negative sample similarity. The contrast loss value is calculated with the objective of maximizing the ratio of the positive sample similarity to the sum of the positive sample similarity and all negative sample similarities. The network parameters of the graph convolutional neural network are updated by backpropagation based on the contrastive loss value. The state vector and the graph structure data are input into the updated graph convolutional neural network for forward propagation to obtain the optimized node embedding features.

[0006] In one embodiment, prior to graph data augmentation of the initial node embedding features, the method further includes: The real-time heat transfer coefficient between adjacent sub-regions is calculated based on the real-time temperature data of each control sub-region. The weight values ​​of the corresponding edges in the graph structure data are dynamically adjusted according to the real-time heat transfer coefficient to update the graph structure data; The updated graph structure data is input into the graph convolutional neural network to re-extract the initial node embedding features of each of the control sub-regions, wherein the re-extracted initial node embedding features are used to augment the graph data.

[0007] In one embodiment, constructing the multi-objective reward function includes: Construct a comfort sub-function whose output value is negatively correlated with the absolute value of the deviation between the actual environmental response data and the preset comfort threshold; Construct an energy-saving sub-function whose output value is negatively correlated with the total energy consumption data of the system. Acquire user-initiated intervention data, which includes records of user corrections to the collaborative control parameters; An intervention penalty subfunction is constructed, the output value of which is negatively correlated with the frequency of the user's active intervention data, and a differentiated penalty weight is set according to the intervention type. The penalty weight includes a first penalty weight corresponding to a set value correction type intervention and a second penalty weight corresponding to a start-stop control type intervention. The second penalty weight is higher than the first penalty weight. The output values ​​of the comfort subfunction, the energy-saving subfunction, and the intervention penalty subfunction are weighted and summed to obtain the output value of the multi-objective reward function.

[0008] In one embodiment, the construction of the intervention penalty subfunction includes: Based on the correction records of the user-initiated intervention data, the intervention type is identified as either the set value correction type intervention or the start / stop control type intervention; Statistical analysis of the frequency of each intervention type within a pre-set period; Based on the frequency and the corresponding penalty weight, an intervention penalty coefficient is calculated, wherein the intervention penalty coefficient is positively correlated with the frequency and the penalty weight. The negative value of the intervention penalty coefficient is used as the output value of the intervention penalty sub-function.

[0009] In one embodiment, the step of inputting the optimized node embedding features into the policy network to generate cooperative control parameters for each of the control sub-regions includes: A hierarchical reinforcement learning framework is constructed, which includes a policy network at the upper layer and multiple execution networks at the lower layer. The policy network outputs the target environment parameters of each control sub-region based on the optimized node embedding features. Each execution network corresponds to a control sub-region. During the current control cycle, it receives the target environment parameters output by the policy network and calculates the corresponding deviation correction amount based on the deviation between the actual environment parameters collected at the end of the previous control cycle and the target environment parameters. The deviation correction amount is superimposed with the target environmental parameters to obtain the cooperative control parameters for the current control cycle.

[0010] In one embodiment, the method further includes: In response to the application request of the new target indoor space, an initial graph convolutional neural network and an initial policy network are constructed using the network parameters of the graph convolutional neural network and policy network of the target indoor space as initial parameters; Obtain the graph structure data and state vector of the new target indoor space; The state vector and graph structure data of the target indoor space are input into the initial graph convolutional neural network to obtain the source space node embedding features; The state vector and graph structure data of the new target indoor space are input into the initial graph convolutional neural network to obtain the new space node embedding features; The source spatial node embedding features and the new spatial node embedding features are input into the domain discriminator. Through adversarial training, the initial graph convolutional neural network learns domain-invariant features. The domain discriminator is used to distinguish the source of its input. Using multi-sensor data from the new target indoor space, the initial policy network and the initial graph convolutional neural network with invariant learning domain features are fine-tuned to obtain a graph convolutional neural network and policy network adapted to the new target indoor space.

[0011] Secondly, an indoor environment self-learning control system based on multi-sensor data fusion is provided, the system comprising: The data acquisition module is used to acquire multi-sensor data of the target indoor space and generate state vectors for each control sub-zone; the target indoor space is pre-divided into multiple control sub-zones. The graph construction module is used to construct graph structure data with the control sub-regions as nodes and the degree of thermophysical correlation between the sub-regions as edge weights. The feature extraction module is used to input the state vector and the graph structure data into the graph convolutional neural network to extract the initial node embedding features of each of the control sub-regions; The contrastive learning optimization module is used to perform graph data augmentation on the initial node embedding features to generate at least two augmented views; and to optimize the graph convolutional neural network using a contrastive loss function, with features of the same control subregion under different augmented views as positive sample pairs and features of different control subregions under the same augmented view as negative sample pairs, to obtain optimized node embedding features. The policy network module is used to input the optimized node embedding features into the policy network to generate the cooperative control parameters for each of the control sub-regions; The execution control module is used to send the collaborative control parameters to the end device controller to drive the end device to perform collaborative regulation; The reinforcement learning optimization module is used to acquire the actual environmental response data and the total system energy consumption data after regulation. Based on the deviation between the actual environmental response data and the preset comfort threshold and the total system energy consumption data, a multi-objective reward function is constructed to calculate the reward value of the current regulation action. With the goal of maximizing the cumulative reward value, the network parameters of the policy network are updated using a reinforcement learning algorithm.

[0012] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory is communicatively connected to the processor, and the memory stores a computer program that can run on the processor, wherein when the processor executes the computer program, it implements the above-described self-learning control method for indoor environment based on multi-sensor data fusion.

[0013] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the above-described method for self-learning control of indoor environment based on multi-sensor data fusion.

[0014] The aforementioned self-learning control method, system, computer equipment, and storage medium for indoor environments based on multi-sensor data fusion divide the target indoor space into multiple control sub-zones and construct graph structure data representing spatial topological relationships. This enables graph convolutional neural networks to accurately extract spatial correlation features between sub-zones, solving the technical problem that traditional methods cannot effectively handle thermal coupling between regions. By augmenting the initial node embedding features with graph data and constructing positive and negative sample pairs, the graph convolutional neural network is optimized using a contrastive loss function, improving the robustness and discriminativeness of feature representation and making subsequent control decisions more accurate. By inputting the optimized node embedding features into the policy network to generate collaborative control parameters, unified control of multiple end devices is achieved. By acquiring the actual environmental response data and total system energy consumption data after control, a multi-objective reward function is constructed. The policy network is updated using a reinforcement learning algorithm with the goal of maximizing the cumulative reward value, achieving adaptive optimization of the control strategy and effectively reducing system energy consumption while meeting comfort requirements. Attached Figure Description

[0015] Figure 1 This is an application environment diagram of an indoor environment self-learning control method based on multi-sensor data fusion in one embodiment; Figure 2 This is a flowchart illustrating an indoor environment self-learning control method based on multi-sensor data fusion in one embodiment. Figure 3 This is a flowchart illustrating the process of obtaining optimized node embedding features in one embodiment; Figure 4 This is a flowchart illustrating the dynamic adjustment of edge weights in one embodiment; Figure 5 This is a schematic diagram illustrating the process of constructing a multi-objective reward function in one embodiment; Figure 6 This is a flowchart illustrating the process of constructing the intervention penalty subfunction in one embodiment; Figure 7 This is a flowchart illustrating the process of generating collaborative control parameters for each control sub-region in one embodiment. Figure 8 This is a schematic diagram of the process of learning to migrate to a new target indoor space in one embodiment; Figure 9 This is a structural block diagram of an indoor environment self-learning control system based on multi-sensor data fusion in one embodiment; Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0017] The indoor environment self-learning control method based on multi-sensor data fusion provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 and server 104 communicate via a network. Terminal 102 is deployed in the target indoor space to collect multi-sensor data and receive collaborative control parameters to drive the end device to perform regulation. Server 104 performs feature extraction, contrastive learning optimization, policy network decision-making, and reinforcement learning updates of a graph convolutional neural network based on the data uploaded by terminal 102, and sends the generated collaborative control parameters back to terminal 102. Terminal 102 can be, but is not limited to, various field controllers, smart gateways, embedded devices, etc., and server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0018] Firstly, traditional indoor environment control methods, such as logic control based on fixed rules or classic PID control algorithms, often struggle to effectively handle the thermal coupling between different areas and lack adaptive optimization capabilities. Therefore, in one embodiment, a self-learning indoor environment control method based on multi-sensor data fusion is provided, which is then applied to… Figure 1 Taking the terminal in the example of this, for example... Figure 2 As shown, the method includes the following steps: Step S1: Acquire multi-sensor data of the target indoor space and generate state vectors for each control sub-zone. The target indoor space is pre-divided into multiple control sub-zones.

[0019] Specifically, firstly, based on factors such as the architectural structure and functional zoning of the target indoor space, the entire space is divided into multiple control sub-zones with physical boundaries, such as separate offices, meeting rooms, or different zones within an open office area. Through a sensor network set up within each control sub-zone, multi-sensor data is collected in real time. This multi-sensor data includes, but is not limited to, temperature data, humidity data, CO2 concentration data, occupant presence data, and differential pressure data. For each control sub-zone, all the sensor data collected at the current moment is organized into a multi-dimensional vector, which is the state vector of that sub-zone. For example, the state vector of sub-zone i at time t can be represented as... ,in For temperature, For humidity, CO2 concentration This pertains to the presence of individuals. All data collection and use must be authorized by the user and comply with relevant privacy regulations.

[0020] Step S2: Construct graph structure data with control sub-regions as nodes and the degree of thermophysical correlation between sub-regions as edge weights.

[0021] Specifically, each control sub-region is abstracted as a node in the graph, and the edges between nodes represent the thermophysical relationships between sub-regions, such as shared walls, floors, ceilings, or airflow through door and window openings. The edge weights quantify the strength of these thermophysical relationships, and their values ​​can be determined based on building structural parameters (such as the heat transfer coefficient and area of ​​wall materials), spatial distance, or temperature correlations in historical operational data. For example, if two adjacent offices share a concrete wall, the edge weight between them can be set to a higher value; if two areas are far apart and not directly connected, the edge weight can be set to zero or a very small value. This constructs a graph structure data that can characterize the topological structure of the target interior space. ,in For a set of nodes, Let be the weighted set of edges.

[0022] Step S3: Input the state vector and graph structure data into the graph convolutional neural network to extract the initial node embedding features of each control sub-region.

[0023] Specifically, graph convolutional neural networks (GNNs) can effectively process graph-structured data in non-Euclidean space. In this step, the graph-structured data constructed in the previous step and the state vectors of each node are used as input to a pre-constructed GNN. This network, through convolution operations on the graph structure, aggregates information about each node itself and its neighboring nodes, thereby extracting initial node embedding features containing spatial association information. The specific computation method of graph convolution can be implemented using any existing graph convolution operator, such as spectral domain-based graph convolution or spatial domain-based graph convolution; this embodiment does not limit this. Through the GNN, each node i obtains its corresponding initial node embedding features. .

[0024] Step S4: Perform graph data augmentation on the initial node embedding features to generate at least two augmented views. Using features of the same control subregion under different augmented views as positive sample pairs and features of different control subregions under the same augmented view as negative sample pairs, optimize the graph convolutional neural network using a contrastive loss function to obtain the optimized node embedding features.

[0025] Specifically, the purpose of graph augmentation is to generate diverse sample views for subsequent contrastive learning, thereby improving the robustness of feature representations. Graph augmentation is performed on the initial node embedding features to generate at least two distinct augmented views. These augmented views can be obtained by randomly masking partial node features. Figure 1 , and the view obtained by randomly discarding some edges Figure 2 Both of these augmented views are derived from the original graph structure data; they represent different perspectives under the same environmental conditions.

[0026] The essence of optimized node embedding feature acquisition is contrastive learning, the goal of which is to teach the model to distinguish between similar and dissimilar features. For the same control sub-region (e.g., sub-region A), its features under two different augmented views (e.g., view...) Figure 1 Features Heshi Figure 2 Features These constitute a positive sample pair because they both essentially describe the state of sub-region A. However, for the same augmented view, the features of two different control sub-regions (e.g., sub-region A and sub-region B) (such as...) and A positive sample pair is formed by comparing positive and negative samples. The graph convolutional neural network is then optimized using a contrastive loss function (e.g., InfoNCE loss), which aims to maximize the similarity between positive sample pairs while minimizing the similarity between negative sample pairs. After contrastive learning optimization, the graph convolutional neural network can output more discriminative and robust optimized node embedding features.

[0027] Step S5: Input the optimized node embedding features into the policy network to generate the collaborative control parameters for each control sub-region.

[0028] Specifically, the policy network is a deep neural network whose input is the embedded features of each node, and whose output is the collaborative control parameters corresponding to each control sub-region. These collaborative control parameters include, but are not limited to, temperature setpoints, humidity setpoints, differential pressure setpoints, and fresh air volume setpoints. These parameters are the target values ​​issued to the specific execution devices.

[0029] Step S6: Send the collaborative control parameters to the terminal device controller to drive the terminal device to perform collaborative control.

[0030] Specifically, the generated collaborative control parameters are distributed via the network to the corresponding terminal device controllers in each control sub-area, such as fan coil unit controllers, fresh air valve controllers, and electric water valve controllers. Based on the received collaborative control parameters, these controllers drive the corresponding terminal devices (such as air conditioners, water pumps, valves, and fresh air units) to perform collaborative regulation, thereby changing the indoor environmental conditions.

[0031] Step S7: Obtain the actual environmental response data and total system energy consumption data after regulation. Construct a multi-objective reward function based on the deviation between the actual environmental response data and the preset comfort threshold, as well as the total system energy consumption data, and calculate the reward value of the current regulation action. With the objective of maximizing the cumulative reward value, update the network parameters of the policy network using a reinforcement learning algorithm.

[0032] Specifically, after one cycle of coordinated control, the actual environmental response data of each control sub-region is acquired again through the sensor network, such as the actual temperature and humidity values. Simultaneously, the total system energy consumption data of the entire system (referring to the equipment system) or major energy-consuming equipment (such as chillers, water pumps, and fans) during this control cycle is recorded. Then, a multi-objective reward function is constructed. This function comprehensively reflects the control effect. The reward value is typically positively correlated with comfort indicators and negatively correlated with energy consumption indicators. The specific construction method of the reward function can employ any existing multi-objective weighted method; this embodiment does not limit this approach. For example, it can be set... The comfort deviation can be the absolute value of the difference between the actual temperature and the set temperature. and These are the corresponding weight coefficients, and the total weight sum is 1.

[0033] The goal of reinforcement learning is to learn an optimal policy that maximizes the cumulative reward over long-term operation. The reward calculated in the previous step is used as a feedback signal, and reinforcement learning algorithms (such as proximal policy optimization or deep deterministic policy gradient algorithms) are employed to update the network parameters of the policy network. The policy network outputs cooperative control parameters based on the current environmental state (i.e., the optimized node embedding features), and the environment returns a new state and reward value. Through iterative processes, the policy network continuously optimizes its control policy, gradually achieving self-learning regulation of the indoor environment.

[0034] Based on the above, this embodiment divides the indoor space into multiple control sub-regions and constructs graph structure data, enabling the graph convolutional neural network to effectively extract spatial correlation features between regions, thus solving the technical problem that traditional methods cannot perceive regional coupling. By comparing and learning to optimize feature representations, the robustness of the model to sensor noise and missing data is enhanced. By using reinforcement learning to adaptively update the policy network, continuous optimization of the control strategy is achieved, ultimately effectively reducing system energy consumption while meeting comfort requirements.

[0035] When optimizing a graph convolutional neural network using a contrastive loss function, it is necessary to accurately calculate the loss value and use it to update the network parameters to obtain more discriminative feature representations, thereby improving model performance. To this end, in one embodiment, such as... Figure 3 As shown, the contrastive loss function is used to optimize the graph convolutional neural network, resulting in optimized node embedding features, including: Step S41: Map the two feature vectors in the positive sample pair to the low-dimensional feature space respectively to obtain the first positive sample mapping vector and the second positive sample mapping vector.

[0036] Specifically, to calculate the similarity described later, a mapping network (e.g., a multilayer perceptron) is typically used to project the high-dimensional node embedding features into a low-dimensional feature space (e.g., a 128-dimensional vector space), where similarity calculation is more efficient. For a positive sample pair consisting of two augmented views of subregion A, its feature vectors are... and Inputting each sample into this mapping network yields the first positive sample mapping vector. Second positive sample mapping vector .

[0037] Step S42: Map each feature vector in the negative sample pair to the same low-dimensional feature space to obtain multiple negative sample mapping vectors.

[0038] Similarly, for all other sub-regions (e.g., sub-regions B, C, D...) in the same augmented view (e.g., view ... Figure 1Multiple negative sample pairs are formed under the given conditions. The feature vectors of these negative samples are also input into the same mapping network to obtain the corresponding negative sample mapping vector. For example, for sub-region B, we get... .

[0039] Step S43: Calculate the cosine similarity between the first positive sample mapping vector and the second positive sample mapping vector as the positive sample similarity, and calculate the cosine similarity between the first positive sample mapping vector and each negative sample mapping vector as the negative sample similarity.

[0040] Specifically, cosine similarity measures how close two vectors are in a direction. Positive sample similarity. The calculation formula is: For each negative sample, such as subregion B, its negative sample similarity is... The calculation formula is: .

[0041] Step S44: Calculate the contrast loss value with the objective of maximizing the ratio of positive sample similarity to the sum of positive sample similarity and all negative sample similarity.

[0042] Specifically, this step aims to define a contrastive loss function that encourages the model to increase the similarity of positive sample pairs while decreasing the similarity of all negative sample pairs. A typical loss function is the InfoNCE loss, which can be expressed as: ,in It is a temperature hyperparameter used to control the degree of concentration of the distribution. This represents the total number of negative samples. As can be seen, this loss function aims to maximize the ratio of the similarity of positive samples to the sum of the similarities of all samples (including the positive samples themselves and all negative samples).

[0043] Step S45: Backpropagate based on the contrast loss value to update the network parameters of the graph convolutional neural network.

[0044] Specifically, after calculating the contrastive loss value, the gradient of the loss value is passed back to the graph convolutional neural network (GCNN) through the backpropagation algorithm, and its internal weight parameters are updated. This process enables the GCNN to generate feature representations that are closer in feature space for the same sub-region and farther apart in feature space for different sub-regions in subsequent forward propagation.

[0045] Step S46: Input the state vector and graph structure data into the updated graph convolutional neural network for forward propagation to obtain the optimized node embedding features.

[0046] Specifically, after updating the parameters of the graph convolutional neural network based on the contrastive loss value, the original (or new) state vector and graph structure data are then input into the parameter-updated graph convolutional neural network for a forward propagation calculation. The node embedding features output by this calculation are the optimized node embedding features, which have better discriminativeness and robustness after contrastive learning optimization.

[0047] Based on the above, this embodiment clarifies the calculation method of the contrastive loss function and how to use it to optimize the graph convolutional neural network, providing a specific technical path for obtaining high-quality feature representations. Therefore, the features extracted by the graph convolutional neural network can more clearly distinguish the states of different sub-regions in the feature space, thereby providing better input for the subsequent policy network and improving the accuracy of the final control decision.

[0048] Because the thermophysical relationships within an indoor space change dynamically over time—e.g., factors like solar radiation and equipment heat dissipation cause localized temperature variations—static graphical data cannot accurately reflect real-time thermal coupling relationships, which affects control accuracy. Therefore, in one embodiment, such as... Figure 4 As shown, before graph data augmentation is performed on the initial node embedding features, the following steps are also included: Step S314: Calculate the real-time heat transfer coefficient between adjacent sub-regions based on the real-time temperature data of each control sub-region.

[0049] Specifically, firstly, real-time temperature data for each control sub-region is acquired via sensors. For any pair of adjacent control sub-regions i and j, their current temperatures are then analyzed. and Based on the physical properties between them (such as spacing, wall type, etc.), a real-time heat transfer coefficient can be calculated. The coefficient can be calculated using any existing thermodynamic model, such as a simplified formula based on Fourier's law of heat conduction, or an empirical model fitted based on historical data. The specific calculation method can be chosen according to actual engineering needs; this embodiment does not limit this, as long as it reflects the influence of real-time temperature difference on heat transfer intensity.

[0050] Step S324: Dynamically adjust the weight values ​​of the corresponding edges in the graph structure data according to the real-time heat transfer coefficient to update the graph structure data.

[0051] Specifically, the calculated real-time heat transfer coefficient As a new dynamic basis, update the original weights of the edges connecting nodes i and j in the graph structure data. The update strategy can be direct replacement, such as adding new edge weights. Alternatively, it can be a weighted fusion with the original weights, for example... ,in It is a balancing factor. In this way, graph-structured data... Dynamically updated to This allows it to more accurately reflect the actual thermophysical correlation state at the current moment. When the real-time temperature difference between any control sub-region and its adjacent sub-region exceeds a preset temperature difference threshold, the weight value of the corresponding edge is increased. This is a more direct and enhanced dynamic adjustment strategy. For example, when the preset temperature difference threshold is 2℃, if the sensor detects a real-time temperature difference between sub-region A and sub-region B... If the temperature exceeds 2°C, it is determined that there is a strong demand for heat exchange between the two. Therefore, when updating the graph structure data, the weight values ​​of the corresponding edges are significantly increased, for example, by doubling the original weights or by a preset upper limit value, in order to strengthen the mutual influence between the two nodes during the feature aggregation process of the graph convolutional neural network.

[0052] Step S334: Input the updated graph structure data into the graph convolutional neural network to re-extract the initial node embedding features of each control sub-region, wherein the re-extracted initial node embedding features are used for graph data augmentation.

[0053] Specifically, using the updated graph structure data Replace the original The data is then fed back into the graph convolutional neural network to re-extract features from the state vectors of each control sub-region. The resulting re-extracted initial node embedding features incorporate the latest real-time thermal coupling information. Subsequent steps such as graph data augmentation and contrastive learning will be based on these features, which better reflect the current physical reality.

[0054] Based on the above, this embodiment dynamically updates the edge weights of the graph structure data, enabling the graph convolutional neural network to perceive changes in spatial thermal coupling in real time, thereby extracting node embedding features that better reflect the current physical reality. This further enhances the model's adaptability to dynamic environmental changes, allowing subsequent contrastive learning and reinforcement learning to make decisions based on more accurate environmental representations, ultimately improving control accuracy and response speed.

[0055] Traditional multi-objective reward functions only consider comfort and energy consumption, neglecting personalized preferences reflected by user intervention, which may lead to control strategies that do not meet user expectations. Therefore, in one embodiment, such as... Figure 5 As shown, a multi-objective reward function is constructed, including: Step S71: Construct a comfort sub-function whose output value is negatively correlated with the absolute value of the deviation between the actual environmental response data and the preset comfort threshold.

[0056] Specifically, the comfort subfunction This is used to quantify the effectiveness of control in meeting user comfort requirements. Its core idea is based on actual environmental parameters (such as actual temperature). The closer it is to the preset comfort threshold (such as the set temperature) The higher the comfort level, the greater the reward value should be. Therefore, Defined as the absolute value of the deviation Inversely proportional. A simple linear implementation could be... ,in It is a positive proportionality coefficient, ensuring that the greater the deviation, the greater the negative reward value (i.e., penalty). For other environmental parameters such as humidity and CO2 concentration, a similar approach can be used to construct corresponding comfort indices, and then weighted and summed to ultimately form a total comfort sub-function.

[0057] Step S72: Construct an energy-saving sub-function whose output value is negatively correlated with the total energy consumption data of the system.

[0058] Specifically, the energy-saving sub-function Used to evaluate the energy-saving performance of control actions. Its core idea is the energy consumed by the system. The less energy is needed, the better the energy-saving effect, and the greater the reward value should be. Therefore, Defined as relative to the total energy consumption data of the system Inversely proportional. A simple linear implementation could be... ,in It is a positive proportionality coefficient, ensuring that the greater the energy consumption, the greater the negative reward (i.e., penalty).

[0059] Step S73: Obtain user-initiated intervention data, which includes records of user corrections to collaborative control parameters.

[0060] Specifically, after the automatically generated collaborative control parameters are issued and executed, users may manually modify the parameters through a human-machine interface (such as a control panel or mobile app) due to discomfort or personal habits. For example, they may increase the temperature setpoint or turn off a fan. These user-initiated intervention data are acquired and recorded, including the intervention time, the control sub-area involved, the type of intervention (whether it modified the setpoint or directly started / stopped the equipment), and the corrected parameter values.

[0061] Step S74: Construct an intervention penalty sub-function whose output value is negatively correlated with the frequency of user active intervention data, and set differentiated penalty weights according to the intervention type. The penalty weights include a first penalty weight corresponding to set value correction intervention and a second penalty weight corresponding to start-stop control intervention. The second penalty weight is higher than the first penalty weight.

[0062] Specifically, user intervention usually means that the current automated control strategy has failed to fully meet the user's expectations. Therefore, an intervention penalty subfunction needs to be constructed. This approach uses negative incentives to address such interventions. The core idea is that the higher the frequency of user intervention, the less the control effect meets user expectations, and the more severe the punishment should be. Intervention frequency Negative correlation. Furthermore, different types of interventions reflect the degree of user dissatisfaction. For example, a user might only slightly adjust the temperature setpoint (setpoint correction intervention), perhaps indicating only a slight difference in preference; however, if the user directly shuts down a running device (start / stop control intervention), it suggests a serious conflict between automatic control and user expectations. Therefore, differentiated penalty weights should be assigned to different types of interventions: a smaller initial penalty weight should be assigned to setpoint correction interventions. A larger second penalty weight is set for start-stop control interventions. The intervention penalty subfunction can be expressed as: ,in and These are the frequencies of the two intervention types occurring within a preset period.

[0063] Step S75: The output values ​​of the comfort subfunction, energy saving subfunction, and intervention penalty subfunction are weighted and summed to obtain the output value of the multi-objective reward function.

[0064] Specifically, the final multi-objective reward function in this embodiment It is a combination of the above three sub-functions, which can be expressed as: ,in , and These are the respective weighting coefficients used to balance the three objectives of comfort, energy efficiency, and user preference. This overall reward value... It will be used in subsequent reinforcement learning processes to guide the policy network to optimize in a direction that simultaneously satisfies multiple objectives.

[0065] Based on the above, this embodiment constructs a multi-objective reward function that includes user intervention penalties, enabling the reinforcement learning process to not only focus on comfort and energy consumption but also proactively learn control strategies to avoid triggering user dissatisfaction. The intervention penalty sub-function sets differentiated weights for different intervention types, allowing the model to more precisely understand the severity of user interventions. This enables more effective behavior adjustment during optimization, ultimately generating control parameters that better reflect users' true preferences, thus improving the user acceptance and automation level of the control method.

[0066] To ensure that reward signals accurately reflect the level of user dissatisfaction, it is necessary to clarify how to calculate penalty values ​​based on intervention type and frequency. Therefore, in one embodiment, such as... Figure 6 As shown, the intervention penalty sub-function is constructed, including: Step S741: Based on the correction records of user-initiated intervention data, identify whether the intervention type is a setpoint correction intervention or a start / stop control intervention.

[0067] Specifically, firstly, the acquired user-initiated intervention data is analyzed. By analyzing the specific content in the correction records, for example, if the record shows that the user changed the "temperature setpoint" from 24℃ to 25℃, this intervention record is identified as a setpoint correction intervention; if the record shows that the user issued a "shut down" command to a fan or the entire air conditioning system, this intervention record is identified as a start / stop control intervention.

[0068] Step S742: Calculate the frequency of each intervention type within the preset period.

[0069] Specifically, a statistical period is set, such as the past hour or the past 24 hours. Within this preset period, the total number of times identified as a set-value correction intervention is counted. and the total number of interventions identified as start-stop control interventions. .

[0070] Step S743: Calculate the intervention penalty coefficient based on the frequency and the corresponding penalty weight. The intervention penalty coefficient is positively correlated with the frequency and the penalty weight.

[0071] Specifically, the intervention penalty coefficient This is used to quantify the total punitive force resulting from intervention behaviors. Its calculation follows a basic principle: the higher the frequency of a certain type of intervention, the greater the punitive weight of that type, and thus the greater the total intervention punitive coefficient. A direct calculation formula is... It can be clearly seen from the formula that... The value and and It is positively correlated with, and also with and They are positively correlated.

[0072] Step S744: Use the negative value of the intervention penalty coefficient as the output value of the intervention penalty sub-function.

[0073] Specifically, since the output value of the intervention penalty subfunction needs to be positively correlated with user satisfaction (i.e., the more intervention, the smaller the output value, indicating a greater penalty), and the intervention penalty coefficient... It is a positive measure of the degree of intervention. Therefore, The negative value is used as the intervention penalty subfunction. The output value, i.e. This ensures that as the frequency of interventions or the penalty weight increases, the output value of the intervention penalty subfunction becomes smaller (i.e., more negative), thus forming a stronger penalty term in the total reward function.

[0074] Based on the above, this embodiment provides clear and executable computational logic for constructing the intervention penalty sub-function, ensuring that the reward signal can accurately reflect the impact of user intervention, thereby guiding the policy network to effectively learn and avoid user-unsatisfactory control behaviors, and improving the personalized adaptability of the control method.

[0075] In high-dimensional continuous action spaces, end-to-end reinforcement learning suffers from training difficulties and slow convergence, necessitating a more efficient control architecture. Therefore, in one embodiment, such as... Figure 7 As shown, the optimized node embedding features are input into the policy network to generate cooperative control parameters for each control sub-region, including: Step S51: Construct a hierarchical reinforcement learning framework, which includes a policy network at the upper layer and multiple execution networks at the lower layer.

[0076] Specifically, to address the control challenges in high-dimensional continuous action spaces, this embodiment employs a hierarchical reinforcement learning approach. The upper-layer policy network acts as the "decision-maker," responsible for formulating macro-level goals based on global information. The lower-layer multiple execution networks act as "executors," each responsible for a control sub-region, generating the final execution instructions based on the goals issued by the upper layer and local real-time feedback.

[0077] Step S52: The policy network outputs the target environment parameters of each control sub-region based on the optimized node embedding features.

[0078] Specifically, the upper-layer policy network receives optimized node embedding features containing spatial correlation information from the graph convolutional neural network. Based on this global and local information, the policy network generates a target environmental parameter for each control sub-region. This target environmental parameter is a relatively stable, macroscopic control objective, such as "maintain the temperature of sub-region A at 24°C for the next 15 minutes."

[0079] Step S53: Each execution network corresponds to a control sub-region. During the current control cycle, it receives the target environment parameters output by the policy network and calculates the corresponding deviation correction amount based on the deviation between the actual environment parameters collected at the end of the previous control cycle and the target environment parameters.

[0080] Specifically, the lower-level execution network is a relatively simple control network. At the beginning of each short control cycle (e.g., 1 minute), it receives the current target environment parameters set for its corresponding sub-region from the upper-level policy network. Simultaneously, it acquires the actual environmental parameters of the sub-region collected by the sensors at the end of the previous control cycle. Then, the network is executed based on the deviation. A bias correction amount is calculated using a built-in proportional-integral-derivative controller or its learned network weights. The deviation correction can be calculated using any existing control algorithm, such as the PID control algorithm.

[0081] Step S54: Superimpose the deviation correction amount with the target environmental parameters to obtain the coordinated control parameters for the current control cycle.

[0082] Specifically, the calculated deviation correction amount With the target environment parameters given by the upper layer By superimposing these parameters, we obtain the final coordinated control parameters to be issued in the current control cycle. ,Right now .this It is sent to the end device controller to guide the device to operate accurately in the next short cycle.

[0083] Based on the above, the upper-level policy network can focus on long-term, macro-level policy optimization, while the lower-level execution network is responsible for handling short-term, localized precise tracking and disturbance suppression, thereby improving the control accuracy and stability of the entire controlled system (equipment system). Based on this approach, reinforcement learning can learn effective macro-level policies at a higher level, while utilizing the rapid response capability of the lower-level controller to ensure the real-time performance and accuracy of control, thus solving the problems of difficult training and slow convergence in end-to-end reinforcement learning for large-scale continuous control problems.

[0084] If a trained model needs to be deployed to a new indoor space, training it from scratch is costly and requires a large amount of data. Therefore, such as... Figure 8 As shown, in one embodiment, the method further includes: Step S101: In response to the application request of the new target indoor space, construct the initial graph convolutional neural network and the initial policy network using the network parameters of the graph convolutional neural network and the policy network of the target indoor space as the initial parameters.

[0085] Specifically, when it is necessary to deploy a control policy from a fully trained and running source space (i.e., the aforementioned "target indoor space") to a completely new target indoor space with scarce data, the first step is to copy all the network parameters of the pre-trained graph convolutional neural network and policy network from the source space. Using these parameters, an initial graph convolutional neural network and an initial policy network are constructed for the new target indoor space. This is equivalent to allowing the new network to inherit the general environmental feature extraction capabilities and basic control logic learned from the source space.

[0086] Step S102: Obtain the graph structure data and state vector of the new target indoor space.

[0087] Specifically, following the aforementioned corresponding method, a unique graph structure data is constructed for the new target indoor space, and its current multi-sensor data is acquired to generate state vectors for each control sub-area of ​​the new target indoor space.

[0088] Step S103: Input the state vector and graph structure data of the target indoor space into the initial graph convolutional neural network to obtain the source space node embedding features. Input the state vector and graph structure data of the new target indoor space into the initial graph convolutional neural network to obtain the new space node embedding features.

[0089] Specifically, the historical (or real-time) data of the source space and the newly acquired spatial data are respectively input into the initial graph convolutional neural network with the same parameters (but about to begin differentiation), resulting in two sets of node embedding features: labeled as source space node embedding features and new space node embedding features, respectively.

[0090] Step S104: Input the source spatial node embedding features and the new spatial node embedding features into the domain discriminator. Through adversarial training, the initial graph convolutional neural network learns domain-invariant features. The domain discriminator is used to distinguish the source of its input.

[0091] Specifically, a domain discriminator network is constructed, whose task is to determine whether a feature originates from the source space or the new target space based on the embedded features of the input nodes. Simultaneously, during training, the initial goal of the graph convolutional neural network is not only to extract features but also to confuse the domain discriminator, preventing it from accurately distinguishing the source of the features. Through this adversarial training, the graph convolutional neural network is forced to learn domain-invariant features that are universal in both spaces and independent of the specific layout of the environment and sensor configuration. This process can be iterated until the discriminator's accuracy drops to a level close to random guessing. The domain discriminator and adversarial training can be implemented using any existing generative adversarial network method; this embodiment does not limit this approach.

[0092] Step S105: Using multi-sensor data from the new target indoor space, fine-tune the initial policy network and the initial graph convolutional neural network after learning domain-invariant features to obtain a graph convolutional neural network and policy network adapted to the new target indoor space.

[0093] Specifically, after adversarial training, the initial graph convolutional neural network (GCNN) already possesses the ability to extract domain-invariant features. At this point, using a small subset of real multi-sensor data collected in the new target indoor space, the entire model (including the initial GCNN and initial policy network after learning domain-invariant features) is jointly fine-tuned. The purpose of this fine-tuning is to allow the model to further adapt to the subtle control logic and user preferences specific to the new space, building upon the domain-invariant features. After fine-tuning, a GCNN and policy network fully adapted to the new target indoor space are obtained.

[0094] Based on the above, this transfer learning method can reduce the amount of data and time required to train the model in a new space, enabling efficient reuse of existing knowledge. According to the solution in this embodiment, when faced with new buildings or renovated spaces, only a small amount of new data is needed to quickly deploy a high-performance control model, improving the practicality and scalability of the solution.

[0095] It should be understood that, although Figure 2-8 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2-8 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0096] Secondly, this embodiment provides an indoor environment self-learning control system based on multi-sensor data fusion, such as... Figure 9 As shown, it includes a data acquisition module, a graph construction module, a feature extraction module, a contrastive learning optimization module, a policy network module, an execution control module, and a reinforcement learning optimization module.

[0097] The data acquisition module is used to acquire multi-sensor data of the target indoor space and generate state vectors for each control sub-zone; the target indoor space is pre-divided into multiple control sub-zones.

[0098] The graph construction module is used to construct graph structure data using control sub-regions as nodes and the degree of thermophysical correlation between sub-regions as edge weights.

[0099] The feature extraction module is used to input the state vector and graph structure data into the graph convolutional neural network to extract the initial node embedding features of each control subregion.

[0100] The contrastive learning optimization module is used to augment the initial node embedding features with graph data, generating at least two augmented views. The contrastive loss function is used to optimize the graph convolutional neural network, with features of the same control subregion under different augmented views as positive sample pairs and features of different control subregions under the same augmented view as negative sample pairs, to obtain the optimized node embedding features.

[0101] The policy network module is used to input the optimized node embedding features into the policy network to generate the cooperative control parameters for each control sub-region.

[0102] The execution control module is used to send collaborative control parameters to the end device controller, driving the end device to perform collaborative regulation.

[0103] The reinforcement learning optimization module is used to acquire the actual environmental response data and total system energy consumption data after regulation. Based on the deviation between the actual environmental response data and the preset comfort threshold and the total system energy consumption data, a multi-objective reward function is constructed to calculate the reward value of the current regulation action. With the goal of maximizing the cumulative reward value, the network parameters of the policy network are updated using a reinforcement learning algorithm.

[0104] In one embodiment, the contrastive learning optimization module includes a contrastive learning unit.

[0105] Specifically, the contrastive learning unit is used to map the two feature vectors in a positive sample pair to a low-dimensional feature space, respectively, to obtain a first positive sample mapping vector and a second positive sample mapping vector. It is also used to map each feature vector in a negative sample pair to the same low-dimensional feature space, resulting in multiple negative sample mapping vectors. The unit calculates the cosine similarity between the first and second positive sample mapping vectors as the positive sample similarity, and calculates the cosine similarity between the first positive sample mapping vector and each negative sample mapping vector as the negative sample similarity. It calculates the contrastive loss value with the objective of maximizing the ratio of the positive sample similarity to the sum of the positive and negative sample similarities. It performs backpropagation based on the contrastive loss value to update the network parameters of the graph convolutional neural network. Finally, it inputs the state vector and graph structure data into the updated graph convolutional neural network for forward propagation to obtain optimized node embedding features.

[0106] In one embodiment, the contrastive learning optimization module is further configured to calculate the real-time heat transfer coefficient between adjacent sub-regions based on the real-time temperature data of each control sub-region before performing graph data augmentation on the initial node embedding features; dynamically adjust the weight values ​​of corresponding edges in the graph structure data according to the real-time heat transfer coefficients to update the graph structure data; and input the updated graph structure data into the graph convolutional neural network to re-extract the initial node embedding features of each control sub-region, wherein the re-extracted initial node embedding features are used for graph data augmentation.

[0107] In one embodiment, the reinforcement learning optimization module includes a reward function construction unit.

[0108] Specifically, the reward function construction unit is used to construct a comfort subfunction, whose output value is negatively correlated with the absolute value of the deviation between the actual environmental response data and the preset comfort threshold. It is also used to construct an energy-saving subfunction, whose output value is negatively correlated with the total system energy consumption data. Furthermore, it is used to acquire user-initiated intervention data, including records of user corrections to collaborative control parameters. Finally, it is used to construct an intervention penalty subfunction, whose output value is negatively correlated with the frequency of user-initiated intervention data, and differentiated penalty weights are set according to the intervention type. These penalty weights include a first penalty weight corresponding to setpoint correction interventions and a second penalty weight corresponding to start-stop control interventions, with the second penalty weight being higher than the first. Finally, it is used to perform a weighted summation of the output values ​​of the comfort subfunction, the energy-saving subfunction, and the intervention penalty subfunction, as the output value of the multi-objective reward function.

[0109] Furthermore, the reward function building unit includes an intervention penalty sub-function building unit.

[0110] Specifically, the intervention penalty sub-function construction unit is used to identify the intervention type as either a setpoint correction intervention or a start / stop control intervention based on the correction records of user-initiated intervention data. It is used to count the frequency of each intervention type within a preset period. It is used to calculate the intervention penalty coefficient based on the frequency and the corresponding penalty weight; the intervention penalty coefficient is positively correlated with both the frequency and the penalty weight. Finally, it is used to output the negative value of the intervention penalty coefficient as the intervention penalty sub-function's output value.

[0111] In one embodiment, the policy network module includes a control parameter generation unit.

[0112] Specifically, the control parameter generation unit is used to construct a hierarchical reinforcement learning framework, which includes a policy network at the upper layer and multiple execution networks at the lower layer. The policy network outputs target environment parameters for each control sub-region based on the optimized node embedding features. Each execution network corresponds to one control sub-region. Within the current control cycle, it receives the target environment parameters output by the policy network and calculates the corresponding deviation correction based on the deviation between the actual environment parameters collected at the end of the previous control cycle and the target environment parameters. Finally, it superimposes the deviation correction with the target environment parameters to obtain the collaborative control parameters for the current control cycle.

[0113] In one embodiment, the system also includes a transfer learning module.

[0114] Specifically, the transfer learning module is used to respond to application requests from new target indoor spaces. It constructs an initial graph convolutional neural network (GCNN) and an initial policy network using the network parameters of the target indoor space's graph convolutional neural network and policy network as initial parameters. It acquires the graph structure data and state vector of the new target indoor space. It inputs the state vector and graph structure data of the target indoor space into the initial GCNN to obtain source space node embedding features. It inputs the state vector and graph structure data of the new target indoor space into the initial GCNN to obtain new space node embedding features. It inputs the source space node embedding features and the new space node embedding features into a domain discriminator, and through adversarial training, enables the initial GCNN to learn domain-invariant features. The domain discriminator is used to distinguish the source of its input. Finally, it uses multi-sensor data from the new target indoor space to fine-tune the initial policy network and the initial GCNN after learning domain-invariant features, resulting in a GCNN and policy network adapted to the new target indoor space.

[0115] Specific limitations regarding the self-learning indoor environment control system based on multi-sensor data fusion can be found in the limitations of the self-learning indoor environment control method based on multi-sensor data fusion described above, and will not be repeated here. Each module in the aforementioned self-learning indoor environment control system based on multi-sensor data fusion can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0116] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores all data related to the indoor environment self-learning control system. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements an indoor environment self-learning control method based on multi-sensor data fusion.

[0117] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0118] Thirdly, a computer device is provided, including a memory and a processor. The memory is communicatively connected to the processor, and the memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the aforementioned self-learning control method for indoor environment based on multi-sensor data fusion. Furthermore, the specific limitations of this computer device in implementing the self-learning control method for indoor environment based on multi-sensor data fusion can be found in the limitations of the self-learning control method for indoor environment based on multi-sensor data fusion described above, and will not be repeated here.

[0119] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the aforementioned self-learning control method for indoor environment based on multi-sensor data fusion. Furthermore, the specific limitations of the computer-readable storage medium in implementing the self-learning control method for indoor environment based on multi-sensor data fusion can be found in the limitations of the self-learning control method for indoor environment based on multi-sensor data fusion described above, and will not be repeated here.

[0120] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0121] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0122] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A self-learning control method for indoor environment based on multi-sensor data fusion, characterized in that, The method includes: Acquire multi-sensor data of the target indoor space and generate state vectors for each control sub-zone; the target indoor space is pre-divided into multiple control sub-zones. A graph structure data is constructed using the control sub-regions as nodes and the degree of thermophysical correlation between the sub-regions as edge weights; The state vector and the graph structure data are input into a graph convolutional neural network to extract the initial node embedding features of each control sub-region; The initial node embedding features are augmented with graph data to generate at least two augmented views. The graph convolutional neural network is optimized using a contrastive loss function, with features of the same control subregion under different augmented views as positive sample pairs and features of different control subregions under the same augmented view as negative sample pairs. The optimized node embedding features are obtained. The optimized node embedding features are input into the policy network to generate the cooperative control parameters for each of the control sub-regions; The coordinated control parameters are sent to the end device controller to drive the end device to perform coordinated control. Acquire the actual environmental response data and total system energy consumption data after regulation, and construct a multi-objective reward function based on the deviation between the actual environmental response data and the preset comfort threshold and the total system energy consumption data to calculate the reward value of the current regulation action; with the goal of maximizing the cumulative reward value, use a reinforcement learning algorithm to update the network parameters of the policy network.

2. The method according to claim 1, characterized in that, The step of optimizing the graph convolutional neural network using a contrastive loss function to obtain optimized node embedding features includes: The two feature vectors in the positive sample pair are mapped to the low-dimensional feature space respectively to obtain the first positive sample mapping vector and the second positive sample mapping vector. Each feature vector in the negative sample pair is mapped to the same low-dimensional feature space to obtain multiple negative sample mapping vectors; The cosine similarity between the first positive sample mapping vector and the second positive sample mapping vector is calculated as the positive sample similarity, and the cosine similarity between the first positive sample mapping vector and each of the negative sample mapping vectors is calculated as the negative sample similarity. The contrast loss value is calculated with the objective of maximizing the ratio of the positive sample similarity to the sum of the positive sample similarity and all negative sample similarities. The network parameters of the graph convolutional neural network are updated by backpropagation based on the contrastive loss value. The state vector and the graph structure data are input into the updated graph convolutional neural network for forward propagation to obtain the optimized node embedding features.

3. The method according to claim 1, characterized in that, Before performing graph data augmentation on the initial node embedding features, the process also includes: The real-time heat transfer coefficient between adjacent sub-regions is calculated based on the real-time temperature data of each control sub-region. The weight values ​​of the corresponding edges in the graph structure data are dynamically adjusted according to the real-time heat transfer coefficient to update the graph structure data; The updated graph structure data is input into the graph convolutional neural network to re-extract the initial node embedding features of each of the control sub-regions, wherein the re-extracted initial node embedding features are used to augment the graph data.

4. The method according to claim 1, characterized in that, The construction of the multi-objective reward function includes: Construct a comfort sub-function whose output value is negatively correlated with the absolute value of the deviation between the actual environmental response data and the preset comfort threshold; Construct an energy-saving sub-function whose output value is negatively correlated with the total energy consumption data of the system. Acquire user-initiated intervention data, which includes records of user corrections to the collaborative control parameters; An intervention penalty subfunction is constructed, the output value of which is negatively correlated with the frequency of the user's active intervention data, and a differentiated penalty weight is set according to the intervention type. The penalty weight includes a first penalty weight corresponding to a set value correction type intervention and a second penalty weight corresponding to a start-stop control type intervention. The second penalty weight is higher than the first penalty weight. The output values ​​of the comfort subfunction, the energy-saving subfunction, and the intervention penalty subfunction are weighted and summed to obtain the output value of the multi-objective reward function.

5. The method according to claim 4, characterized in that, The construction of the intervention penalty sub-function includes: Based on the correction records of the user-initiated intervention data, the intervention type is identified as either the set value correction type intervention or the start / stop control type intervention; Statistical analysis of the frequency of each intervention type within a pre-set period; Based on the frequency and the corresponding penalty weight, an intervention penalty coefficient is calculated, wherein the intervention penalty coefficient is positively correlated with the frequency and the penalty weight. The negative value of the intervention penalty coefficient is used as the output value of the intervention penalty sub-function.

6. The method according to claim 1, characterized in that, The step of inputting the optimized node embedding features into the policy network to generate cooperative control parameters for each control sub-region includes: A hierarchical reinforcement learning framework is constructed, which includes a policy network at the upper layer and multiple execution networks at the lower layer. The policy network outputs the target environment parameters of each control sub-region based on the optimized node embedding features. Each execution network corresponds to a control sub-region. During the current control cycle, it receives the target environment parameters output by the policy network and calculates the corresponding deviation correction amount based on the deviation between the actual environment parameters collected at the end of the previous control cycle and the target environment parameters. The deviation correction amount is superimposed with the target environmental parameters to obtain the cooperative control parameters for the current control cycle.

7. The method according to any one of claims 1 to 6, characterized in that, The method also includes: In response to the application request of the new target indoor space, an initial graph convolutional neural network and an initial policy network are constructed using the network parameters of the graph convolutional neural network and policy network of the target indoor space as initial parameters; Obtain the graph structure data and state vector of the new target indoor space; The state vector and graph structure data of the target indoor space are input into the initial graph convolutional neural network to obtain the source space node embedding features; The state vector and graph structure data of the new target indoor space are input into the initial graph convolutional neural network to obtain the new space node embedding features; The source spatial node embedding features and the new spatial node embedding features are input into the domain discriminator. Through adversarial training, the initial graph convolutional neural network learns domain-invariant features. The domain discriminator is used to distinguish the source of its input. Using multi-sensor data from the new target indoor space, the initial policy network and the initial graph convolutional neural network with invariant learning domain features are fine-tuned to obtain a graph convolutional neural network and policy network adapted to the new target indoor space.

8. An indoor environment self-learning control system based on multi-sensor data fusion, characterized in that, The system includes: The data acquisition module is used to acquire multi-sensor data of the target indoor space and generate state vectors for each control sub-zone; the target indoor space is pre-divided into multiple control sub-zones. The graph construction module is used to construct graph structure data with the control sub-regions as nodes and the degree of thermophysical correlation between the sub-regions as edge weights. The feature extraction module is used to input the state vector and the graph structure data into the graph convolutional neural network to extract the initial node embedding features of each of the control sub-regions; The contrastive learning optimization module is used to perform graph data augmentation on the initial node embedding features to generate at least two augmented views; and to optimize the graph convolutional neural network using a contrastive loss function, with features of the same control subregion under different augmented views as positive sample pairs and features of different control subregions under the same augmented view as negative sample pairs, to obtain optimized node embedding features. The policy network module is used to input the optimized node embedding features into the policy network to generate the cooperative control parameters for each of the control sub-regions; The execution control module is used to send the collaborative control parameters to the end device controller to drive the end device to perform collaborative regulation; The reinforcement learning optimization module is used to acquire the actual environmental response data and the total system energy consumption data after regulation; construct a multi-objective reward function based on the deviation between the actual environmental response data and the preset comfort threshold and the total system energy consumption data, calculate the reward value of the current regulation action, and update the network parameters of the policy network using a reinforcement learning algorithm with the goal of maximizing the cumulative reward value.

9. A computer device comprising a memory and a processor, the memory being communicatively connected to the processor, and the memory storing a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the indoor environment self-learning control method based on multi-sensor data fusion as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the indoor environment self-learning control method based on multi-sensor data fusion as described in any one of claims 1 to 7.