Multi-objective reinforcement learning air conditioner control method and system based on forgetting model and storage medium

By using a multi-objective reinforcement learning method based on a forgetting model, the thermal environment state field is reconstructed and the perception decay mapping and cumulative effect superposition are performed. Combined with the state compression coding and policy optimization of the multi-objective reinforcement learning agent, control actions that take into account both user thermal comfort and system energy consumption are generated. This solves the problems of thermal environment perception disconnect and multi-objective optimization imbalance in air conditioning control, and achieves dynamic balance between thermal comfort and energy consumption.

CN122429464APending Publication Date: 2026-07-21SICHUAN HONGMEI INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN HONGMEI INTELLIGENT TECH CO LTD
Filing Date
2026-05-09
Publication Date
2026-07-21

Smart Images

  • Figure CN122429464A_ABST
    Figure CN122429464A_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of based on forgetting model's multi-objective reinforcement learning air conditioner control method, system and storage medium, method includes: original environment physical quantity measurement set is carried out time sequence synchronization and space field structure reconstruction, generates current thermal environment state field;According to somatic forgetfulness cumulative model, perception attenuation mapping and cumulative effect superposition are obtained thermal sensation perception cumulative state field;The state compression coding network of multi-objective reinforcement learning intelligent agent is started, and multi-level spatial feature abstraction and feature channel dimension reduction compression are carried out to generate compressed state representation vector;Based on compressed state representation vector, the multi-dimensional control action vector of heat comfort and energy consumption target is output;It is changed into air conditioner operating parameter variable quantity by executing mechanism action mapping layer conversion and executes update, and the attenuation rate constant of forgetting attenuation mapping function is adaptively corrected using the subjective thermal sensation feedback record collected after updating, effectively improve the individuality and dynamic balance ability of comfort and energy efficiency of air conditioner control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent control technology, specifically to a multi-objective reinforcement learning air conditioning control method, system, and storage medium based on a forgetting model. Background Technology

[0002] In the field of air conditioning control, feedback control methods based on environmental physical quantity measurements have been widely researched and applied. A common approach involves deploying multiple environmental measurement nodes, such as those for temperature and wind speed, within the controlled space. The physical quantity readings at each node are collected, and the current thermal environment distribution field is constructed using spatial interpolation or data fusion methods. For control decision-making, a preset temperature setpoint curve or a single-objective optimization controller is typically relied upon. The deviation between the desired environmental state and the current thermal environment distribution field is used as input, and algorithms such as proportional-integral-derivative control or model predictive control are used to calculate and execute adjustments to the air conditioning operating parameters.

[0003] However, the aforementioned existing technical solutions generally treat thermal environment control as an instantaneous physical quantity tracking problem, without modeling the memory decay and cumulative effects of users' perception of the thermal environment over continuous time. This results in control strategies failing to align with the actual dynamic process of subjective thermal perception. Furthermore, in terms of multi-objective optimization, it is difficult to achieve a dynamic balance between thermal comfort and energy consumption based on personalized user perception. Summary of the Invention

[0004] This application provides a multi-objective reinforcement learning-based air conditioning control method, system, and storage medium based on a forgetting model.

[0005] This application provides a multi-objective reinforcement learning-based air conditioning control method based on a forgetting model, including: The original environmental physical quantity measurement sets collected by multiple environmental measurement nodes distributed within the target space are processed by time synchronization and spatial field structure reconstruction to generate the current thermal environment state field based on a unified spatial coordinate system. Based on a preset somatosensory forgetting accumulation model, the current thermal environment state field is subjected to perceptual attenuation mapping and cumulative effect superposition processing to obtain the thermal perception accumulation state field. A state compression coding network is initiated for the multi-objective reinforcement learning agent to perform multi-level spatial feature abstraction and feature channel dimensionality reduction compression processing on the thermal sensing cumulative state field, generating a compressed state representation vector suitable for policy optimization. The multi-objective policy network of the multi-objective reinforcement learning agent performs forward inference based on the compressed state representation vector, and outputs a multi-dimensional control action vector that takes into account both the user thermal comfort optimization objective and the system energy consumption optimization objective. The multidimensional control action vector is converted into an operating parameter change that the air conditioning equipment can recognize through the actuator action mapping layer and the air conditioning operating status is updated. After the update, the subjective thermal feedback record of the current user is collected, and the decay rate constant of the forgetting decay mapping function in the preset somatic forgetting accumulation model is adaptively corrected using the subjective thermal feedback record.

[0006] This application also provides a multi-objective reinforcement learning air conditioning control system, including: processor; Storage device, on which computer programs are stored, When the computer program is executed by the processor, the processor implements any of the aforementioned multi-objective reinforcement learning air conditioning control methods based on the forgetting model.

[0007] This application embodiment also provides a readable storage medium storing a program or instructions, which, when executed by a processor, implement the steps of the multi-objective reinforcement learning air conditioning control method based on a forgetting model.

[0008] Therefore, the embodiments of this application have the following beneficial effects: By reconstructing the unified thermal environment state field of the target space and introducing a somatosensory forgetting accumulation model to perform perceptual attenuation mapping and cumulative effect superposition of thermal environment physical quantities, the embodiments of this application enable the control system to continuously simulate the user's subjective feelings about changes in the thermal environment and their memory attenuation characteristics, rather than only responding to instantaneous physical measurement values. Furthermore, the multi-objective reinforcement learning agent uses a state compression coding network to perform multi-level spatial feature abstraction and feature channel dimensionality reduction compression on the thermal perception accumulation state field, generating a highly condensed compressed state representation vector that retains key environmental context information, effectively reducing the state space complexity of policy optimization. Based on this compressed state representation vector, the multi-objective policy network simultaneously outputs a multi-dimensional control action vector that takes into account both the user's thermal comfort optimization objective and the system's energy consumption optimization objective, which is converted into changes in air conditioning operating parameters through the actuator action mapping layer to complete the control closed loop. Based on this, the decay rate constant of the forgetting decay mapping function is adaptively corrected by using the subjective thermal feedback records collected after the control update, so that the somatosensory forgetting accumulation model continuously approximates the actual perceived dynamic characteristics of individual users. This enables refined energy consumption management while ensuring thermal comfort, and solves the problems of disconnect between thermal environment control and user subjective feelings and imbalance of multi-objective optimization in existing technologies. Attached Figure Description

[0009] Figure 1 This is a flowchart of a multi-objective reinforcement learning air conditioning control method based on a forgetting model, provided as an embodiment of this application.

[0010] Figure 2This is a schematic diagram of the basic structure of a multi-objective reinforcement learning air conditioning control system provided in an embodiment of this application.

[0011] Figure 3 This is a functional block diagram of a multi-objective reinforcement learning air conditioning control device provided in an embodiment of this application. Detailed Implementation

[0012] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0013] See Figure 1 As shown, this figure is a flowchart of a multi-objective reinforcement learning air conditioning control method based on a forgetting model provided in an embodiment of this application. This method is executed by a multi-objective reinforcement learning air conditioning control system. Figure 1 As shown, the method includes steps 110-150.

[0014] This application provides a multi-objective reinforcement learning air conditioning control method based on a forgetting model. This method is applied to an air conditioning control system that includes multiple environmental measurement nodes, an air conditioning execution controller, a multi-objective reinforcement learning agent, and a user feedback acquisition terminal. This method continuously models the user's thermal perception by synchronizing the temporal and spatial reconstruction of the target space thermal environment and combining it with a somatosensory forgetting accumulation model. It uses a multi-objective reinforcement learning agent to optimize the compression state representation, outputs multi-dimensional control actions, and converts them into changes in air conditioning operating parameters for execution control. At the same time, it relies on the user's subjective thermal feedback to adaptively correct the forgetting model parameters, thus realizing a multi-objective optimization control closed loop that takes into account both user thermal comfort and system energy consumption.

[0015] Step 110: Perform time-series synchronization and spatial field structure reconstruction processing on the original environmental physical quantity measurement set collected by multiple environmental measurement nodes distributed in the target space to generate the current thermal environment state field based on a unified spatial coordinate system.

[0016] In this embodiment of the application, the target space can be an indoor space configured with multiple environmental measurement nodes. Each environmental measurement node includes a temperature sensing unit and a wind speed sensing unit, which can periodically collect environmental physical quantity measurements of its location.

[0017] The raw environmental physical quantity measurement set contains measurement data packets reported by each node within the same acquisition cycle. Each measurement data packet carries a node identifier, an acquisition timestamp, and a vector of environmental physical quantity measurement values, including temperature and wind speed readings. Due to slight drift in the local clocks of each node, the timestamps reported by different nodes in the same acquisition batch are not completely consistent. Therefore, it is necessary to perform time synchronization processing on these measurement data.

[0018] First, based on node identifiers, the original measurement set is split into independent time series for each node. A linear interpolation method is then used to resample the time series of each node on a unified time reference axis, ensuring that the measurements of all nodes are aligned to the same time scale point, forming a time-synchronized environmental physical quantity measurement set. Next, a pre-stored spatial location distribution map of the environmental measurement nodes is retrieved. This map defines the spatial grid node position coordinates corresponding to each node identifier using a unified spatial coordinate system. For example, the target space is discretized into a regular three-dimensional grid array, where each grid node coordinate consists of X-axis, Y-axis, and Z-axis indices. Each measurement value vector in the time-synchronized set is mapped to its corresponding grid node position according to its node identifier, forming a spatially discrete measurement point matrix. This spatially discrete measurement point matrix has measurements at only some grid node positions, with other grid node positions being empty.

[0019] To obtain a continuous state field covering all grid node locations in the target space, it is necessary to perform spatial gradient continuous filling of the spatial discrete measurement lattice based on a pre-stored spatial continuous reconstruction function. The spatial continuous reconstruction function can employ a method combining spatial distance-weighted interpolation and a physical diffusion model. Using grid nodes with measured values ​​as anchor points, interpolation weights are calculated along the X, Y, and Z axes based on the physical quantity gradients of adjacent anchor points, gradually filling in the temperature and wind speed components of missing grid nodes to generate the current thermal environment state field. Each spatial grid node location in the current thermal environment state field carries a thermal environment state vector, which contains at least temperature and wind speed components, each associated with specific grid node coordinates in a unified spatial coordinate system.

[0020] In one optional implementation, step 110 can be broken down into the following sub-steps: Step 1101: Receive the original environmental physical quantity measurement set collected by the multiple environmental measurement nodes, and extract the collection timestamp and node identifier attached to each environmental physical quantity measurement value.

[0021] In this embodiment, the data aggregation module of the air conditioning control system receives raw measurement data packets reported by various environmental measurement nodes within the target space at the end of each acquisition cycle. The raw environmental physical quantity measurement set consists of multiple data packets, each corresponding to a single measurement event of a node. The system traverses each data packet in the set, parsing out the node identifier field, acquisition timestamp field, and measurement value payload field. The measurement value payload field stores the temperature measurement value component and the wind speed measurement value component. The extracted data is temporarily organized into an associated structure, with each node identifier associated with its corresponding acquisition timestamp sequence and environmental physical quantity measurement value vector sequence.

[0022] Step 1102: Project all environmental physical quantity measurement values ​​in the original environmental physical quantity measurement set onto a unified time reference axis according to the acquisition timestamp, and perform timestamp interpolation alignment processing on the environmental physical quantity measurement values ​​with different node identifiers to generate a time-synchronized environmental physical quantity measurement set.

[0023] After extracting the acquisition timestamps of all measurements, the system identifies the earliest acquisition timestamp as the starting point of the unified time reference axis and the latest acquisition timestamp as the ending point. It then sets the standard time point sequence on the reference axis according to the desired time resolution of the control cycle. For each node identifier, its corresponding acquisition timestamp and measurement value vector sequence are treated as the original sampling point sequence, and linear interpolation is performed on the standard time point sequence.

[0024] The interpolation process is performed independently for each component in the measurement vector. For a given standard time point, the two closest original sampling points are selected, and the interpolation coefficient is calculated based on the ratio of the time difference between their acquisition timestamps and the standard time point. The temperature components of the two original sampling points are then weighted and averaged to obtain the temperature interpolation result for that standard time point. The wind speed component is interpolated in the same way. The interpolation results of all nodes at the same standard time point constitute a set of synchronous measurements for that standard time point, and all synchronous measurements at all standard time points constitute a time-synchronized environmental physical quantity measurement set.

[0025] Step 1103: Retrieve the preset spatial location distribution map of environmental measurement nodes, and map each environmental physical quantity measurement value in the time-synchronized environmental physical quantity measurement set to the spatial grid node position in the unified spatial coordinate system according to the corresponding node identifier, thereby creating a spatial discrete measurement point matrix.

[0026] A pre-defined spatial distribution map of environmental measurement nodes records the mapping relationship between each node identifier and the coordinates of spatial grid nodes in a unified spatial coordinate system. The system reads this distribution map and establishes a lookup index from node identifiers to grid node coordinates. It iterates through each standard time point in the time-synchronized environmental physical quantity measurement set, extracts the measurement value vector of each node at that time point, and fills it into the corresponding grid node coordinate position in the spatial discrete measurement point matrix using the lookup index. For any standard time point, the spatial discrete measurement point matrix is ​​a three-dimensional array. The position of each element in the array is determined by an X-axis index, a Y-axis index, and a Z-axis index. The element value is an environmental physical quantity measurement value vector. If a grid node coordinate does not have a corresponding node identifier, the element value remains empty.

[0027] Step 1104: Based on the spatial discrete measurement point array, the environmental physical quantities between adjacent spatial grid nodes are continuously filled with spatial gradients through a pre-stored spatial continuous reconstruction function to generate a current thermal environment state field covering all spatial grid node positions in the target space. Each spatial grid node position in the current thermal environment state field carries a thermal environment state vector based on a unified spatial coordinate system.

[0028] The spatial continuous reconstruction function first processes the spatial discrete measurement point matrix layer by layer according to the Z-axis index, with each layer corresponding to an XY plane slice. For each slice, anchor grid nodes with measurements and empty grid nodes are identified. For each empty grid node, the nearest anchor point is searched along the positive and negative X-axis, and along the positive and negative Y-axis, respectively, to obtain the measurement value vector of the nearest anchor point in these four directions. The spatial weight of each anchor point is determined based on the distance ratio between the empty grid node and the four anchor points, and the sum of the four spatial weights is 1. The temperature measurement components of the four anchor points are weighted and summed according to the spatial weights to serve as the temperature filling value for the empty grid node, and the wind speed measurement components are treated similarly.

[0029] After completing the intraplane interpolation of all XY plane slices, interlayer interpolation is then performed along the Z-axis. For positions where gaps still exist after the intraplane interpolation, linear interpolation is performed based on the thermal environment state vectors of the filled mesh nodes in the adjacent layers. After layer-by-layer intraplane interpolation and interlayer interpolation, all spatial mesh node positions in the target space carry complete thermal environment state vectors, forming the current thermal environment state field.

[0030] Step 120: Based on the preset somatosensory forgetting accumulation model, perform sensory attenuation mapping and cumulative effect superposition processing on the current thermal environment state field to obtain the thermal perception accumulation state field.

[0031] The pre-set somatosensory forgetting accumulation model is a mathematical model used to simulate the user's subjective perception process of changes in the thermal environment and its memory decay characteristics. The model internally stores the forgetting decay mapping function, the cumulative effect superposition function, and the thermal perception accumulation state field generated and saved in the previous control cycle.

[0032] The forgetting decay mapping function describes how the human body's perceptual memory of the thermal environment at previous moments gradually weakens over time, and its core parameter is the decay rate constant. The cumulative effect superposition function describes the interaction between the immediate thermal environment perception at the current moment and the decayed past perceptual memory, which is usually manifested as a nonlinear fusion mechanism.

[0033] In its implementation, step 120 is executed sequentially according to the following sub-steps: Step 1201: Retrieve the forgetting decay mapping function, the cumulative effect superposition function, and the thermal perception cumulative state field of the previous moment corresponding to the spatial coordinate range of the current thermal environment state field from the preset somatosensory forgetting accumulation model.

[0034] The system maintains a model parameter storage area, which persistently stores the type identifier of the forgetting decay mapping function and its current value of the decay rate constant, the type identifier of the cumulative effect superposition function and its built-in fusion weight coefficients, and the thermal perception cumulative state field written at the end of the previous control cycle. Each time this step is executed, these data and functions are loaded into memory from this storage area.

[0035] Step 1202: The current thermal environment state field is analyzed according to the spatial grid cells of a unified spatial coordinate system to obtain the spatial grid cell position identifier and the current thermal environment state vector corresponding to each spatial grid cell position.

[0036] The current thermal environment state field is itself a three-dimensional data field indexed by the coordinates of grid nodes. The system traverses each grid node in this data field, extracts its coordinate index as a spatial grid cell location identifier, and simultaneously extracts the thermal environment state vector stored at that coordinate location as the current thermal environment state vector. This results in a sequence containing pairs of location identifiers and state vectors.

[0037] Step 1203: Extract the previous sensing cumulative state vector corresponding to the position of each spatial grid cell from the previous thermal sensing cumulative state field according to the spatial grid cell position identifier.

[0038] The previous thermal sensing cumulative state field shares the exact same spatial coordinate system and grid division as the current thermal environment state field. The system traverses the location identifier of each spatial grid cell, using this identifier as an index to search for and retrieve the corresponding sensing cumulative state vector in the previous thermal sensing cumulative state field. The dimension of the previous sensing cumulative state vector is not necessarily the same as that of the current thermal environment state vector; it internally encodes historical thermal sensing cumulative information.

[0039] Step 1204: For each spatial grid cell location, input the previously perceived cumulative state vector and the decay rate constant of the forgetting decay mapping function set in the preset somatosensory forgetting cumulative model into the forgetting decay mapping function to generate the decayed perceived memory vector for that spatial grid cell location.

[0040] The forgetting decay mapping function is a time-decay transformation that operates on vectors. For a single spatial grid cell location, this function processes each component of the previously sensed cumulative state vector element-wise. Let p be the value of one of the components in the previously sensed cumulative state vector.prev The decay rate constant is λ, and the forgetting decay mapping function transforms this component into p. decay Its calculation logic is p decay =p prev ×exp lam (Exponential decay factor), where exp lam The correspondence with λ is determined by the type of the forgetting decay mapping function. For example, if the forgetting decay mapping function adopts an exponential decay form, then exp lam It is the negative λ power of the natural constant e multiplied by the control period duration. All components of the sensed cumulative state vector from the previous moment are processed according to this logic, and the resulting vector is the decayed sensed memory vector of the spatial grid cell position.

[0041] Step 1205: For each spatial grid cell location, simultaneously input the attenuated sensing memory vector and the current thermal environment state vector into the cumulative effect superposition function. The cumulative effect superposition function performs element-wise weighted fusion and nonlinear activation on the attenuated sensing memory vector and the current thermal environment state vector to generate the current time-time sensing cumulative state vector of the spatial grid cell location.

[0042] The cumulative effect superposition function, during element-wise processing, first assigns values ​​to the corresponding element components of the decayed sensory memory vector and the current thermal environment state vector. This function incorporates an adjustable fusion weight parameter α, ranging from zero to one, and a nonlinear activation operator. During element-wise fusion, the component value at one position in the decayed sensory memory vector is multiplied by the fusion weight parameter α to obtain the memory contribution term; the value of the corresponding component in the current thermal environment state vector is multiplied by the difference between 1 and α to obtain the immediate sensory contribution term; the two are added together to obtain the linear fusion result. The linear fusion result is then mapped using a nonlinear activation operator, which can be a hyperbolic tangent or sigmoid function. The purpose is to compress the fusion result to a preset sensory intensity range and introduce sensory saturation and threshold effects. This operation is repeated for all component positions of the two input vectors, ultimately obtaining the current-moment sensory cumulative state vector at that spatial grid cell position.

[0043] Step 1206: Combine the current time-based cumulative state vectors of all spatial grid cell locations, and reconstruct the spatial field according to the spatial grid cell distribution of the unified spatial coordinate system to generate a thermal sensing cumulative state field that includes complete spatial coverage and where each spatial grid cell location carries the current time-based cumulative state vector.

[0044] After processing each spatial grid cell location, the system obtains the corresponding current-moment cumulative state vector. These current-moment cumulative state vectors for all locations are then rearranged according to their spatial grid cell location identifiers into a three-dimensional spatial field structure. The spatial resolution of this field is consistent with the current thermal environment state field. This newly created field is the thermal sensing cumulative state field, which completely covers the target space, and each grid node stores the sensing state code after forgetting decay and accumulation fusion.

[0045] Step 1207: Store the thermal perception cumulative state field as the previous thermal perception cumulative state field for use in the next cycle, so that the preset somatosensory forgetting cumulative model can be continuously iterated and called in subsequent control cycles.

[0046] At the end of the current control cycle, the system writes the generated thermal sensing cumulative state field into the model parameter storage area, overwriting the previous thermal sensing cumulative state field. When starting step 1201 in the next control cycle, this field will be retrieved again as the previous thermal sensing cumulative state field, thus forming a continuous sensing memory iteration in the control closed loop.

[0047] Step 130: Start the state compression coding network of the multi-objective reinforcement learning agent, perform multi-level spatial feature abstraction and feature channel dimensionality reduction compression processing on the thermal sensing cumulative state field, and generate a compressed state representation vector suitable for policy optimization.

[0048] The multi-objective reinforcement learning agent includes a state compression coding network specifically designed to compress high-dimensional spatial field inputs into low-dimensional vector representations. This network extracts spatial features step by step from the thermal perception accumulated state field and, through channel compression and fully connected projection, compresses the high-dimensional perception state codes, which were originally distributed across a large number of spatial grid nodes, into a fixed-length compressed state representation vector. This significantly reduces the computational burden on subsequent policy networks while preserving core environmental context information closely related to thermal comfort and energy consumption decisions.

[0049] In its specific implementation, this step can be further broken down into the following sub-steps: Step 1301: Load the network weight parameters of the state compression coding network from the network parameter storage unit of the multi-objective reinforcement learning agent. The state compression coding network includes a first spatial feature abstraction layer, a spatial downsampling transition layer, a second spatial feature abstraction layer, a feature channel compression layer, and a fully connected projection layer, all connected in series.

[0050] The multi-objective reinforcement learning agent maintains a dedicated network parameter storage unit, which stores the weight matrices and bias vectors of each layer of the state compression coding network. This step loads these weight parameters into the inference computation unit, completing the network initialization. The state compression coding network adopts a hierarchical architecture: the first spatial feature abstraction layer captures local spatial patterns from the input field; the spatial downsampling transition layer compresses the spatial resolution; the second spatial feature abstraction layer extracts high-order semantic features in the compressed feature space; the feature channel compression layer removes redundant channels; and the fully connected projection layer maps the final spatial features to a one-dimensional vector.

[0051] Step 1302: Arrange the current time perception cumulative state vector of each spatial grid cell position in the thermal perception cumulative state field into a multi-channel input feature tensor according to the spatial grid cell distribution. Input the multi-channel input feature tensor into the first spatial feature abstraction layer. Perform local spatial pattern extraction and nonlinear mapping on the multi-channel input feature tensor through the local receptive field coding unit in the first spatial feature abstraction layer to obtain a primary spatial feature mapping set.

[0052] The thermal sensing cumulative state field is arranged according to its original X-axis, Y-axis, and Z-axis spatial grid distribution. Using the dimension of each sensing cumulative state vector as the number of channels, a four-dimensional multi-channel input feature tensor is formed. Its shape is defined by the number of X-axis, Y-axis, and Z-axis grids and the number of channels. The first spatial feature abstraction layer contains several local receptive field encoding units, each corresponding to a learnable small-sized three-dimensional convolutional kernel. This kernel slides along the X, Y, and Z axes on the input tensor. At each sliding window position, the weights of the convolutional kernel are element-wise multiplied and summed with the corresponding channel components of the sensing cumulative state vectors of each grid node covered within the window. A bias term is then added to calculate a response value. A nonlinear activation function is then applied to this response value to generate the output feature value for that channel.

[0053] Multiple coding units work in parallel, with each coding unit producing an independent output channel. These output channels together form a primary spatial feature map set. Each feature map in this set has a spatial dimension that is roughly equivalent to the input field, but the semantic level of the features has been elevated from the original perceptual state vector to an abstract expression of local spatial patterns.

[0054] Step 1303: Input the primary spatial feature map set into the spatial downsampling transition layer, and generate a spatially compressed intermediate feature map set with reduced spatial resolution and increased number of feature channels by performing spatial neighborhood compression and feature channel reorganization on the primary spatial feature map set.

[0055] The spatial downsampling transition layer first performs spatial neighborhood compression on the primary spatial feature map set. For the two spatial dimensions of X and Y, this layer slides with a window with a stride greater than 1, or uses a feature fusion downsampling method between adjacent grid nodes, such as averaging or maximizing the corresponding components of the feature vectors of two adjacent grid nodes, thereby reducing the number of grid nodes in space.

[0056] Simultaneously, this layer introduces new feature encoding channels. By linearly combining and redistributing multiple channels of the input primary spatial feature map set, the number of output channels is significantly increased compared to the number of input channels. After processing by the spatial downsampling transition layer, the number of spatial grid points in the X and Y axes of the generated spatially compressed intermediate feature map set is reduced, while the number of channels in the feature vector corresponding to each spatial location increases, realizing the conversion and compression of spatial information into channel information.

[0057] Step 1304: Input the spatial compression intermediate feature map set into the second spatial feature abstraction layer, and perform high-order spatial correlation feature extraction on the spatial compression intermediate feature map set through the extended receptive field coding unit in the second spatial feature abstraction layer to capture the large-scale spatial context dependency relationship in the thermal perception cumulative state field and generate a high-level spatial feature map set.

[0058] Since the spatial resolution of the spatially compressed intermediate feature map set has been reduced, the extended receptive field coding unit in the second spatial feature abstraction layer will use dilated convolution or a larger convolution kernel size, so that its effective receptive field can cover a larger spatial range on the input feature map. At each spatial grid node, the extended receptive field coding unit comprehensively considers the features of other nodes in a large neighborhood centered on that node. Through weighted summation of convolution weights and these node features and nonlinear transformation, it extracts high-order feature representations containing large-scale spatial context dependencies. Compared with the spatially compressed intermediate feature map set, the high-level spatial feature map set further improves the level of semantic abstraction. The feature vector of each spatial location not only represents the thermal perception state of the local area, but also embeds spatial correlation information with other areas, such as implicit structures such as temperature gradient direction and airflow path.

[0059] Step 1305: Input the high-level spatial feature map set into the feature channel compression layer, use the channel importance weight parameters stored in the feature channel compression layer to perform weighted recombination on each feature channel in the high-level spatial feature map set, and truncate and retain the weighted recombination feature channels in descending order of channel importance weight to generate a target feature map set with reduced channel dimension.

[0060] The feature channel compression layer internally stores a learnable channel importance weight vector, the length of which is equal to the number of channels in the high-level spatial feature map set. This layer first calculates the average response value across the entire space for each feature channel. This average response value is then multiplied by the corresponding weight scalar in the channel importance weight vector to obtain the weighted overall response for each channel. Subsequently, all channels are sorted from largest to smallest according to their weighted overall response, and the top few channels are selected as retained channels, discarding the rest. The indices of the retained channels correspond to the original channels, and feature map slices of these channels are directly extracted to form the target feature map set after channel dimensionality reduction. This process eliminates redundant channels that contribute little to subsequent policy decisions and extracts the most discriminative thermal perception feature representations.

[0061] Step 1306: Input the target feature mapping set into the fully connected projection layer. The spatial dimension information and channel dimension information of the target feature mapping set are integrated and mapped into a compressed state representation vector of a set length through the dimension alignment transformation matrix in the fully connected projection layer. Each numerical position in the compressed state representation vector output by the fully connected projection layer corresponds to an environmental perception feature component after spatial feature abstraction and channel compression.

[0062] The target feature map set still retains its spatial grid structure and channel structure. The fully connected projection layer first flattens the spatial and channel dimensions of this set into a one-dimensional long vector, and then performs a linear transformation on this long vector using a dimension alignment transformation matrix and a bias vector. The dimension alignment transformation matrix is ​​a learnable parameter matrix, with the number of rows equal to the preset length of the compressed state representation vector and the number of columns equal to the length of the flattened long vector. Matrix multiplication maps the flattened long vector into a shorter compressed vector. The value at each position of the compressed vector is obtained by weighted summing of the weights of the corresponding rows in the matrix and the elements of the flattened vector, plus the bias. Each component of the final compressed state representation vector is an environmental perception feature component that has been abstracted, selected, and integrated through the previous processing stages, comprehensively reflecting the global key information of the current thermal perception cumulative state.

[0063] Step 140: The multi-objective policy network of the multi-objective reinforcement learning agent performs forward inference based on the compressed state representation vector, and outputs a multi-dimensional control action vector that takes into account both the user thermal comfort optimization objective and the system energy consumption optimization objective.

[0064] The multi-objective policy network is the core decision-making module of the multi-objective reinforcement learning agent. It takes the compressed state representation vector from the state compression coding network as input and achieves a balanced solution for two potentially conflicting objectives, thermal comfort optimization and energy consumption optimization, through a specific architecture design. Finally, it outputs a multi-dimensional control action vector, in which each dimension corresponds to different adjustable variables of the air conditioning system.

[0065] In its specific implementation, this step can be further broken down into the following sub-steps: Step 1401: Load the network weight parameters of the multi-objective policy network from the network parameter storage unit of the multi-objective reinforcement learning agent. The multi-objective policy network includes a shared feature representation layer, a thermal comfort optimization sub-policy branch, an energy consumption optimization sub-policy branch, and a multi-objective action fusion layer.

[0066] The weight parameters of the multi-objective policy network are also stored in the agent's network parameter storage unit. The shared feature representation layer is located at the front end of the network and is responsible for extracting the hidden features shared by the two optimization objectives from the compressed state representation vector. The thermal comfort optimization sub-policy branch and the energy consumption optimization sub-policy branch are connected in parallel after the shared feature representation layer and are responsible for generating candidate control action components for a single objective, respectively. The multi-objective action fusion layer is located at the end of the network and is responsible for integrating the action components generated by the two branches into the final multi-dimensional control action vector.

[0067] Step 1402: Input the compressed state representation vector into the shared feature representation layer. The shared feature representation layer performs a nonlinear transformation and feature decoupling on the compressed state representation vector to generate a shared hidden layer feature vector. The shared hidden layer feature vector retains general environmental context information that is related to both the thermal comfort optimization objective and the system energy consumption optimization objective.

[0068] The shared feature representation layer typically consists of multiple fully connected networks. The compressed state representation vector sequentially passes through the first fully connected sub-layer, which multiplies the input vector by a first weight matrix and adds a first bias vector to obtain an initial transformed vector. Then, a non-linear activation is applied through an activation function sub-layer. Afterward, it passes through a second fully connected sub-layer, where it is multiplied by a second weight matrix and added a second bias vector, and activated again, ultimately generating the shared hidden layer feature vector. Each dimension of the shared hidden layer feature vector does not encode purely thermal comfort information or purely energy consumption information, but rather fuses related environmental context factors into a unified feature representation.

[0069] Step 1403: Input the shared hidden layer feature vector into the thermal comfort optimization sub-strategy branch and the energy consumption optimization sub-strategy branch simultaneously.

[0070] The shared hidden layer feature vector is copied into two copies. One copy is sent to the thermal comfort optimization sub-policy branch, and the other copy is sent to the energy consumption optimization sub-policy branch. The two branches are computed in parallel.

[0071] Step 1404: In the thermal comfort optimization sub-strategy branch, the shared hidden layer feature vector is refined through the thermal comfort feature transformation sub-layer to obtain the thermal comfort sensitive hidden layer vector. The thermal comfort action generation sub-layer outputs thermal comfort related control action components based on the thermal comfort sensitive hidden layer vector. The thermal comfort related control action components include the temperature setting change direction and the relative intensity of the temperature setting change, as well as the wind speed change direction and the relative intensity of the wind speed change.

[0072] The thermal comfort feature transformation sublayer consists of at least one fully connected sublayer and an activation function sublayer. It performs linear and nonlinear transformations on the shared hidden layer feature vector, enhancing the expression of feature components directly related to thermal comfort assessment and suppressing the expression of feature components related to energy consumption, outputting a thermal comfort-sensitive hidden layer vector. The thermal comfort action generation sublayer receives the thermal comfort-sensitive hidden layer vector and maps it to thermal comfort-related control action components through the output layer. The number of neurons in the output layer matches the number of thermal comfort control variables; for example, it outputs four scalars. The first scalar, after processing with a sign function, indicates the direction of temperature setting change; the second scalar, after processing with an absolute value activation function, indicates the relative intensity of temperature setting change; and the third and fourth scalars similarly correspond to the direction and relative intensity of wind speed change, respectively.

[0073] Step 1405: In the energy consumption optimization sub-strategy branch, the shared hidden layer feature vector is refined through the energy consumption-specific feature transformation sub-layer to obtain the energy consumption-sensitive hidden layer vector. The energy consumption-related control action component is output by the energy consumption action generation sub-layer based on the energy consumption-sensitive hidden layer vector. The energy consumption-related control action component includes the operating mode adjustment tendency and the compressor frequency adjustment tendency.

[0074] The energy consumption-specific feature transformation sublayer architecture is similar to the thermal comfort feature transformation sublayer, but with independent weight parameters. It reprocesses features to generate an energy consumption-sensitive hidden layer vector based on the energy consumption optimization objective. The energy consumption action generation sublayer maps the energy consumption-sensitive hidden layer vector into energy consumption-related control action components. For example, it outputs two scalars. The first scalar, after activation, represents the operating mode adjustment tendency value, which determines whether the air conditioner switches between cooling, heating, and ventilation modes, and the direction of the switch. The second scalar represents the compressor frequency adjustment tendency value, indicating the trend and magnitude of the compressor frequency increase or decrease relative to the current frequency.

[0075] Step 1406: The thermal comfort-related control action component and the energy consumption-related control action component are simultaneously input into the multi-objective action fusion layer. The thermal comfort-related control action component and the energy consumption-related control action component are weighted and combined using the multi-objective fusion weight parameters stored in the multi-objective action fusion layer. A multi-dimensional control action vector is generated through the fusion activation function. Each dimension of the multi-dimensional control action vector corresponds to an adjustable control variable of the air conditioner.

[0076] The multi-objective action fusion layer first reads the corresponding fusion weight pairs from the stored multi-objective fusion weight parameter matrix for each specific adjustable control variable of the air conditioner. For example, for temperature setting-related variables, it reads the temperature weight w from the thermal comfort branch. ct and energy consumption branch temperature weight w et The sum of the two weights is 1. The direction and relative intensity of the temperature setting change output from the thermal comfort branch are combined into a thermal comfort temperature action value. The potential temperature regulation tendency generated by the energy consumption branch is defined as the energy consumption temperature action value. The fused temperature control action value is then w. ct Multiply by thermal comfort temperature and add w et Multiply by the energy consumption and temperature action values. Similarly, perform this weighting process for each control variable such as wind speed, wind direction, operating mode, and compressor frequency. The fusion activation function applies a nonlinear transformation to the weighted result to ensure that the output action value is within the effective range that the air conditioner can execute. Finally, all the fused action values ​​are concatenated into a multi-dimensional control action vector.

[0077] Step 150: The multidimensional control action vector is converted into an operating parameter change that the air conditioning equipment can recognize through the actuator action mapping layer and the air conditioning operating status is updated. After the update, the subjective thermal feedback record of the current user is collected, and the decay rate constant of the forgetting decay mapping function in the preset somatosensory forgetting accumulation model is adaptively corrected using the subjective thermal feedback record.

[0078] The numerical values ​​in the multidimensional control action vector cannot be directly executed by the air conditioning hardware. They need to be converted into changes in operating parameters that the air conditioner can understand through the actuator action mapping layer. Meanwhile, the objective effect of control execution is monitored by environmental measurement nodes, but the user's subjective thermal comfort needs to be acquired through user feedback collection terminals and used to correct key parameters of the forgetting model, ensuring that the model continuously reflects the individual user's true perceptual characteristics. The specific implementation of this step is as follows: Step 1501: Analyze the correspondence between the values ​​of each dimension in the multidimensional control action vector and the adjustable control variables of the air conditioner, and separate the multidimensional change values ​​and the operating mode switching instruction codes.

[0079] The system maintains an action dimension mapping table, which defines the types of air conditioning control variables corresponding to the first to Nth dimensions in the multi-dimensional control action vector. According to this mapping table, the multi-dimensional control action vector is divided into two sets of data: the first set is the multi-dimensional change value, which includes the change amount of continuously adjustable variables such as temperature setpoint change value, fan speed change value, and wind direction change value; the second set is the operating mode switching instruction code, which is usually a discrete code vector indicating the target operating mode to which the system wants to switch.

[0080] Step 1502: Based on the mapping function processing result of the multidimensional change value and the mode switching control signal corresponding to the operation mode switching instruction code, generate an air conditioner operation parameter change instruction package and send it to the air conditioner execution controller to perform air conditioner operation status update.

[0081] In one alternative implementation, this sub-step can be further refined.

[0082] In this embodiment, the multidimensional change values ​​specifically include temperature setting change values, wind speed change values, and wind direction change values. The processing procedure is as follows.

[0083] Step 15021: Input the temperature setting change value into the temperature setting mapping function, linearly convert the temperature setting change value into a temperature adjustment step that is acceptable to the air conditioner temperature setting parameter register, and add it to the current air conditioner temperature setting value to generate the updated temperature setting value.

[0084] The temperature setting mapping function is a linear mapping equipped with a scaling factor and an offset. The temperature setting change value is multiplied by the scaling factor, rounded to the minimum temperature adjustment step allowed by the air conditioner, and then algebraically summed with the current temperature setting value of the air conditioner to obtain the updated temperature setting value.

[0085] Step 15022: Input the wind speed change value into the wind speed level mapping function, determine the target wind speed level according to the positive and negative direction and magnitude of the wind speed change value, and encode the target wind speed level into an air conditioning wind speed control signal.

[0086] The wind speed level mapping function pre-stores a segmented mapping table from a continuous range of wind speed changes to discrete wind speed levels. If the wind speed change is positive and its absolute value exceeds a threshold, the target wind speed level is increased by one level; if the wind speed change is negative and its absolute value exceeds the threshold, it is decreased by one level; if the absolute value is within the threshold range, the current level is maintained. The determined target wind speed level number is encoded into the corresponding wind speed control signal codeword.

[0087] Step 15023: Input the wind direction change value into the wind direction adjustment mapping function, determine the target wind direction blade position based on the angle deviation information of the wind direction change value, and encode the target wind direction blade position into an air conditioning wind direction control signal.

[0088] Wind direction change values ​​typically carry angular bias information. The wind direction adjustment mapping function constrains and maps this angular bias value according to the swing range that the guide vanes can support, obtains the target angular position of the vanes, and encodes it into a wind direction control signal according to the communication protocol.

[0089] Step 15024: Look up the air conditioner operating mode switching table according to the operating mode switching instruction code, and generate a mode switching control signal corresponding to the target operating mode.

[0090] The operating mode switching instruction is encoded as an integer. The system uses this integer to index the preset air conditioner operating mode switching table and retrieve the corresponding mode control signal encoding format.

[0091] Step 15025: Combine the updated temperature setpoint, air conditioner fan speed control signal, air conditioner fan direction control signal, and mode switching control signal to generate an air conditioner operating parameter change instruction package, and send the air conditioner operating parameter change instruction package to the air conditioner execution controller to perform the air conditioner operating status update.

[0092] The aforementioned control signal fields are combined according to the frame format specified in the communication protocol, and fields such as frame header, device address, and checksum are added to generate a complete air conditioning operating parameter change instruction packet. This packet is then sent to the air conditioning actuator controller via the control bus or wireless communication link. After parsing the instruction packet, the air conditioning actuator controller executes the corresponding operating parameter change operation.

[0093] Step 1503: Record the completion time of the air conditioner operation status update, and after a preset feedback acquisition delay time after the completion time, trigger the user feedback acquisition terminal arranged in the target space to acquire the current user's subjective thermal feedback record. The subjective thermal feedback record includes the user's thermal perception category identifier of the thermal environment at this moment.

[0094] After the air conditioner's operating status is updated, it takes some time for physical quantities such as ambient temperature to stabilize and affect user comfort. The system records a completion timestamp at the moment the status update is complete and starts a timer. The timer duration is equal to a preset feedback acquisition delay time, which is dynamically set based on the space volume and the air conditioner's cooling and heating capacity. Once the timer is triggered, the system sends a acquisition command to the user feedback acquisition terminal. The terminal collects the user's subjective feelings about the current thermal environment through an interactive interface. The user can select thermal perception category identifiers such as "slightly cold," "relatively cool," "comfortable," "warm," and "slightly hot." This identifier, along with the user identifier and the acquisition time, is recorded and transmitted back as subjective thermal perception feedback.

[0095] Step 1504: Perform deviation analysis between the subjective thermal feedback record and the current time perception cumulative state vector of the key spatial grid unit position in the thermal perception cumulative state field, calculate the thermal deviation direction and degree, and generate thermal deviation descriptive quantity.

[0096] First, determine the location of key spatial grid cells, such as the spatial grid cells corresponding to the user's usual activity area. Extract the current time-to-time perception cumulative state vector of this grid cell location from the current thermal perception cumulative state field. Convert each component of this vector into the expected thermal sensation value according to a preset mapping rule from perception state to thermal sensation. Quantify the thermal sensation category identifier provided by the user into a set of thermal sensation score values. Calculate the deviation between the two; the direction of the deviation indicates whether the perception model is biased towards heat estimation or cold estimation, and the degree of deviation indicates the magnitude of the deviation. This generates a thermal sensation deviation descriptor, which can be a two-dimensional vector.

[0097] Step 1505: Input the thermal deviation description quantity into the attenuation rate constant correction unit. The attenuation rate constant correction unit generates an attenuation rate constant adjustment quantity according to the mapping logic between the thermal deviation description quantity and the current value of the attenuation rate constant. The current value of the attenuation rate constant is superimposed with the attenuation rate constant adjustment quantity to obtain the corrected attenuation rate constant.

[0098] The decay rate constant correction unit determines the sign of the adjustment amount based on the direction of the deviation of the thermal perception deviation descriptor. If the perception model continuously underestimates the changes in thermal perception caused by environmental changes, the decay rate constant needs to be appropriately increased to accelerate the forgetting rate; conversely, it should be decreased. The degree of deviation is linearly scaled by a preset correction gain coefficient and used as the magnitude of the adjustment amount. This adjustment amount is added to the current value of the decay rate constant, and upper and lower limits of the value range are applied to obtain the corrected decay rate constant.

[0099] Step 1506: Write the corrected decay rate constant into the preset somatosensory forgetting accumulation model to replace the original decay rate constant, so that it can be called in the next control cycle.

[0100] The corrected decay rate constant obtained in step 1505 is persistently written to the model parameter storage area, overwriting the original decay rate constant value. When step 1204 is executed in the next control cycle, the somatosensory forgetting accumulation model will use the updated decay rate constant to perform forgetting decay mapping, thereby achieving continuous adaptation to the user's somatosensory characteristics.

[0101] Based on the correction of the forgetting model parameters using feedback, the embodiments of this application can further utilize data accumulated over multiple control cycles to perform higher-level adaptive optimization of the system: Step 210: Obtain the accumulated subjective thermal feedback recording sequence and the corresponding corrected decay rate constant sequence within multiple control cycles.

[0102] After completing parameter calibration for at least one full control cycle, the system stores the calibrated decay rate constant and the collected subjective thermal feedback record of the current cycle as a record pair in the historical experience database. The historical experience database is designed as a persistent storage structure indexed by user identifier and spatial identifier. Each record pair contains the control cycle number, timestamp, thermal category identifier in the subjective thermal feedback record, the complete value of the calibrated decay rate constant, and the location identifier of the corresponding key spatial grid cell for that cycle. When the system has completed a preset number of control cycles, such as accumulating dozens of control cycle record pairs, the sequence extraction process in this step is triggered. The system queries and sorts the historical experience database according to the user identifier and spatial identifier, organizing all record pairs into two aligned sequences according to the chronological order of the control cycles. The first sequence is the subjective thermal feedback record sequence, where each element is a thermal category identifier, such as "cool," "comfortable," or "warm," describing the user's subjective feeling at the current moment. The second sequence is the corresponding calibrated decay rate constant sequence, where each element is the specific value of the decay rate constant of the forgetting decay mapping function that was calibrated and stored in the same control cycle. Both sequences are equal in length to the cumulative number of control period samples and are strictly aligned periodically. This serialized data organization provides a complete temporal observation window for subsequent analysis of individual perception time-varying characteristics.

[0103] Step 220: Encode the thermal sensation category identification conversion relationship between adjacent control cycles in the subjective thermal sensation feedback recording sequence into thermal sensation state transition features, and divide the corrected decay rate constant sequence into decay rate fluctuation segments according to the time window.

[0104] For a subjective thermal feedback recording sequence, the system iterates through two adjacent elements in the sequence, i.e., the thermal category identifiers of the t-th control cycle and the (t+1)-th control cycle constitute an identifier transition pair. A preset set of thermal category identifiers contains a finite number of discrete categories, for example, five thermal categories, which are mapped to integer indices 0 to 4. For each extracted identifier transition pair, the system encodes the transition pair into a one-hot vector based on the integer index values ​​of the preceding and following identifiers using a 5×5 two-dimensional encoding table. This one-hot vector has a length of 25 and corresponds to all possible transition combinations, with a value of 1 only at the position corresponding to the transition pair and 0 at the rest. By performing the same one-hot encoding on all identifier transition pairs in the sequence, a sequence of one-hot vectors is obtained, which constitutes the thermal state transition feature. This feature compactly represents the direction and frequency of the user's subjective thermal experience state transitions within the existing control cycle span.

[0105] For the corrected decay rate constant sequence, the system uses a fixed-length time window for sliding segmentation. The time window length is set to include a fixed number of control period samples, and the sliding step size is one control period. Several consecutive decay rate constant values ​​in the decay rate constant sequence are considered as a decay rate fluctuation segment, with each segment containing the change in the decay rate constant within the window's coverage area. To eliminate the influence of absolute value differences between different time window segments on subsequent analysis, the system performs numerical normalization processing within each decay rate fluctuation segment. This involves subtracting the mean of all decay rate constants in the segment from each decay rate constant in the segment, and then dividing by the range or standard deviation of the segment. Each processed decay rate fluctuation segment is transformed into a dimensionless normalized fluctuation vector. The set of normalized fluctuation vectors generated by all time windows together constitutes the set of segmented decay rate fluctuation segments.

[0106] Step 230: By inputting the thermal state transition features and the decay rate fluctuation segments into a preset individual perception drift analysis network, an individual perception drift pattern vector characterizing the time-varying characteristics of the user's thermal perception is generated.

[0107] The Individual Perceptual Drift Analysis Network is a time-series processing network specifically designed for perceptual dynamics analysis. Its architecture consists of a cascaded input alignment layer, a sequence encoding layer, and a projection output layer. The input alignment layer is responsible for pairing and aligning the thermal state transition feature sequences and decay rate fluctuation segment sequences along the time dimension. Due to the segmentation of the time window, the number of decay rate fluctuation segments is slightly less than the number of transition pairs in the thermal state transition features. The alignment method uses the center time point of the time window corresponding to the decay rate fluctuation segment as a reference, searches for thermal state transition features in the time period near that time window, and selects two transition pairs adjacent to the time period covered by that time window as the aligned pairing input.

[0108] The sequence coding layer employs a bidirectional gated recurrent unit (RRN) network, comprising forward and backward propagation channels. At each time step, the forward propagation channel receives the aligned paired input of the current time step in ascending chronological order, and updates the forward hidden state vector of the current time step by combining it with the forward hidden state vector of the previous time step. The backward propagation channel receives the paired input in reverse chronological order, and updates the backward hidden state vector of the current time step by combining it with the backward hidden state vector of the next time step. The complete hidden state vector of each time step is obtained by concatenating the forward and backward hidden state vectors along the feature dimension. The gating structure of the bidirectional gated RRN network includes update and reset gates. The update gate controls the proportion of previous hidden state information incorporated into the current hidden state, while the reset gate controls the degree to which previous hidden state information is ignored. Both gating signals are calculated using an activation function based on the current input and the previous hidden state. Through this bidirectional temporal modeling approach, the sequence coding layer can simultaneously capture the dependencies of user-perceived characteristics in both historical and future directions.

[0109] The projection output layer receives the complete hidden state vector output from the last time step of the sequence coding layer, and projects this vector into a vector of a set dimension through a fully connected transformation. This vector is the individual perception drift pattern vector. The dimensions of the individual perception drift pattern vector do not directly correspond to any physical semantics, but are a compact implicit encoding of the time-varying characteristics of the user's subjective thermal perception, such as drift tendency, drift rate, and drift direction, which are spontaneously formed during the training process.

[0110] Step 240: Align the individual perception drift pattern vector with the forgetting decay mapping function parameters at the current moment in the preset somatosensory forgetting accumulation model by context association, and synchronously update the decay rate constant in the forgetting decay mapping function and the fusion weight coefficient in the cumulative effect superposition function according to the alignment result.

[0111] After obtaining the individual perceptual drift pattern vector, the system initiates the parameter association alignment and joint update process. First, it reads the currently effective forgetting decay mapping function parameters, including the decay rate constant λ, from the model parameter storage area of ​​the somatosensory forgetting accumulation model. current The fusion weight parameter α in the superposition function of cumulative effects current Simultaneously, the auxiliary state parameters of the forgetting decay mapping function and the cumulative effect superposition function at the current moment are read, such as the saturation state quantity of the cumulative effect, etc. These parameters and state parameters are concatenated into a context parameter vector. The dimension of the context parameter vector may not be consistent with the dimension of the individual perception drift pattern vector, so dimension matching needs to be performed through the context alignment module.

[0112] The context alignment module employs a cross-attention alignment mechanism. Specifically, the context parameter vector is transformed into a query vector sequence using a linear transformation matrix, and the individual perception drift pattern vector is transformed into key and value vectors using another linear transformation matrix. By calculating the similarity of the inner product between the query and key vectors, an attention weight distribution is obtained. This distribution characterizes which components of the context parameter vector require adjustment and which dimensions of the individual perception drift pattern vector encode the drift characteristics most relevant to them. The value vector is then weighted and summed using these attention weights to generate a transfer adjustment vector with the same dimension as the context parameter vector. This transfer adjustment vector integrates information about the direction and magnitude of adjustments needed for each model parameter based on the time-varying characteristics revealed by the individual perception drift pattern.

[0113] Based on the migration adjustment vector, the system parses two sub-adjustment vectors. According to a preset dimensionality partitioning rule, the first half of the migration adjustment vector corresponds to the decay rate constant adjustment increment δ_λ, and the second half corresponds to the fusion weight coefficient adjustment increment δ_α. δ_λ is then compared with the current decay rate constant λ. current Algebraic addition yields the updated decay rate constant λ. updated The δ_α is compressed using an activation function and then combined with the current fusion weight parameter α. current Algebraic addition yields the updated fusion weight coefficient α. updated For λ updated and α updated Apply predefined valid value range constraints respectively. If the updated value exceeds the upper limit of the valid value range, force it to be set to the upper limit. If it is lower than the lower limit, set it to the lower limit to ensure the physical validity of the parameter.

[0114] Finally, the corrected λ updated and α updated Together with the preset model parameter storage area of ​​the somatosensory forgetting accumulation model, the parameters are written to substantially replace the original decay rate constant and fusion weight coefficients. After this step, the somatosensory forgetting accumulation model not only corrects the decay rate constant through real-time feedback during continuous control, but also synchronously optimizes the decay rate constant and fusion weight parameters based on the long-term accumulated individual perception drift pattern, thereby further enhancing the model's adaptability to changes in users' long-term perception characteristics.

[0115] As another optional embodiment, the method further includes: Step 310: Extract the thermal sensing state evolution trajectory at a specified spatial grid cell location from the accumulated thermal sensing state field sequence accumulated over multiple control cycles.

[0116] In the historical state database of the control system, the thermal sensing cumulative state field generated at the end of each control cycle is completely and persistently stored, associated with the control cycle number and metadata description of a unified spatial coordinate system. To analyze the dynamic characteristics of thermal sensing at a specific spatial location, the system first needs to determine at least one specified spatial grid cell location. This location is typically selected as the centroid grid coordinates of the user's long-term residence area, or as the grid coordinates of the high-frequency human activity coverage obtained through statistical analysis of occupancy detection data from multiple cycles. After determining the specified spatial grid cell location, the system traverses all thermal sensing cumulative state field records in the historical state database, sorted in ascending order by control cycle number. For each thermal sensing cumulative state field, the system uses the spatial coordinate system corresponding to that field as an index to directly extract the current-moment sensing cumulative state vector at the grid node whose location coordinates are equal to the specified spatial grid cell location. Arranging the current moment-accumulated state vectors of all control cycles sequentially in time sequence constitutes the thermal sensing state evolution trajectory at the location of the spatial grid cell. Each element of this trajectory is a multi-dimensional vector, the dimension of which is equal to the dimension of the accumulated state vector. The time span of this trajectory covers all historical control cycles involved in the analysis, thus providing a continuous perspective for observing the evolution of thermal sensing state over time at a fixed point in space.

[0117] Step 320: Input the thermal sensing state evolution trajectory into a preset sensing state differential analysis network. Through the sensing state differential analysis network, extract temporal difference features and analyze nonlinear dynamic patterns of the thermal sensing state evolution trajectory to generate a sensing manifold structure description vector that characterizes the spatiotemporal evolution law of the current thermal sensing cumulative state field.

[0118] The perceptual state differential analysis network is a type of neural network specifically designed to analyze the local variation features and global dynamic patterns of temporal trajectories. It comprises two main computational stages: a temporal differential encoding layer and a multi-scale dynamic analysis layer. The thermal perceptual state evolution trajectory is first fed into the temporal differential encoding layer as a time-step sequence. The temporal differential encoding layer performs a time-step-by-time adjacent difference operation on this trajectory. Specifically, for the perceptual cumulative state vector at the k-th time step and the perceptual cumulative state vector at the (k+1)-th time step, the difference between the previous and subsequent values ​​of each component is calculated, resulting in a temporal difference vector with the same dimension as the perceptual cumulative state vector. To capture higher-order change information, the temporal differential encoding layer further performs a difference operation on the obtained first-order difference vector sequence, generating a second-order difference vector sequence to represent the state change features at the acceleration level. The temporal differential encoding layer concatenates the original trajectory vector, the first-order difference vector, and the second-order difference vector along the feature dimension to form an augmented trajectory feature sequence.

[0119] The augmented trajectory feature sequence is then fed into a multi-scale dynamics analysis layer. This layer contains several parallel temporal convolutional coding branches, each employing a causal dilated convolutional kernel with a different dilation rate. Branches with smaller dilation rates focus on characterizing local short-term state fluctuation patterns, while branches with larger dilation rates emphasize modeling long-term trends and periodic dynamic behaviors. Each branch performs convolution and nonlinear activation processing on the augmented trajectory feature sequence at its respective time scale, outputting a dynamics hidden layer feature sequence at the corresponding scale. All the dynamics hidden layer feature sequences output by the branches are then compressed into a fixed-length branch aggregation vector via global average pooling in the time dimension. These branch aggregation vectors are then concatenated along the feature dimension and integrated through a fully connected fusion sublayer to generate a single vector—the perceptual manifold structure description vector. This vector implicitly encodes the intrinsic pattern of the thermal perception cumulative state field's evolution over time at specified key spatial locations, including the rate of state change, the persistence of change, and the presence or absence of oscillating and recurring manifold structure properties.

[0120] Step 330: Based on the perceptual manifold structure description vector, the spatial neighborhood compression range parameter of the spatial downsampling transition layer and the extended receptive field encoding range parameter of the second spatial feature abstraction layer in the state compression coding network of the multi-objective reinforcement learning agent are jointly reconfigured so that the feature abstraction granularity of the state compression coding network is adaptively matched with the spatiotemporal evolution rate of the current thermal perception cumulative state field.

[0121] After generation, the perceptual manifold structure description vector is fed into a configuration parameter mapping unit. This unit consists of two parallel lightweight mapping networks: a spatial neighborhood compression branch and a receptive field branch. The spatial neighborhood compression branch receives the perceptual manifold structure description vector and maps it to a set of spatial neighborhood compression range adjustment factors through a fully connected transformation and activation function. These factors include downsampling step size adjustment factors for the X and Y axes. When the perceptual manifold structure description vector represents a rapidly changing spatiotemporal evolution phase of the thermal perception state field, these adjustment factors reduce the sliding window step size of the spatial downsampling transition layer, even approaching 1, to preserve spatial detail differences as much as possible. When the state field evolves smoothly and slowly, the adjustment factors appropriately increase the step size to compress spatial information over a larger range and highlight macroscopic features.

[0122] The receptive field branch also receives the perceptual manifold structure description vector, mapping it to a set of dilation coefficient adjustment factors. These factors directly affect the extended receptive field encoding units in the second spatial feature abstraction layer. The extended receptive field encoding units internally use dilated convolutions, and their dilation rate parameter is modulated by multiplication using these adjustment factors. If the perceptual state evolves rapidly, the dilation rate is lowered, making the receptive field relatively focused on local detail relationships; if the evolution is slow and the spatial pattern is stable, the dilation rate is increased, expanding the receptive field to capture a wide range of global contextual dependencies.

[0123] The spatial neighborhood compression branch and the receptive field branch share information from the perceptual manifold structure description vector, but have independent mapping weights. This ensures that the spatial downsampling and receptive field structural components can change collaboratively without needing to be completely synchronized. After calculating the new spatial neighborhood compression range adjustment factor and expansion coefficient adjustment factor, the configuration parameter mapping unit writes these parameters into the parameter register of the corresponding layer of the state compression coding network through the network parameter update interface, replacing the old spatial neighborhood compression range parameters and expanded receptive field coding range parameters. After the replacement is completed, at the next control cycle, the multi-objective reinforcement learning agent reloads the state compression coding network parameters, and the feature abstraction granularity automatically adapts to the spatiotemporal evolution rate of the current thermal perception cumulative state field.

[0124] As another optional embodiment, the method further includes: Step 410: Obtain the multi-dimensional control action vector sequence output by the multi-objective reinforcement learning agent within multiple control cycles and the multi-objective fusion weight parameter sequence used by the multi-objective action fusion layer in the multi-objective policy network at the corresponding time.

[0125] In the decision data recording module of the multi-objective reinforcement learning agent, the multi-dimensional control action vector output by the multi-objective policy network for each control cycle, as well as the multi-objective fusion weight parameters actually invoked by the multi-objective action fusion layer during that forward inference, are all recorded as log entries. Each log entry includes at least the control cycle number, timestamp, a complete list of values ​​for each dimension of the multi-dimensional control action vector, and a snapshot of the weight coefficients for each action variable in the multi-objective fusion weight parameter matrix. When the system needs to analyze the multi-objective optimization equilibrium situation, it reads log entries covering several recent control cycles from this decision data recording module and arranges these entries in ascending order of control cycle number. From these log entries, the multi-dimensional control action vectors for each control cycle are extracted and arranged sequentially to form a multi-dimensional control action vector sequence; simultaneously, snapshots of the multi-objective fusion weight parameters for each control cycle are extracted and arranged sequentially to form a multi-objective fusion weight parameter sequence. These two sequences are perfectly aligned in time, and their lengths are equal to the number of control cycles read.

[0126] Step 420: Encode the difference in action change direction between adjacent control cycles in the multidimensional control action vector sequence as a strategy decision fluctuation feature, and encode the weight change direction between adjacent control cycles in the multi-objective fusion weight parameter sequence as a multi-objective preference shift feature.

[0127] Each action vector in the multidimensional control action vector sequence contains values ​​for multiple dimensions, including temperature setting-related action components, wind speed-related action components, wind direction-related action components, operating mode-related action components, and compressor frequency-related action components. Traversing two adjacent control cycles (cycle t and t+1), the values ​​of corresponding dimensions are subtracted to obtain an action difference vector with the same dimensions as the multidimensional control action vector. For each component in the action difference vector, its sign is extracted: a positive value represents a positive change in that dimension, a negative value represents a negative change, and a zero value represents no change, thus obtaining a sign vector. The sign vector sequences of all adjacent cycles are integrated. The system forms a local distribution using the nearest neighbor sign vectors. The frequency of various sign change patterns within each local sequence is statistically analyzed. The normalized statistical distribution vector is used as a strategy decision fluctuation feature, reflecting the fluctuation pattern and stability of the control action decision direction under multi-objective strategy driving in continuous control.

[0128] For the multi-objective fusion weight parameter sequence, each snapshot records the fusion weight pairs corresponding to each action variable, such as the thermal comfort weight and energy consumption weight for the temperature action. For weight snapshots of two adjacent periods, the difference between each weight pair is calculated to obtain the direction of weight change for each action variable: if the thermal comfort weight increases, the direction of change is marked as positive; if the thermal comfort weight decreases, the direction of change is marked as negative. The direction of weight change for all action variables is co-encoded into a multi-objective preference shift feature vector, the length of which is equal to the number of action variables. Each element can take the value representing whether the user's preference shifts towards thermal comfort, energy consumption, or remains unchanged for that action. By summarizing the multi-objective preference shift features of multiple consecutive adjacent periods, the system further obtains the preference change trend at the multi-objective balance level.

[0129] Step 430: By inputting the strategy decision fluctuation characteristics and the multi-objective preference offset characteristics into a preset multi-objective trade-off analysis network, a multi-objective trade-off situation vector representing the current multi-objective optimization balance situation is generated.

[0130] The multi-objective tradeoff analysis network employs a dual-input channel architecture, consisting of a feature fusion layer and a situation classification layer. Policy decision fluctuation features and multi-objective preference shift features are first concatenated in the feature fusion layer, with the components of the policy decision fluctuation features added first, followed by the components of the multi-objective preference shift features, forming a combined feature vector. This combined feature vector is then fed into the situation classification layer, which is constructed from multiple fully connected networks. Each fully connected network layer is followed by a normalization layer and an activation layer, using a linear rectified function with leakage as the activation function. In the final layer of the situation classification layer, the transformed features are mapped to a multi-objective tradeoff situation vector equal to the number of action variables through a fully connected matrix. Each element in the multi-objective tradeoff situation vector corresponds to an adjustable control variable, and its output value is mapped through a hyperbolic tangent activation function, taking values ​​between -1 and 1. When the value of an element approaches 1, it indicates that the current strategy is significantly biased towards thermal comfort optimization on that control variable; when the value approaches -1, it indicates that the current strategy is significantly biased towards energy consumption optimization on that control variable; when the value is close to 0, it indicates that the state is relatively balanced.

[0131] Step 440: Feed the multi-objective trade-off situation vector back to the thermal comfort optimization sub-strategy branch and the energy consumption optimization sub-strategy branch of the multi-objective strategy network, and perform coordinated fine-tuning of the action generation strategy parameters within the thermal comfort optimization sub-strategy branch and the energy consumption optimization sub-strategy branch to adjust the action synthesis balance of the multi-objective strategy network in the multi-objective optimization process.

[0132] The multi-objective tradeoff state vector is propagated back into the multi-objective policy network as a guiding signal for fine-tuning the action generation policy parameters. For each action variable, the system reads the element value corresponding to that variable from the multi-objective tradeoff state vector. Based on the sign and absolute value of this value, it calculates the gain adjustment factor for that variable in the thermal comfort optimization sub-policy branch and the energy consumption optimization sub-policy branch. If the tradeoff state value of a certain action variable is biased towards thermal comfort, the gain adjustment unit calculates a positive gain factor for the action generation sub-layer of the thermal comfort branch and a negative gain factor for the action generation sub-layer of the energy consumption branch. The positive gain factor is 1 plus a scaling factor based on the absolute value of the tradeoff state, and the negative gain factor is 1 minus the same scaling factor. These two gain factors are multiplied by the connection weights responsible for outputting the action variable in the corresponding branch action generation sub-layer, thereby enhancing the relative influence of the thermal comfort side on the control variable without changing the network structure, while weakening the adversarial output of the energy consumption side.

[0133] Cooperative fine-tuning is performed in parallel across all action variables, with the gain factor for each action variable calculated independently. To maintain policy stability and avoid overcorrection, the system sets smoothing limits on the magnitude of gain factor changes, and the upper limit of scaling in a single adjustment is constrained by hyperparameters. In addition to gain adjustment, when certain components of the multi-objective tradeoff situation vector exhibit extreme bias over multiple consecutive cycles, the fine-tuning mechanism selectively adjusts the initial values ​​of the fusion weight parameters corresponding to these actions in the next cycle. This causes the weight snapshot of the multi-objective action fusion layer to be corrected towards balance in advance, thereby gradually restoring the equilibrium of subsequent action synthesis to the preset ideal range. After multi-level cooperative fine-tuning, the multi-dimensional control action vector output by the multi-objective policy network in the next forward inference will reflect a balance between thermal comfort and energy consumption that better meets the current global actual needs of the system.

[0134] Optionally, the specific architecture, key module technical implementation, training process, and application methods of the various artificial intelligence models and networks appearing in the embodiments of this application are further explained as follows: For the state compression coding network of a multi-objective reinforcement learning agent, its specific architecture adopts a hierarchical convolutional encoder structure. The first spatial feature abstraction layer consists of multiple parallel 3D convolutional coding units, each containing a 3D convolutional kernel and a channel-level nonlinear activation function. The spatial downsampling transition layer uses strided 3D convolution operations instead of pooling operations, with the stride size set along the spatial dimension, and simultaneously expands and reorganizes the number of channels in the channel direction through 1×1×1 convolutions. The second spatial feature abstraction layer consists of 3D convolutional coding units with dilation rates, and the dilation rate parameter can be dynamically modulated according to the perceptual manifold description vector. The feature channel compression layer internally maintains a learnable channel importance weight vector, where each element of the vector corresponds to a channel in the high-level spatial feature map, and channel pruning is achieved during forward propagation through weighted sorting and truncation. The fully connected projection layer consists of a single-layer fully connected transformation matrix, which linearly maps the flattened multi-dimensional tensor to a compressed state representation vector of a preset length. The training of the state compression coding network is performed independently in the first stage of the overall multi-objective reinforcement learning training, using a reconstruction-based self-supervised pre-training strategy.

[0135] The training data comes from thermal sensing cumulative state field samples recorded by air conditioning systems operating under various climatic conditions and in different simulated building spaces, totaling approximately hundreds of thousands of samples. The input is a real or simulated thermal sensing cumulative state field, and the output is a compressed state representation vector. The training objective is to reconstruct the input sensing cumulative state field with high accuracy using the compressed state representation vector through a mirrored decoder. An adaptive momentum estimation optimizer is used, with an initial learning rate set to a small order of magnitude. The batch size is adjusted according to computational resources. The training epochs are stopped when the reconstruction loss no longer decreases on the validation set. The reconstruction loss is the mean square error of each component of the sensing state vector. After training, the decoder is removed, and the weights of the compressed coding network are frozen or only fine-tuned. During inference, the thermal sensing cumulative state field generated in each control cycle is directly fed into each layer of the state compressed coding network, sequentially passing through 3D convolutional coding, strided downsampling, dilated convolutional coding, channel-weighted pruning, and fully connected projection to obtain the compressed state representation vector for subsequent policy networks.

[0136] For the multi-objective policy network, its overall architecture is a multi-branch fully connected policy network. The shared feature representation layer consists of two cascaded fully connected sub-layers and corresponding nonlinear activation layers. The first fully connected sub-layer maps the compressed state representation vector to an intermediate dimension, which is typically set to about twice the dimension of the compressed state representation vector. The second fully connected sub-layer maps the intermediate dimension to the dimension of the shared hidden layer feature vector. The thermal comfort optimization sub-policy branch and the energy consumption optimization sub-policy branch are both independent fully connected branch networks, each containing a feature transformation sub-layer and an action generation sub-layer. The feature transformation sub-layer consists of a single fully connected layer and an activation function. The action generation sub-layer maps the output of the component values ​​corresponding to the action space dimension using a single fully connected layer. The thermal comfort branch outputs four components: the direction of temperature setting change, the intensity of temperature setting change, the direction of wind speed change, and the intensity of wind speed change. The energy consumption branch outputs two components: the tendency to adjust the operating mode and the tendency to adjust the compressor frequency.

[0137] The multi-objective action fusion layer internally stores fusion weight parameters in the form of a diagonal matrix. After weighted summation of each action variable, a fusion activation function is applied. The fusion activation function uses a scaled hyperbolic tangent function to ensure the output remains within the effective action boundary. The multi-objective policy network is trained within a multi-objective reinforcement learning framework, employing a training algorithm based on multi-objective Q-learning. Training data comes from online interactions between the air conditioning control environment and a simulated user thermal comfort model. One experience sample is generated at each time step, and the experience replay buffer stores approximately hundreds of thousands of experiences. Each experience includes a compressed state representation vector, a multi-dimensional control action vector, an environmental feedback thermal comfort reward value, an energy consumption penalty value, and the next compressed state representation vector. Training uses two identical target networks with asynchronously updated parameters, used for thermal comfort Q-value estimation and energy consumption Q-value estimation, respectively. The loss function for single-step optimization is the weighted sum of the temporal difference errors of the two Q-value branches. The thermal comfort weights and energy consumption weights are dynamically adjusted according to the multi-objective trade-offs.

[0138] The optimizer also employs an adaptive momentum estimation optimizer. The learning rate decays during training using a cosine annealing strategy. The batch size for each sample from the buffer is a preset fixed value, and the target network soft update coefficient is set to a relatively small value. During inference, forward propagation is performed directly based on the current compressed state representation vector, without sampling or exploration. The output multidimensional control action vector is converted into air conditioning execution instructions via an action mapping layer.

[0139] The architecture and training details of the adaptive update of the forgetting decay mapping function and the cumulative effect superposition function parameters in the pre-defined somatosensory forgetting accumulation model, as well as auxiliary analysis networks such as the individual perceptual drift analysis network, the perceptual state differential analysis network, and the multi-objective tradeoff analysis network, are as follows. The individual perceptual drift analysis network employs a bidirectional gated recurrent unit sequence encoding network with a single layer. The hidden state dimension is determined by the thermal state transition feature dimension and the length of the decay rate fluctuation segment. The projection output layer following the sequence encoding layer is a single fully connected layer, outputting an individual perceptual drift pattern vector with a preset fixed length. The perceptual state differential analysis network consists of a temporal difference encoding layer and a multi-scale dynamic analysis layer. The temporal difference encoding layer implements first-order and second-order difference operations through programming logic. The multi-scale dynamic analysis layer uses multiple one-dimensional dilated convolution branches with dilation rates distributed in a geometric progression within a preset range. The kernel length and output channel number of each branch are consistent. After aggregation, the output perceptual manifold structure description vector is fused through pooling and a fully connected layer.

[0140] The multi-objective tradeoff analysis network employs a dual-input concatenation followed by a three-layer fully connected network structure, with the number of neurons in each layer progressively decreasing. The final output dimension corresponds to the number of action variables. The hidden layers use a linear rectified activation function with leakage, and the output layer uses a hyperbolic tangent activation function. These auxiliary analysis networks are trained using offline supervised or self-supervised methods. The training data consists of parameter and state sequences recorded during the long-term operation of the air conditioning system, with a sample size covering operational data from multiple user models and various seasonal conditions. The training objective of the individual perception drift analysis network is to predict the changing trend of the decay rate constant over several future periods. The training objective of the perception state differential analysis network is to classify the evolutionary dynamics type of perception states. The training objective of the multi-objective tradeoff analysis network is to fit a manually calibrated multi-objective equilibrium state score. Each auxiliary network uses an adaptive momentum estimation optimizer for parameter updates, with batch size and learning rate set according to the network size. During training, the performance on the validation set is monitored, and the network is terminated early after performance stabilizes. In inference applications, these networks all operate in forward inference mode. The input is preprocessed into sequence tensors or aligned vector pairs as required, and the vectors output by the network are directly used for parameter updates or network reconfiguration without additional inverse transformations or thresholding.

[0141] This application's embodiments reconstruct a unified thermal environment state field in the target space and introduce a somatosensory forgetting accumulation model to perform perceptual attenuation mapping and cumulative effect superposition of thermal environment physical quantities. This enables the control system to continuously simulate the user's subjective perception of changes in the thermal environment and its memory attenuation characteristics, rather than merely responding to instantaneous physical measurements. Furthermore, the multi-objective reinforcement learning agent utilizes a state compression coding network to perform multi-level spatial feature abstraction and feature channel dimensionality reduction compression on the thermal perception accumulation state field, generating a highly condensed compressed state representation vector that retains key environmental context information, effectively reducing the state space complexity of policy optimization. Based on this compressed state representation vector, the multi-objective policy network simultaneously outputs a multi-dimensional control action vector that considers both user thermal comfort optimization and system energy consumption optimization objectives. This vector is then converted into changes in air conditioning operating parameters through an actuator action mapping layer to complete the control closed loop. Based on this, the decay rate constant of the forgetting decay mapping function is adaptively corrected by using the subjective thermal feedback records collected after the control update, so that the somatosensory forgetting accumulation model continuously approximates the actual perceived dynamic characteristics of individual users. This enables refined energy consumption management while ensuring thermal comfort, and solves the problems of disconnect between thermal environment control and user subjective feelings and imbalance of multi-objective optimization in existing technologies.

[0142] See Figure 2 As shown in the figure, this is a schematic diagram of the basic structure of a multi-objective reinforcement learning air conditioning control system 20 provided in an embodiment of this application. The multi-objective reinforcement learning air conditioning control system 20 includes: Processor 201; Storage device 202, on which computer program 2020 is stored; When the computer program 2020 is executed by the processor 201, the processor 201 implements any of the aforementioned multi-objective reinforcement learning air conditioning control methods based on the forgetting model.

[0143] Based on the above, a readable storage medium is provided, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the above method are implemented.

[0144] See Figure 3 As shown, this figure is a functional block diagram of a multi-objective reinforcement learning air conditioning control device provided in an embodiment of this application. The multi-objective reinforcement learning air conditioning control device includes: The environmental state reconstruction module is used to perform time-series synchronization and spatial field structure reconstruction processing on the original environmental physical quantity measurement set collected by multiple environmental measurement nodes distributed in the target space, and generate the current thermal environment state field based on a unified spatial coordinate system. The cumulative effect superposition module is used to perform perception attenuation mapping and cumulative effect superposition processing on the current thermal environment state field according to the preset somatosensory forgetting accumulation model, so as to obtain the thermal perception cumulative state field. The state compression coding module is used to start the state compression coding network of the multi-objective reinforcement learning agent, perform multi-level spatial feature abstraction and feature channel dimensionality reduction compression processing on the thermal sensing cumulative state field, and generate a compressed state representation vector suitable for policy optimization. The forward inference processing module is used to perform forward inference based on the compressed state representation vector by the multi-objective policy network of the multi-objective reinforcement learning agent, and output a multi-dimensional control action vector that takes into account both the user thermal comfort optimization objective and the system energy consumption optimization objective. The air conditioning status update module is used to convert the multi-dimensional control action vector into an operating parameter change that the air conditioning equipment can recognize through the actuator action mapping layer and perform an air conditioning operation status update. After the update, the module collects the subjective thermal feedback record of the current user and uses the subjective thermal feedback record to adaptively correct the decay rate constant of the forgetting decay mapping function in the preset somatic forgetting accumulation model.

[0145] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0147] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0148] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0149] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A multi-objective reinforcement learning-based air conditioning control method based on a forgetting model, characterized in that, include: The original environmental physical quantity measurement sets collected by multiple environmental measurement nodes distributed within the target space are processed by time synchronization and spatial field structure reconstruction to generate the current thermal environment state field based on a unified spatial coordinate system. Based on a preset somatosensory forgetting accumulation model, the current thermal environment state field is subjected to perceptual attenuation mapping and cumulative effect superposition processing to obtain the thermal perception accumulation state field. A state compression coding network is initiated for the multi-objective reinforcement learning agent to perform multi-level spatial feature abstraction and feature channel dimensionality reduction compression processing on the thermal sensing cumulative state field, generating a compressed state representation vector suitable for policy optimization. The multi-objective policy network of the multi-objective reinforcement learning agent performs forward inference based on the compressed state representation vector, and outputs a multi-dimensional control action vector that takes into account both the user thermal comfort optimization objective and the system energy consumption optimization objective. The multidimensional control action vector is converted into an operating parameter change that the air conditioning equipment can recognize through the actuator action mapping layer and the air conditioning operating status is updated. After the update, the subjective thermal feedback record of the current user is collected, and the decay rate constant of the forgetting decay mapping function in the preset somatic forgetting accumulation model is adaptively corrected using the subjective thermal feedback record.

2. The method according to claim 1, characterized in that, The process of performing sensory attenuation mapping and cumulative effect superposition on the current thermal environment state field based on a preset somatosensory forgetting accumulation model to obtain a thermal sensory cumulative state field includes: Retrieve the forgetting decay mapping function, the cumulative effect superposition function, and the thermal perception cumulative state field of the previous moment corresponding to the spatial coordinate range of the current thermal environment state field from the preset somatosensory forgetting accumulation model. The current thermal environment state field is analyzed according to the spatial grid cells of a unified spatial coordinate system to obtain the spatial grid cell location identifier and the current thermal environment state vector corresponding to each spatial grid cell location. Based on the spatial grid cell location identifier, extract the previous time sensing cumulative state vector corresponding to the location of each spatial grid cell from the previous time sensing cumulative state field; For each spatial grid cell location, the previously perceived cumulative state vector and the decay rate constant of the forgetting decay mapping function set in the preset somatosensory forgetting cumulative model are input into the forgetting decay mapping function to generate the decayed perceived memory vector for that spatial grid cell location. For each spatial grid cell location, the attenuated sensing memory vector and the current thermal environment state vector are simultaneously input into the cumulative effect superposition function. The cumulative effect superposition function performs element-wise weighted fusion and nonlinear activation on the attenuated sensing memory vector and the current thermal environment state vector to generate the current time-time sensing cumulative state vector for that spatial grid cell location. By combining the current-moment perception cumulative state vectors of all spatial grid cell locations, and reconstructing the spatial field according to the spatial grid cell distribution of the unified spatial coordinate system, a thermal perception cumulative state field is generated that includes complete spatial coverage and carries the current-moment perception cumulative state vector at each spatial grid cell location. The thermal perception cumulative state field is stored as the previous thermal perception cumulative state field for use in the next cycle, so that the preset somatosensory forgetting cumulative model can be continuously iterated and called in subsequent control cycles.

3. The method according to claim 1 or 2, characterized in that, The state compression coding network for initiating the multi-objective reinforcement learning agent performs multi-level spatial feature abstraction and feature channel dimensionality reduction compression processing on the thermal perception accumulated state field to generate a compressed state representation vector suitable for policy optimization, including: The network weight parameters of the state compression coding network are loaded from the network parameter storage unit of the multi-objective reinforcement learning agent. The state compression coding network includes a first spatial feature abstraction layer, a spatial downsampling transition layer, a second spatial feature abstraction layer, a feature channel compression layer, and a fully connected projection layer in series. The current time-based perception cumulative state vector of each spatial grid cell in the thermal sensing cumulative state field is arranged into a multi-channel input feature tensor according to the spatial grid cell distribution. The multi-channel input feature tensor is input into the first spatial feature abstraction layer. The local receptive field coding unit in the first spatial feature abstraction layer performs local spatial pattern extraction and nonlinear mapping on the multi-channel input feature tensor to obtain a primary spatial feature mapping set. The primary spatial feature map set is input into the spatial downsampling transition layer. By performing spatial neighborhood compression and feature channel recombination on the primary spatial feature map set, a spatially compressed intermediate feature map set with reduced spatial resolution and increased number of feature channels is generated. The spatial compression intermediate feature map set is input into the second spatial feature abstraction layer. The extended receptive field coding unit in the second spatial feature abstraction layer performs high-order spatial correlation feature extraction on the spatial compression intermediate feature map set to capture the large-scale spatial context dependency relationship in the thermal perception cumulative state field and generate a high-level spatial feature map set. The high-level spatial feature map set is input into the feature channel compression layer. The channel importance weight parameters stored in the feature channel compression layer are used to perform weighted recombination on each feature channel in the high-level spatial feature map set. The weighted recombination feature channels are then truncated and retained in descending order of channel importance weight to generate a target feature map set with reduced channel dimension. The target feature mapping set is input into the fully connected projection layer. The spatial dimension information and channel dimension information of the target feature mapping set are integrated and mapped into a compressed state representation vector of a set length through the dimension alignment transformation matrix in the fully connected projection layer. Each numerical position in the compressed state representation vector output by the fully connected projection layer corresponds to an environmental perception feature component after spatial feature abstraction and channel compression.

4. The method according to claim 1, characterized in that, The multi-objective policy network of the multi-objective reinforcement learning agent performs forward inference based on the compressed state representation vector, and outputs a multi-dimensional control action vector that takes into account both the user's thermal comfort optimization objective and the system's energy consumption optimization objective, including: The network weight parameters of the multi-objective policy network are loaded from the network parameter storage unit of the multi-objective reinforcement learning agent. The multi-objective policy network includes a shared feature representation layer, a thermal comfort optimization sub-policy branch, an energy consumption optimization sub-policy branch, and a multi-objective action fusion layer. The compressed state representation vector is input into the shared feature representation layer, and the shared feature representation layer performs a nonlinear transformation and feature decoupling on the compressed state representation vector to generate a shared hidden layer feature vector. The shared hidden layer feature vector retains general environmental context information that is related to both the thermal comfort optimization objective and the system energy consumption optimization objective. The shared hidden layer feature vector is simultaneously input into the thermal comfort optimization sub-strategy branch and the energy consumption optimization sub-strategy branch, respectively. In the thermal comfort optimization sub-strategy branch, the shared hidden layer feature vector is refined through the thermal comfort feature transformation sub-layer to obtain the thermal comfort sensitive hidden layer vector. The thermal comfort action generation sub-layer outputs thermal comfort related control action components based on the thermal comfort sensitive hidden layer vector. The thermal comfort related control action components include the temperature setting change direction and the relative intensity of the temperature setting change, as well as the wind speed change direction and the relative intensity of the wind speed change. In the energy consumption optimization sub-strategy branch, the shared hidden layer feature vector is refined by the energy consumption dedicated feature transformation sub-layer to obtain the energy consumption sensitive hidden layer vector. The energy consumption action generation sub-layer outputs energy consumption related control action components based on the energy consumption sensitive hidden layer vector. The energy consumption related control action components include the operating mode adjustment tendency and the compressor frequency adjustment tendency. The thermal comfort-related control action components and the energy consumption-related control action components are simultaneously input into the multi-objective action fusion layer. The thermal comfort-related control action components and the energy consumption-related control action components are weighted and combined using the multi-objective fusion weight parameters stored in the multi-objective action fusion layer. A multi-dimensional control action vector is generated through a fusion activation function, where each dimension of the multi-dimensional control action vector corresponds to an adjustable control variable for air conditioning.

5. The method according to claim 1 or 4, characterized in that, The process of converting the multidimensional control action vector into an operating parameter change that the air conditioning equipment can recognize through the actuator action mapping layer and updating the air conditioning operating status, and then collecting the current user's subjective thermal feedback record after the update, and using the subjective thermal feedback record to adaptively correct the decay rate constant of the forgetting decay mapping function in the preset somatic forgetting accumulation model, includes: The correspondence between the values ​​of each dimension in the multidimensional control action vector and the adjustable control variables of the air conditioner is analyzed, and the multidimensional change values ​​and the operation mode switching instruction codes are separated. Based on the mapping function processing result of the multidimensional changed values ​​and the mode switching control signal corresponding to the operation mode switching instruction code, an air conditioning operation parameter change instruction package is generated and sent to the air conditioning execution controller to perform air conditioning operation status update. Record the completion time of the air conditioner operation status update, and after a preset feedback collection delay time after the completion time, trigger the user feedback collection terminal arranged in the target space to obtain the current user's subjective thermal feedback record. The subjective thermal feedback record includes the user's thermal perception category identifier of the thermal environment at this moment. The subjective thermal feedback record is compared with the current time perception cumulative state vector of the key spatial grid cell in the thermal perception cumulative state field. The direction and degree of thermal deviation are calculated, and a thermal deviation descriptive quantity is generated. The thermal deviation descriptor is input into the attenuation rate constant correction unit. The attenuation rate constant correction unit generates an attenuation rate constant adjustment amount based on the mapping logic between the thermal deviation descriptor and the current value of the attenuation rate constant. The current value of the attenuation rate constant is then superimposed with the attenuation rate constant adjustment amount to obtain the corrected attenuation rate constant. The corrected decay rate constant is written into the preset somatosensory forgetting accumulation model to replace the original decay rate constant, so that it can be called in the next control cycle.

6. The method according to claim 5, characterized in that, The multidimensional change values ​​include temperature setting change values, wind speed change values, and wind direction change values. The mapping function processing result based on the multidimensional change values, and the mode switching control signal corresponding to the operating mode switching instruction code, generate an air conditioning operating parameter change instruction package and send it to the air conditioning execution controller to perform an air conditioning operating status update, including: The temperature setting change value is input into the temperature setting mapping function, which linearly converts the temperature setting change value into a temperature adjustment step that the air conditioner temperature setting parameter register can accept, and then adds it to the current air conditioner temperature setting value to generate the updated temperature setting value. The wind speed change value is input into the wind speed level mapping function. The target wind speed level is determined according to the positive and negative direction and magnitude of the wind speed change value. The target wind speed level is then encoded as an air conditioning fan speed control signal. The wind direction change value is input into the wind direction adjustment mapping function, the target wind direction blade position is determined according to the angle deviation information of the wind direction change value, and the target wind direction blade position is encoded into an air conditioning wind direction control signal. The air conditioner operation mode switching table is searched according to the operation mode switching instruction code, and a mode switching control signal corresponding to the target operation mode is generated. Combining the updated temperature setpoint, air conditioner fan speed control signal, air conditioner air direction control signal, and mode switching control signal, an air conditioner operating parameter change instruction package is generated, and the air conditioner operating parameter change instruction package is sent to the air conditioner execution controller to perform the air conditioner operating status update.

7. The method according to claim 1, characterized in that, The process of performing time-series synchronization and spatial field structure reconstruction on the original environmental physical quantity measurement set collected by multiple environmental measurement nodes distributed within the target space to generate the current thermal environment state field based on a unified spatial coordinate system includes: Receive the raw environmental physical quantity measurement set collected by the multiple environmental measurement nodes, and extract the collection timestamp and node identifier attached to each environmental physical quantity measurement value; All environmental physical quantity measurements in the original environmental physical quantity measurement set are projected onto a unified time reference axis according to the acquisition timestamp. Timestamp interpolation and alignment are performed on the environmental physical quantity measurements with different node identifiers to generate a time-synchronized environmental physical quantity measurement set. Retrieve the preset spatial location distribution map of environmental measurement nodes, and map each environmental physical quantity measurement value in the time-synchronized environmental physical quantity measurement set to the spatial grid node position in the unified spatial coordinate system according to the corresponding node identifier, thereby creating a spatial discrete measurement point matrix. Based on the spatial discrete measurement point array, the environmental physical quantities between adjacent spatial grid nodes are continuously filled with spatial gradients through a pre-stored spatial continuous reconstruction function, generating a current thermal environment state field that covers all spatial grid node positions in the target space. Each spatial grid node position in the current thermal environment state field carries a thermal environment state vector based on a unified spatial coordinate system.

8. The method according to claim 1, characterized in that, After adaptively correcting the decay rate constant of the forgetting decay mapping function in the preset somatosensory forgetting accumulation model using the subjective thermal feedback record, the method further includes: Obtain the accumulated subjective thermal feedback record sequence and the corresponding corrected decay rate constant sequence within multiple control cycles; The thermal sensation category identification conversion relationship between adjacent control cycles in the subjective thermal sensation feedback recording sequence is encoded as thermal sensation state transition feature, and the corrected decay rate constant sequence is divided into decay rate fluctuation segments according to the time window. By inputting the thermal state transition features and the decay rate fluctuation segments into a preset individual perception drift analysis network, an individual perception drift pattern vector characterizing the time-varying characteristics of the user's thermal perception is generated. The individual perception drift pattern vector is context-aligned with the forgetting decay mapping function parameters at the current moment in the preset somatosensory forgetting accumulation model, and the decay rate constant in the forgetting decay mapping function and the fusion weight coefficient in the cumulative effect superposition function are updated synchronously based on the alignment result.

9. A multi-objective reinforcement learning air conditioning control system, characterized in that, include: processor; A storage device storing a computer program, which, when executed by the processor, causes the processor to implement the multi-objective reinforcement learning air conditioning control method based on a forgetting model as described in any one of claims 1-8.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions, which, when executed by a processor, implement the multi-objective reinforcement learning air conditioning control method based on a forgetting model as described in any one of claims 1-8.