Building energy consumption dynamic optimization method and system based on BIM and reinforcement learning
By using a BIM-based and reinforcement learning approach, a monitoring grid and dimming loop were established, a light transmission correlation model was generated and converted to the spectral domain, key data were screened, a differentiable digital twin predictor was established, and a reinforcement learning network was driven to generate dimming control commands. This solved the lighting control problem caused by the complexity of the light environment in large-span steel structure buildings, and achieved energy consumption optimization and improved illuminance uniformity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU YUANLUO CONSTRUCTION ENGINEERING CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-19
AI Technical Summary
Existing lighting control systems are ill-suited to the complexity and dynamism of the lighting environment in large-span steel structure industrial buildings, leading to frequent adjustments in lamp brightness and a decline in lighting quality, and failing to achieve high-precision energy consumption optimization control.
Based on BIM and reinforcement learning, an initial light transmission correlation model is generated by establishing a monitoring grid and dimming loop with spatial alignment of structural span modulus. This model is then converted to a spatial distribution spectral domain, key spectral domain data is screened, a differentiable digital twin predictor is established, and a reinforcement learning policy network is driven to generate dimming control commands. Parameters are updated by combining energy consumption cost and illuminance deviation sensitivity.
It effectively eliminates random environmental interference, accurately locks structured light field fluctuations, eliminates control oscillations, and achieves deep energy savings and improved illuminance uniformity in lighting.
Smart Images

Figure CN122065387A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent building control technology, and more specifically, to a method and system for dynamic optimization of building energy consumption based on BIM and reinforcement learning. Background Technology
[0002] In the field of modern industrial construction, large-span steel portal frame systems are widely used in large factories and high-bay warehouses due to their rapid construction speed and high space utilization. To reduce lighting energy consumption, these buildings typically incorporate skylights in the roof and supplement lighting with artificial lighting systems. Traditional energy-saving lighting control mainly relies on light sensors combined with PID closed-loop or time-based static strategies, aiming to adjust the brightness of indoor luminaires according to changes in outdoor natural light. However, in the actual operating environment of steel structure industrial spaces, changes in the lighting environment are highly complex and dynamic, making it difficult for existing general control methods to adapt. This is mainly reflected in the following aspects: First, steel roof structures are extremely sensitive to temperature changes. Under solar radiation, large-span steel components generate a non-uniform temperature field distribution, causing thermal expansion and contraction and minute deformation of the roof structure. Although this structural deformation is small, it can alter the light transmission path and shading relationship between lighting fixtures and the work surface under the long optical path projection in a high-ceilinged space.
[0003] Secondly, industrial plants are typically equipped with large mobile equipment such as overhead cranes. The reciprocating movement of these cranes on the tracks creates a huge mobile obstruction, which not only intermittently blocks skylights but also dynamically alters the reflection and obstruction conditions of artificial lighting.
[0004] Existing lighting control systems typically treat light field fluctuations caused by structural thermal deformation and equipment movement as random noise or simple environmental disturbances. General reinforcement learning algorithms or feedback control logic lack an understanding of the intrinsic properties of the building structure and cannot distinguish between structural light field modulation and random environmental noise. When faced with these complex non-stationary changes, the control system often misjudges, leading to frequent adjustments in lamp brightness, control oscillations, or failure to respond promptly to local illuminance imbalances, resulting in energy waste and decreased lighting quality. Furthermore, existing digital twin technologies are mostly based on static geometric models, making it difficult to map the real-time morphological changes of steel structures under thermal effects and the high-frequency movement of vehicles. This results in discrepancies between the computational model and the real environment, hindering the support of high-precision energy consumption optimization control. Summary of the Invention
[0005] This invention provides a method and system for dynamic optimization of building energy consumption based on BIM and reinforcement learning, which solves the technical problems mentioned in the background.
[0006] The first aspect is a dynamic optimization method for building energy consumption based on BIM and reinforcement learning, including: Based on the building information model, a monitoring grid and dimming loop are established that are spatially aligned with the structural span modulus of the steel portal frame, and an initial optical transmission correlation model is generated. The initial optical transmission correlation model is transformed into the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermal deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the runway position, key spectral domain data are screened and assembled into a structured sparse spectral state set. A differentiable digital twin predictor is established based on the structured sparse spectral state set, and the spectral intensity data in the structured sparse spectral state set is mapped to the state input vector of reinforcement learning. The reinforcement learning policy network is driven to generate dimming control commands in response to the state input vector, and the parameters of the reinforcement learning policy network are updated based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor.
[0007] Secondly, a building energy consumption dynamic optimization system based on BIM and reinforcement learning, applied to any of the aforementioned building energy consumption dynamic optimization methods based on BIM and reinforcement learning, includes: The initial modeling module is used to establish a monitoring grid and dimming loop that are aligned with the structural span modulus space of the steel portal frame based on the building information model, and to generate an initial optical transmission correlation model. The spectral domain feature extraction module is used to convert the initial optical transmission correlation model to the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermal deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the vehicle runway position, the module filters key spectral domain data and assembles them into a structured sparse spectral state set. The digital twin and state mapping module is used to establish a differentiable digital twin predictor based on the structured sparse spectral state set, and to map the spectral intensity data in the structured sparse spectral state set into a state input vector for reinforcement learning. The reinforcement learning control and update module is used to drive the reinforcement learning policy network to generate dimming control commands in response to the state input vector, and to perform parameter updates of the reinforcement learning policy network based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor.
[0008] The beneficial effects of this invention include: by transforming the initial light transmission correlation model to the spatial distribution spectral domain, the light field data is sparsified using the inherent cross-frame repetition pattern of the steel structure portal frame, the roof thermal deformation distribution mode, and the displacement parameters of the vehicle movement, effectively eliminating random environmental interference and accurately locking the structured light field fluctuations caused by the operation of the structure and equipment; at the same time, the reinforcement learning strategy network is updated in a targeted manner by combining the illuminance deviation sensitivity provided by the differentiable digital twin predictor, thereby eliminating the control oscillation phenomenon caused by the inability to identify structured fluctuations in traditional lighting control, and realizing a deep saving of lighting energy consumption and a significant improvement in illuminance uniformity in steel structure industrial buildings under complex working conditions. Attached Figure Description
[0009] Figure 1 This is a flowchart of the dynamic optimization method for building energy consumption based on BIM and reinforcement learning of the present invention; Figure 2 This is a schematic diagram illustrating a specific implementation of the present invention. Detailed Implementation
[0010] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0011] Example 1: As Figure 1 As shown, the dynamic optimization method for building energy consumption based on BIM and reinforcement learning includes: Based on the building information model, a monitoring grid and dimming loop are established that are spatially aligned with the structural span modulus of the steel portal frame, and an initial optical transmission correlation model is generated. The initial optical transmission correlation model is transformed into the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermal deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the runway position, key spectral domain data are screened and assembled into a structured sparse spectral state set. A differentiable digital twin predictor is established based on the structured sparse spectral state set, and the spectral intensity data in the structured sparse spectral state set is mapped to the state input vector of reinforcement learning. The reinforcement learning policy network is driven to generate dimming control commands in response to the state input vector, and the parameters of the reinforcement learning policy network are updated based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor.
[0012] Preferably, a monitoring grid and dimming loop are established based on the building information model and aligned with the structural span modulus of the steel portal frame, and an initial optical transmission correlation model is generated, including: The structural span modulus of the steel portal frame in the building information model is analyzed, and a monitoring grid is divided along the length of the building. The illuminance monitoring point data in the monitoring grid is serialized into an illuminance observation sequence with the dimension of the total number of monitoring points, and the control signal of the dimming loop is used as the dimming input vector with the dimension of the total number of loops. Read the standard photometric distribution data file of the luminaires in the dimming loop, and calculate the correlation coupling coefficient between the nth illuminance monitoring point and the kth dimming loop in the initial light transmission correlation model using the following formula: in, The coefficient represents the correlation coupling coefficient in the nth row and kth column of the initial optical transmission correlation model, and represents the contribution weight of the kth dimming loop to the direct illuminance of the nth illuminance monitoring point. This represents the set of all physical luminaires belonging to the k-th dimming loop; Represents a set The first in lamp fixtures; This indicates the first value obtained from the standard photometric distribution data file. The luminous intensity value of a lamp in the direction pointing towards the nth illuminance monitoring point; Indicates the first The vertical luminous intensity angle of the light ray from one lamp to the nth illuminance monitoring point in the local coordinate system of the lamp; Indicates the first The horizontal luminous intensity angle of the light ray from one lamp to the nth illuminance monitoring point in the local coordinate system of the lamp; Indicates the first The angle of incidence between the incident ray from one lamp toward the nth illuminance monitoring point and the normal to the working surface of the monitoring grid; Indicates the first The Euclidean spatial distance between the optical center of a lamp and the nth illuminance monitoring point; This indicates the maximum value operation, used to eliminate the influence of backlighting.
[0013] The structural span modulus of a steel portal frame is the distance between two adjacent frames, which can be obtained by analyzing the structural design data in the Building Information Model (BIM). The monitoring grid is a two-dimensional grid aligned along the building's length using the structural span modulus as a reference, used to arrange illuminance monitoring points. Illuminance monitoring point data consists of the illuminance intensity data collected at each monitoring point on the monitoring grid, acquired by illuminance sensors installed at the grid nodes. The illuminance observation sequence is a one-dimensional data sequence formed by sequentially arranging the illuminance monitoring point data on the two-dimensional monitoring grid in a preset order. A dimming circuit is an independent control unit formed by aggregating multiple luminaires according to the electrical circuit design, used to uniformly adjust the brightness of its respective luminaires. The standard photometric distribution data file is a standardized document recording the photometric characteristics of luminaires, such as luminous intensity in different directions, which can be obtained from the product technical data provided by the luminaire manufacturer. Luminous intensity distribution data is the luminous intensity value of the luminaire in each direction recorded in the standard photometric distribution data file. Spatial distance is the straight-line distance between the optical center of the luminaire and the illuminance monitoring point. The incident angle is the angle between the light ray from the luminaire to the illuminance monitoring point and the normal to the working surface of the monitoring grid. The single-lamp illuminance contribution value is the illuminance intensity contribution value of a single luminaire to a given illuminance monitoring point. The correlation coupling coefficient is the result of the linear superposition of the single-lamp illuminance contribution values of all luminaires in the same dimming loop to a given illuminance monitoring point. The initial light transmission correlation model is a matrix composed of the correlation coupling coefficients between all dimming loops and all illuminance monitoring points, used to characterize the light transmission relationship between the dimming loops and the illuminance monitoring points. The nth illuminance monitoring point is the illuminance acquisition point numbered n in a preset order in the monitoring grid; n is a positive integer. The kth dimming loop is an independent controlled luminaire unit numbered k according to a preset rule; k is a positive integer. The set of luminaires belonging to the kth dimming loop is the sum of all physical luminaires assigned to and uniformly controlled by the kth dimming loop. The ℓth luminaire is the luminaire numbered ℓ in a preset order within the set of luminaires belonging to the kth dimming loop; ℓ is a positive integer. The vertical luminous intensity angle is the angle between the ray of light from the ℓth luminaire to the nth illuminance monitoring point and the luminaire's vertical direction in the luminaire's local coordinate system. The horizontal luminous intensity angle is the angle between the ray of light from the ℓth luminaire to the nth illuminance monitoring point and the luminaire's horizontal reference direction in the luminaire's local coordinate system. The Euclidean spatial distance is the straight-line distance between the optical center of the ℓth luminaire and the nth illuminance monitoring point, calculated using the Euclidean distance formula.
[0014] Cross-module alignment for monitoring grid division and data serialization: First, the structural span module in the BIM is analyzed. For example, a span of 6 meters is divided into monitoring grids with a grid spacing of 0.5 meters, each 6 meters along the building length direction, ensuring uniform grid distribution within each span unit. Data serialization adopts a column-priority order, that is, the data of all monitoring points in the lateral direction within the first span unit are arranged sequentially, then the data of the second span unit is arranged, and so on, to form an illumination observation sequence, ensuring that the data corresponds to the structural span module, which facilitates subsequent spectral domain analysis.
[0015] Lighting fixtures are grouped into dimming circuits and models are constructed as follows: Based on the electrical circuit design, lighting fixtures driven by the same controller output port are grouped into one dimming circuit. For example, if a factory has 20 lighting fixtures driven by 5 controllers, it is divided into 5 dimming circuits. The correlation coupling coefficient of each circuit is obtained by superimposing the contribution values of all lighting fixtures in the circuit to the individual lamps at each monitoring point, so that the model can accurately reflect the overall impact of the circuit on illumination and avoid the complexity of individual lamp control.
[0016] Calculation of single-lamp illuminance contribution: Combining the luminous intensity value in the standard photometric distribution data, for example, a lamp with a luminous intensity of 5000 candela in a certain direction, a distance of 20 meters from the lamp to the monitoring point, and an incident angle of 30 degrees, first calculate the square of the distance as 400, the cosine value of which is 0.866. Then, calculate the single-lamp contribution value according to the formula: 5000 × 0.866 ÷ 400 ≈ 10.83 lux. This calculation method considers the effects of luminous intensity, distance attenuation, and incident angle to ensure accurate calculation of illuminance contribution.
[0017] Standard photometric distribution data file: Utilizing the IESLM-63 format, an internationally recognized standard for luminaire photometric data. During parsing, luminous intensity data corresponding to vertical and horizontal angles is extracted. For example, vertical angles range from 0 to 180 degrees, and horizontal angles from 0 to 360 degrees, with bilinear interpolation used to obtain luminous intensity values in any direction.
[0018] Data serialization order: column first. For example, if the monitoring grid is 10 rows and 20 columns, first arrange the data of the 10 monitoring points in the first column, then arrange the data of the second column, and so on up to the 20th column.
[0019] Lighting aggregation standard: According to the electrical circuit design, lighting fixtures driven by the same DALI group, 0-10V circuit or the same gateway channel constitute a dimming circuit. For example, if channel 1 of a certain DALI controller drives 4 lighting fixtures, then these 4 lighting fixtures constitute a dimming circuit, ensuring that the circuit division is consistent with the actual control logic.
[0020] Angle transformation rule: Transform the light direction vector from the building's global coordinate system to the luminaire's local coordinate system. First, eliminate the influence of the luminaire's installation angle by using a rotation matrix. For example, if the luminaire is installed at a 30-degree angle around the z-axis, rotate the light direction vector in the global coordinate system by -30 degrees, and then calculate the vertical and horizontal luminous angles.
[0021] The max function is used when the cosine of the angle between the incident ray and the normal to the working surface is negative, indicating that the ray is facing away from the working surface and cannot produce effective illumination. In this case, the value is set to 0. For example, if the angle is 100 degrees and the cosine value is -0.1736, then the value is set to 0 to eliminate the influence of invalid rays.
[0022] Preferably, transforming the initial optical transmission correlation model to the spatial distribution spectral domain includes: We construct longitudinal spatial transformation components corresponding to the building's length and lateral spatial transformation components corresponding to the building's span, and then construct a global two-dimensional spatial orthogonal transformation operator using the Kronecker product operation. The calculation formula is as follows: The initial optical transmission correlation model is transformed using the global two-dimensional spatial orthogonal transformation operator to obtain the global spatial distribution spectrum model, and the calculation formula is as follows: in, This represents a global two-dimensional spatial orthogonal transformation operator used to map the illuminance response of physical space to the spatial distribution spectral domain. The vertical spatial transformation component is represented by a square matrix whose dimension is the total number of vertical grid nodes; The component representing the horizontal spatial transformation is a square matrix whose dimension is the total number of horizontal grid nodes; This represents the Kronecker product operator, used to extend two low-dimensional transformation components into a high-dimensional transformation operator; This represents the vertical spatial frequency index, with a value ranging from zero to the total number of vertical grid nodes minus one. This represents the vertical physical space grid index, with a value ranging from zero to the total number of vertical grid nodes minus one. This represents the total number of vertical grid nodes, corresponding to the number of discrete sampling points along the length of the steel portal frame. This represents the horizontal spatial frequency index, with a value ranging from zero to the total number of horizontal grid nodes minus one. In exponential functions, it represents the imaginary unit; in indexes, it represents the horizontal physical space grid index. This represents the total number of horizontal grid nodes, corresponding to the number of discrete sampling points along the horizontal span direction of the steel portal frame. Pi is a constant. This represents a global spatial distribution spectrum model, which includes the optical transmission complex spectral coefficients across the entire frequency domain; The initial optical transmission correlation model represents the optical transmission coupling relationship in physical space.
[0023] The total number of longitudinal grid nodes is the number of monitoring grid nodes divided along the length of the steel portal frame. The total number of transverse grid nodes is the number of monitoring grid nodes divided along the transverse span of the steel portal frame. The longitudinal spatial transformation component is a square matrix constructed based on the total number of longitudinal grid nodes, consisting of a sequence of discrete complex exponential functions, used to achieve spatial transformation along the length direction. The transverse spatial transformation component is a square matrix constructed based on the total number of transverse grid nodes, consisting of a sequence of discrete complex exponential functions, used to achieve spatial transformation along the transverse span direction. The unit orthogonality normalization rule is a standardization rule used to construct the sequence of discrete complex exponential functions. It is preferably normalized according to the total number of nodes in the direction corresponding to the square root to ensure the orthogonality of the transformation operator and not change the energy distribution of the signal. The discrete complex exponential function sequence is the basic function sequence constituting the longitudinal and transverse spatial transformation components, and each function satisfies the unit orthogonality normalization rule. The global two-dimensional spatial orthogonal transformation operator is a high-dimensional operator obtained by performing the Kronecker product operation on the transverse and longitudinal spatial transformation components, used to achieve the conversion from physical space to the spectral domain. The initial optical transmission correlation model is a matrix characterizing the optical transmission coupling relationship between the down-adjustment optical loop in physical space and the illuminance monitoring point. The global spatial distribution spectrum model is the spectral domain model obtained after transforming the initial optical transmission correlation model using a global two-dimensional spatial orthogonal transformation operator, containing complex amplitude and phase information across the entire frequency domain. The complex amplitude is the magnitude composed of the real and imaginary complex parts corresponding to each frequency point in the global spatial distribution spectrum model. The phase information is the complex phase angle corresponding to each frequency point in the global spatial distribution spectrum model. The vertical spatial frequency index is used to identify different frequency components in the vertical spatial transformation component, with values ranging from zero to the total number of vertical grid nodes minus one. The vertical physical space grid index is used to identify the location of vertical physical space monitoring grid nodes, with values ranging from zero to the total number of vertical grid nodes minus one. The horizontal spatial frequency index is used to identify different frequency components in the horizontal spatial transformation component, with values ranging from zero to the total number of horizontal grid nodes minus one. The horizontal physical space grid index is used to identify the location of nodes in the horizontal physical space monitoring grid, and its value ranges from zero to the total number of horizontal grid nodes minus one. The pi constant is a fixed constant with a value of approximately 3.1416.
[0024] Construction of spatial transformation components: Both longitudinal and transverse spatial transformation components are constructed from a sequence of orthogonally normalized discrete complex exponential functions. Specifically, the exponent of each function is negative 2 multiplied by pi, multiplied by the frequency index, multiplied by the grid index, and then divided by the total number of nodes in the corresponding direction. For example, if the total number of grid nodes in the longitudinal direction is 100, the longitudinal spatial frequency index is 5, and the longitudinal physical space grid index is 20, then the value of the discrete complex exponential function is 1 divided by the square root of 100, multiplied by the negative j of the natural exponent, multiplied by 2, multiplied by 3.1416, multiplied by 5, multiplied by 20, and then divided by 100, ensuring that the transformation components are orthogonal.
[0025] Generation of global two-dimensional orthogonal transformation operators: The Kronecker product operation expands the two low-dimensional transformation components, horizontal and vertical, into high-dimensional operators. Specifically, each element of the horizontal transformation component is multiplied by the vertical transformation component using a matrix multiplication, and then the elements are arranged and combined in order. For example, if the horizontal transformation component is a 20x20 matrix and the vertical component is a 100x100 matrix, the Kronecker product yields a 2000x2000 global operator, achieving a global transformation in two-dimensional space.
[0026] Transformation from physical space to spectral domain: The initial optical transmission correlation model is linearly multiplied by a global two-dimensional spatial orthogonal transformation operator. That is, each row of the operator matrix is multiplied by each column of the initial model and then summed. This transforms the complex optical transmission coupling coefficients in physical space into complex amplitude and phase information in the spectral domain, thereby enabling the extraction of structured features.
[0027] Construction rules for discrete complex exponential function sequences: Each function follows unit orthogonal normalization, with the numerator fixed at 1, the denominator being the total number of nodes in the corresponding direction under the square root, and the exponent being negative 2 multiplied by pi multiplied by the frequency index multiplied by the grid index, and then divided by the total number of nodes in the corresponding direction, ensuring that each function is orthogonal to each other, and that the signal energy is conserved after transformation.
[0028] The storage association between complex amplitude and phase information: The complex information of each frequency point is stored as real part and imaginary part. The complex amplitude is calculated by taking the square root of the sum of the square of the real part and the square of the imaginary part. The phase information is calculated by the arctangent of the imaginary part and the real part. The two are stored in a one-to-one correspondence to ensure the integrity of the spectral domain information.
[0029] Preferably, the initial optical transmission correlation model is transformed into the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermally induced deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the runway position, key spectral domain data are screened and assembled into a structured sparse spectral state set, including: Calculate the cross-frame repetition frequency index and define a key spectral domain screening index set. The calculation formula is as follows: The vehicle motion displacement parameters are calculated based on the location of the running track, using the following formula: The structured sparse spectral state set is assembled using the following calculation formula: in, This indicates the structural span modulus of a steel portal frame. This indicates the longitudinal grid spacing of the monitoring grid along the length of the building; This indicates the rounding operation; This represents the discrete grid period corresponding to the structural span modulus; Indicates the total number of vertical grid nodes; The cross-frame repetition frequency index represents the dominant frequency position of the structure's periodicity in the spectral domain; This represents the set of frequency points in the sideband of vehicle movement, including the adjacent frequency points on both sides of the main frequency and the harmonics; This represents the set of frequency points for the distribution modes of thermally induced deformation, which consists of high-energy characteristic frequencies determined by offline simulation. This represents the key spectral domain filtering index set, which is the union of all the feature frequency points that need to be tracked; This represents the physical coordinates of the train on the runway at time t; This represents the normalized position of the vehicle at time t relative to the structural span modulus. This indicates the frequency index. The vehicle motion displacement parameters represent the phase rotation introduced by the position movement; Represents the imaginary unit; The structured sparse spectral state set refers to low-dimensional spectral data after feature filtering. Indicates based on set The constructed filtering operator is used to extract rows at specified indices from the full-frequency domain matrix; This represents the global spatial distribution spectrum model of the initial optical transmission correlation model; This represents the total number of thermally induced deformation distribution modes; The modal coordinate amplitude represents the r-th order thermally induced deformation distribution mode at time t; This represents a pre-set r-th order thermally induced deformation distribution mode template, characterizing the spectral domain perturbation texture under unit mode amplitude; This indicates the operation of taking the real part; This indicates the preset frequency index. The sideband spectrum template for vehicle motion is used to characterize the spectral domain sideband response under unit phase change.
[0030] The longitudinal grid spacing is the distance between adjacent nodes along the building's length, preferably 0.5 meters, a commonly used grid spacing in industrial scenarios, balancing monitoring accuracy and computational efficiency. The discrete period is the ratio of the structural span modulus of the steel portal frame to the longitudinal grid spacing, rounded to the nearest integer. The span repetition frequency index is the ratio of the total number of longitudinal grid nodes to the discrete period, rounded to the nearest integer, characterizing the dominant frequency position of the structure's periodicity in the spectral domain. The basic structure frequency set consists of the frequency points corresponding to the span repetition frequency index and its second harmonic index. The vehicle motion sideband frequency set is formed by offsetting adjacent frequency points from the basic structure frequency set. The thermo-deformation distribution mode frequency set consists of the high-energy characteristic frequencies determined by offline thermo-deformation simulation. The key spectral domain screening index set is the union of the basic structure frequency set, the vehicle motion sideband frequency set, and the thermo-deformation distribution mode frequency set. The physical coordinates of the train are the actual position coordinates of the train on the runway at time t, which can be acquired by the train encoder or programmable logic controller. The relative position of the train is the normalized value of the physical coordinates of the train relative to the span modulus of the steel portal frame structure. The base of the natural logarithm is a fixed constant in mathematics, with a value of approximately 2.7183. The imaginary unit is the unit used in mathematics to represent complex numbers, denoted as j. The train motion displacement parameter is a complex exponential function value with the base of the natural logarithm, the exponent of the product of the imaginary unit, the relative position of the train, and the frequency index, representing the phase rotation introduced by the position movement. The thermo-deformation distribution modal spectrum template is a preset template data representing the spectral domain perturbation texture under the unit modal amplitude. The train motion sideband spectrum template is a preset template data representing the spectral domain sideband response under the unit phase change. The real-time thermo-deformation modal coordinates are the amplitude of the r-th order thermo-deformation distribution mode at time t, which can be calculated by acquiring data from temperature sensors installed at key locations on the roof. The thermally induced deformation perturbation term is the product of the real-time thermally induced deformation modal coordinates and the thermally induced deformation distribution mode spectral template. The vehicle motion perturbation term is the product of the real part of the vehicle motion displacement parameter and the vehicle motion sideband spectral template. The global spatial distribution spectral model is the spectral domain model obtained after transforming the initial optical transmission correlation model through a global two-dimensional spatial orthogonal transformation operator. The filtering operator is a matrix constructed based on the key spectral domain filtering index set to extract specified index rows from the full-frequency domain matrix. The structured sparse spectral state set is the low-dimensional spectral data obtained by dimensionality reduction extraction after linearly superimposing the global spatial distribution spectral model and the two types of perturbation terms through the filtering operator. The total number of thermally induced deformation distribution modes is a preset number, preferably 3, as the main influence dimension of structural thermal deformation can usually be covered by 3rd-order modes. The frequency index is a parameter that takes the value of the cross-frame repetition frequency index or its second harmonic index.
[0031] Calculation of the span repetition frequency index and discrete period: First, obtain the structural span modulus of the steel portal frame, for example, 6 meters, with a longitudinal grid spacing of 0.5 meters. The ratio of the two is 12, which is rounded to obtain the discrete period of 12. If the total number of longitudinal grid nodes is 120, the ratio of the total number of longitudinal grid nodes to the discrete period is 10, which is rounded to obtain the span repetition frequency index of 10. This calculation directly binds to the span modulus characteristics of the steel structure, providing a structural basis for subsequent spectral domain selection.
[0032] Construction of the key spectral domain screening index set: The basic structure frequency point set consists of the frequency points corresponding to the cross-frame repetition frequency index 10 and its second harmonic 20, i.e., ±10 and ±20. The vehicle motion sideband frequency point set is obtained by offsetting the basic frequency points adjacently, resulting in ±9, ±11, ±19, and ±21. The thermally induced deformation distribution mode frequency point set consists of three high-energy frequency points obtained from offline simulation. The union of these three sets constitutes the key spectral domain screening index set, enabling precise locking of structured feature frequency points.
[0033] Construction of vehicle motion displacement parameters: Divide the vehicle's physical coordinates (e.g., 3 meters) by the structural span modulus (6 meters) to obtain a normalized relative vehicle position of 0.5. When the frequency index is 10, the exponent of the complex exponential function is the imaginary unit multiplied by 2, pi, 10, and 0.5, then divided by the total number of vertical grid nodes (120). The calculated complex exponential function value is the vehicle motion displacement parameter, accurately representing the phase change caused by vehicle movement.
[0034] Assembly of two types of perturbation terms and a structured sparse spectral state set: A pre-set thermo-deformation distribution mode spectrum template is called, and the real-time thermo-deformation mode coordinates 0.8 are multiplied by the template to obtain the thermo-deformation perturbation term. The real part 0.3 of the vehicle motion displacement parameter is multiplied by the vehicle motion sideband spectrum template to obtain the vehicle motion perturbation term. The global spatial distribution spectrum model is linearly superimposed with the two types of perturbation terms, and then a filtering operator is used to extract key spectral domain data to obtain a structured sparse spectral state set, achieving accurate characterization of structural fluctuations.
[0035] Offline thermal deformation simulation method: Finite element simulation software is used to input the geometric parameters and thermophysical properties of the steel portal frame, and set boundary conditions such as solar radiation intensity of 800 watts per square meter and ambient temperature of 25 degrees Celsius. The temperature field distribution of the roof at different times is obtained through simulation, and high-energy characteristic frequency points are extracted through modal analysis to form a set of thermal deformation distribution modal frequency points.
[0036] High-energy characteristic frequency point selection criteria: Calculate the energy percentage of each frequency point, select the frequency points whose energy accounts for more than 80% of the total energy, and if the number of frequency points exceeds 5, select the 5 with the highest energy to ensure that the set size is appropriate and covers the main thermal deformation characteristics.
[0037] The filtering operator construction logic is as follows: The filtering operator is a 0-1 matrix with the number of rows equal to the number of frequency points in the key spectral domain filtering index set and the number of columns equal to the total number of frequency points in the entire frequency domain. The column position corresponding to the frequency point in the key spectral domain filtering index set is set to 1, and the other column positions are set to 0, so as to achieve accurate extraction of the specified frequency points.
[0038] Preset spectral template generation process: The thermo-deformation distribution mode spectral template is generated through a unit thermal mode perturbation experiment. A unit thermal deformation is applied to the roof, the change in the optical transmission matrix is measured and converted to the spectral domain to obtain template data. The vehicle motion sideband spectral template is generated through vehicle position movement simulation. The occlusion effect of the vehicle at different positions is simulated, the spectral domain sideband response is calculated and normalized to obtain template data.
[0039] The calculation priority of the complex exponential function is as follows: first calculate the product of the imaginary unit and the relative position of the vehicle, then multiply by the frequency index, and finally divide by the total number of vertical grid nodes to obtain the exponential part, ensuring consistent calculation logic.
[0040] Linear superposition dimension matching rule: The matrix dimensions of the global spatial distribution spectrum model, thermal deformation disturbance term, and vehicle motion disturbance term are all unified as the number of rows in the full frequency domain multiplied by the number of dimming loops to ensure that the linear superposition operation can be executed.
[0041] Preferably, establishing a differentiable digital twin predictor based on the structured sparse spectral state set includes: The real-time restored optical transmission coupling model is recovered using inverse transform, and the calculation formula is as follows: The external solar field distribution vector is calculated using the following formula: The prediction equation for the differentiable digital twin predictor is constructed, and the calculation formula is as follows: in, The transpose of the filtering operator is used to map the low-dimensional structured sparse spectral state set back to the full-dimensional spectral domain space, with unfiltered frequency points automatically padded with zeros. This represents the structured sparse spectral state set; This represents the full-size spectral domain data after dimensionality upsizing and filling. The inverse operator of the global two-dimensional space orthogonal transformation operator is used to perform the two-dimensional discrete Fourier inverse transform; This indicates the operation of taking the real part, used to eliminate the imaginary noise generated in numerical calculations; The real-time restoration optical transmission coupling model represents the relationship between thermal deformation and physical space optical transmission after vehicle obstruction at time t. This represents a preset solar response basis function matrix, where each column corresponds to an indoor illuminance response distribution under a standard general sky model. The vector of sky brightness weighting coefficients at time t is a combination of coefficients obtained by fitting outdoor observation data to the International Commission on Illumination standard general sky model using the least squares method. The vector representing the external sunlight field distribution characterizes the indoor illuminance distribution introduced by pure sunlight at time t. This represents the dimming control command vector, which is the output action of the reinforcement learning policy network; The digital twin predicted illuminance vector is the estimate of the working surface illuminance by the differentiable digital twin predictor at the next moment.
[0042] The transpose matrix of the screening operator is constructed based on the key spectral domain screening index set. It is used to upscale the low-dimensional structured sparse spectral state set to full-size spectral domain data. The structured sparse spectral state set is a low-dimensional spectral model containing key spectral domain data after dimensionality reduction extraction by the screening operator. The full-size spectral domain data is the complete spectral domain data recovered after dimensionality upscaling and padding of the structured sparse spectral state set, with zeros automatically padded at unscreened frequency points. The inverse operator of the global two-dimensional spatial orthogonal transformation operator is the inverse matrix of the global two-dimensional spatial orthogonal transformation operator, used to inversely transform the spectral domain data back to physical space. The real-time restored optical transmission coupling model is a matrix obtained by taking the real part of the full-size spectral domain data after inverse transformation, representing the physical space optical transmission relationship at time t considering thermal deformation and vehicle obstruction. Outdoor sky brightness observation data is brightness measurement data from different sky directions outdoors, which can be collected by a fisheye camera or sky brightness scanner installed in an unobstructed outdoor location. The sky brightness weighting coefficient vector is a combined coefficient vector obtained by fitting outdoor sky brightness observation data to the International Commission on Illumination (ICI) standard general sky model using the least squares method. The daylight response basis function matrix is a pre-defined matrix representing the indoor illuminance response distribution under different standard general sky models. It is preferably a matrix with 15 columns and rows equal to the total number of monitoring grid nodes, as the ICI standard general sky model includes 15 sky types, comprehensively covering various daylight environments. The external daylight field distribution vector is the product of the sky brightness weighting coefficient vector and the daylight response basis function matrix, representing the indoor illuminance distribution introduced by pure sunlight at time t. The real-time restored light transmission coupling model is the light transmission coupling matrix between the dimming loops and illuminance monitoring points in the physical space, considering thermal deformation and vehicle obstruction. The dimming control command vector is the command vector output by the reinforcement learning policy network for adjusting the brightness of each dimming loop. The digital twin predicted illuminance vector is the predicted illuminance value obtained by multiplying the real-time restored light transmission coupling model and the dimming control command vector, and then superimposing the external daylight field distribution vector. This vector is completely differentiable with respect to the dimming control command vector. The preset target illuminance vector is a vector with a target illuminance value of 300 lux for each monitoring point on the pre-set working surface grid, based on the illuminance requirements of routine operations in industrial plants.
[0043] Dimensional Upsizing and Inverse Transformation Recovery: The transpose of the filtering operator can map low-dimensional sparse spectral data back to the full-size spectral domain. For example, if the filtering operator is 10 rows and 100 columns (corresponding to 10 key frequency points), its transpose is 100 rows and 10 columns. Multiplying this by a 10-row, K-column structured sparse spectral state set yields 100-row, K-column full-size spectral domain data, with the 90 unfiltered frequency point columns automatically padded with zeros. Then, the inverse transformation is performed using the inverse operator of the global two-dimensional spatial orthogonal transformation operator, converting the spectral domain data back to physical space. The real part is taken to remove the imaginary noise generated by numerical calculations, resulting in a real-time restored optical transmission coupling model, ensuring accurate restoration of optical transmission relationships.
[0044] Standardized Generation of External Sunlight Field: The International Commission on Illumination (ICI) standard general sky model includes 15 sky types, covering various sunlight environments from cloudy to sunny days. Brightness data from 500 different directions were collected outdoors using a fisheye camera. A design matrix was constructed, and a 15-dimensional sky brightness weighting coefficient vector was obtained by fitting the matrix using the least squares method. This vector was multiplied by a 15-column sunlight response basis function matrix, generating an external sunlight field distribution vector that can standardize and characterize the real-time sunlight contribution, avoiding random interference caused by sunlight fluctuations.
[0045] Construction of the differentiable predictor: The calculation process of the digital twin predicted illuminance vector is a linear operation superposition. The matrix multiplication of the real-time restored light transmission coupling model and the dimming control command vector, as well as the superposition of the external sunlight field distribution vector, all satisfy the differentiability property, making the predicted illuminance vector differentiable with respect to the dimming control command vector throughout the process. This provides key support for the gradient update of the subsequent reinforcement learning policy network.
[0046] The fitting method for the sky brightness weight coefficient vector is as follows: Each row of the design matrix corresponds to an outdoor sky sampling direction, and each column corresponds to the relative brightness distribution of a standard general sky model, totaling 500 rows and 15 columns. Using the brightness values of 500 directions obtained from outdoor observations as the dependent variable, the product of the design matrix and the weight coefficient vector is solved using the least squares method to minimize the sum of squared errors between the result and the observed values. Simultaneously, a constraint is added that the weight coefficients are non-negative and their sum is 1 to ensure that the fitting result conforms to physical meaning.
[0047] The pre-set process for the daylight response basis function matrix is as follows: Using daylighting simulation software, input the building's geometric parameters, skylight dimensions and transmittance, internal surface reflectance (wall 0.5, ground 0.2, roof 0.7), and geometric data of the roof shading components. Daylighting simulations are performed on 15 types of standard general sky models. For each type of sky, output a vector containing the illuminance values of all monitoring points. Arrange the 15 vectors column-wise to form the daylight response basis function matrix.
[0048] The accuracy control method of inverse transform: The inverse fast Fourier transform algorithm is used to perform the inverse transform of the global two-dimensional space orthogonal transform operator. The numerical accuracy threshold is set to 1e-6. When the absolute value of the imaginary part of the complex number after inverse transform is less than the threshold, the real part is directly taken; when it is greater than the threshold, the real part is taken after linear interpolation correction, so as to ensure the numerical stability of the optical transmission coupling model in real time.
[0049] The dimension matching rule between the real-time restored optical transmission coupling model and the dimming control command vector is as follows: the real-time restored optical transmission coupling model is N rows and K columns (N is the total number of monitoring grid nodes, and K is the number of dimming loops), and the dimming control command vector is K rows and 1 column. After matrix multiplication, an N rows and 1 column vector is obtained, which is consistent with the dimension of the external sunlight field distribution vector (N rows and 1 column), ensuring that the superposition operation can be executed.
[0050] Preferably, mapping the spectral intensity data within the structured sparse spectral state set to a state input vector for reinforcement learning includes: A complex eigenvectorization operator is defined to perform real-imaginary separation and serialization concatenation on the structured sparse spectral state set. The calculation formula is as follows: The state input vector is constructed using the following formula: in, This represents the complex eigenvectorization operator, used to convert a complex matrix into a long vector in the real field; This represents the structured sparse spectral state set, which includes the filtered complex spectral coefficients; This represents a matrix vectorization operator, used to flatten a matrix into column vectors in column order. This indicates the operation of extracting the real part, which extracts the real matrix components of a complex matrix; This indicates the operation of extracting the imaginary part, which extracts the imaginary matrix components of a complex matrix; The state input vector represents the input layer data of the reinforcement learning policy network; This represents the sky brightness weighting coefficient vector, which characterizes the external sunlight environment at the current moment; The vector represents the coordinate amplitude of the thermally induced deformation modes, which consists of the coordinate amplitudes of all orders of the thermally induced deformation modes and characterizes the thermal deformation state of the roof at the current moment. It represents the scalar of the relative position of the train, which indicates the normalized position of the train on the runway at the current moment.
[0051] The complex spectral feature enhancement vector is a high-dimensional real vector formed by concatenating the real and imaginary eigenvectors of the structured sparse spectral state set. The real matrix components are the real parts of the complex matrices contained in the structured sparse spectral state set. The imaginary matrix components are the imaginary parts of the complex matrices contained in the structured sparse spectral state set. The real eigenvector is a one-dimensional vector obtained by performing a column-wise expansion serialization operation on the real matrix components. The imaginary eigenvector is a one-dimensional vector obtained by performing a column-wise expansion serialization operation on the imaginary matrix components. The structured sparse spectral state set is a low-dimensional spectral model containing key spectral domain data after dimensionality reduction extraction using a screening operator. The sky brightness weight coefficient vector is a combined coefficient vector obtained by fitting outdoor sky brightness observation data to the International Commission on Illumination (ICIDI) standard general sky model using the least squares method. The thermo-deformation mode coordinate amplitude vector is a vector composed of the real-time thermo-deformation mode coordinate amplitudes of all orders, representing the current thermal deformation state of the roof. The relative position scalar of the train is a normalized value of the train's physical coordinates relative to the span modulus of the steel portal frame structure, representing the train's current position on the runway. The state input vector is obtained by concatenating the complex spectral feature enhancement vector, the sky brightness weight coefficient vector, the thermal deformation mode coordinate amplitude vector, and the relative position scalar of the train in a predetermined dimensional order, and is used as input to the reinforcement learning policy network. The complex feature vectorization operator is used to separate a complex matrix into real and imaginary matrix components, then serialize them separately and concatenate them into a real vector. Preferably, the concatenation is done with the real part first and the imaginary part last to ensure the integrity of the complex information and facilitate subsequent network computation. The matrix vectorization operator is used to straighten a matrix into a one-dimensional column vector in a predetermined order. Preferably, it is done in column-major order to conform to mainstream matrix operation standards and ensure operational consistency.
[0052] Realization of Complex Spectral State Sets: Structured sparse spectral state sets contain complex information and cannot be directly input into reinforcement learning networks. By using a complex feature vectorization operator, the real and imaginary matrix components are first separated. For example, a structured sparse spectral state set might be a 10x5 complex matrix; after separation, this yields two 10x5 real matrices. These two real matrices are then expanded in column-major order: the real matrix expands into a 50-dimensional vector, and the imaginary matrix expands into a 50-dimensional vector. These vectors are then concatenated to generate a 100-dimensional complex spectral feature enhancement vector, thus realizing the realization of complex information and adapting it to network input requirements.
[0053] Cascaded integration of multi-dimensional state information: Reinforcement learning strategy networks require comprehensive state information to support decision-making, thus integrating four types of key data. For example, the complex spectral feature enhancement vector is 100-dimensional, the sky brightness weight coefficient vector is 15-dimensional, the thermal deformation mode coordinate amplitude vector is 3-dimensional, and the vehicle relative position scalar is 1-dimensional. These are cascaded in the order of complex spectral feature enhancement vector → sky brightness weight coefficient vector → thermal deformation mode coordinate amplitude vector → vehicle relative position scalar, resulting in a 119-dimensional state input vector. This ensures that the network can simultaneously acquire spectral features, sunlight environment, thermal deformation state, and vehicle position information.
[0054] The specific order of column-wise serialization is as follows: starting from the first column of the matrix, extract the elements of each column in sequence and arrange them from top to bottom. For example, in a 3x2 matrix, if the first column elements are a1, a2, a3 and the second column elements are b1, b2, b3, the expanded order will be a1, a2, a3, b1, b2, b3, ensuring that the serialization rules are unambiguous.
[0055] The predetermined order of dimensional concatenation is fixed as follows: complex spectrum feature enhancement vector → sky brightness weight coefficient vector → thermal deformation mode coordinate amplitude vector → vehicle relative position scalar. This order is arranged logically according to core features → external environment → structural state → equipment position, ensuring the rationality of information transmission and facilitating subsequent network parameter adjustments.
[0056] The order of the coordinate amplitude vectors of the thermo-deformation modes is arranged from low to high according to the order of the thermo-deformation distribution modes. For example, the coordinate amplitudes of the 3rd mode are q1, q2, and q3, and the vector order is q1, q2, and q3, which is consistent with the mode order in the offline simulation to ensure the accuracy of the state information.
[0057] The dimension calculation rule for complex spectral feature enhancement vectors is as follows: the dimension is equal to the number of rows of the structured sparse spectral state set multiplied by the number of columns and then multiplied by 2. For example, if the structured sparse spectral state set has M rows and K columns, the real part is expanded to M×K dimensions, the imaginary part is expanded to M×K dimensions, and the dimension after concatenation is 2×M×K. This clear dimension calculation method facilitates the design of the network input layer.
[0058] Preferably, the reinforcement learning policy network includes: The hidden layer feature vector is calculated using the following formula: The dimming control command is calculated using the following formula: in, The state input vector represents the input data of the reinforcement learning policy network; This represents the first weight matrix, used to map the state input vector to the hidden layer feature space; This represents the first bias vector, used to adjust the activation threshold in the hidden layer feature space; This represents the hyperbolic tangent nonlinear activation function, used to introduce nonlinear features and output results with values ranging from negative one to positive one. This represents the hidden layer feature vector, which characterizes the intermediate layer state features after nonlinear extraction. This represents the second weight matrix, used to map the hidden layer feature vectors to the output action space; This represents the second bias vector, used to adjust the activation threshold of the output action space; This indicates that the dimming control command corresponds to the normalized brightness setting value of each dimming loop; This represents the sigmoid nonlinear activation function, used to strictly limit the output result to an open interval between zero and one; The input variables represent the sigmoid nonlinear activation function; It represents the base of the natural logarithm.
[0059] The first weight matrix is the weight matrix of the hidden layer of the reinforcement learning policy network, used to map the state input vector to the hidden layer feature space. It is preferably a matrix with a dimension equal to the hidden layer feature vector dimension multiplied by the state input vector dimension, initialized with a zero-mean Gaussian distribution and a standard deviation of 0.01 to balance feature extraction capability and training stability. The first bias vector is the bias vector of the hidden layer of the reinforcement learning policy network, used to adjust the activation threshold of the hidden layer feature space. It is preferably a zero vector with the same dimension as the hidden layer feature vector to avoid excessively large initial biases affecting network training convergence. The hidden layer feature vector is the vector obtained after the state input vector undergoes linear transformation by the first weight matrix, adjustment by the first bias vector, and mapping by the hyperbolic tangent activation function, used to represent the extracted nonlinear intermediate features. The second weight matrix is the weight matrix of the output layer of the reinforcement learning policy network, used to map the hidden layer feature vector to the output action space. It is preferably a matrix with a dimension equal to the number of dimming loops multiplied by the hidden layer feature vector dimension, initialized with a zero-mean Gaussian distribution and a standard deviation of 0.01 to adapt to the output dimension and ensure the rationality of instruction generation. The second bias vector is the bias vector of the output layer of the reinforcement learning policy network. It is used to adjust the activation threshold of the output action space. Preferably, the zero vector dimension is consistent with the number of dimming loops to ensure initial stability and facilitate subsequent gradient updates. The hyperbolic tangent nonlinear activation function is used to perform nonlinear mapping on the result of the linear transformation of the hidden layer, with an output value ranging from -1 to +1. The sigmoid nonlinear activation function is used to perform element-wise mapping on the result of the linear transformation of the output layer, which can strictly limit the output to an open interval between zero and one. The input variable is the result of the input of the sigmoid nonlinear activation function, i.e., the hidden layer feature vector, after linear transformation by the second weight matrix and adjustment by the second bias vector. The base of the natural logarithm is a fixed constant with a value of approximately 2.7183.
[0060] The design logic of the two-layer network structure: The reinforcement learning policy network adopts a two-layer structure combining hidden layers and output layers, which ensures non-linear feature extraction capability while avoiding training complexity caused by excessive network depth. For example, the state input vector is 119-dimensional, the first weight matrix is set to 64×119, and the first bias vector is 64-dimensional. After linear transformation and hyperbolic tangent activation, a 64-dimensional hidden layer feature vector is obtained. This dimension can effectively extract key features without increasing the computational cost too much.
[0061] Combination of activation functions: The hidden layer uses the hyperbolic tangent activation function because it introduces strong nonlinearity and has a symmetrical output range, which helps alleviate the gradient vanishing problem and allows the network to learn the complex mapping relationship between the state and the dimming command. The output layer uses the sigmoid activation function because the dimming control command needs to be between zero and one, corresponding to 0% to 100% of the lamp's brightness. For example, the linear transformation result of the output layer is 3.0, which becomes 0.95 after sigmoid activation, corresponding to 95% brightness, meeting the actual dimming requirements.
[0062] Initialization strategy for weights and biases: The weight matrix is initialized with a zero-mean Gaussian distribution and a standard deviation of 0.01 to avoid the activation function from being too large at the beginning. The bias vector is initialized to a zero vector so that the initial output of the network is in an intermediate state, providing a stable starting point for subsequent training. For example, the first bias vector b1 is a 64-dimensional zero vector to ensure that the initial hidden layer feature vectors are distributed near zero, which is convenient for gradient updates.
[0063] The dimensions of the weight matrices are set according to the following criteria: The number of rows in the first weight matrix A1 is equal to the dimension of the hidden layer feature vector, preferably 64 or 128, and the number of columns is equal to the dimension of the state input vector. For example, if the state input vector is 119-dimensional, then A1 is a 64×119 matrix. The number of rows in the second weight matrix A2 is equal to the number of dimming loops K, and the number of columns is equal to the dimension of the hidden layer feature vector. For example, if K=5, then A2 is a 5×64 matrix. The dimensions are set to adapt to the input and output dimensions and balance feature transfer and computational efficiency.
[0064] The dimension selection of the hidden layer feature vector is preferably 64 dimensions. If the dimension of the state input vector is high (such as more than 200 dimensions), it can be increased to 128 dimensions. The basis for selection is that the number of dimming loops in industrial scenarios is usually small (5-20). A 64-dimensional hidden layer is sufficient to capture the mapping relationship between the state and the command, while controlling the amount of computation and avoiding excessive latency in real-time control.
[0065] Numerical precision control of the S-type activation function: A numerically stable implementation method is adopted during calculation. When the input variable x is greater than 20, 0.9999 is directly output; when x is less than -20, 0.0001 is directly output. This avoids exponential overflow and ensures that the output value is stable in the range of zero to one, which meets the precision requirements of the dimming command.
[0066] Network parameters are stored and updated in batches during training. After each update, the rationality of the output instructions is verified. If a value exceeds the range of 0-1 (due to numerical error), it is directly truncated to 0 or 1 to ensure the feasibility of practical applications.
[0067] Preferably, the reinforcement learning policy network is driven to generate dimming control commands in response to the state input vector, and the parameters of the reinforcement learning policy network are updated based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor, including: The overall objective function and the comprehensive cost of a single-step operation are constructed and calculated using the following formulas: The partial derivative vector of the single-step operation comprehensive cost with respect to the dimming control command is calculated based on the differentiable digital twin predictor, using the following formula: The parameter update of the reinforcement learning policy network is performed using the following formula: in, The total objective function represents the cumulative loss within the rolling time window; This represents the set of parameters to be updated in the reinforcement learning policy network; This indicates the length of the rolling time window, i.e., the number of control cycles it contains; The single-step operation cost at time t represents the total cost of the operation. This represents the energy consumption weighting coefficient for lighting fixtures; Indicates the total number of dimming circuits; This indicates that the k-th dimming circuit is in dimming command mode. The real-time operating power was calculated using a linear affine model. This represents the illuminance deviation weighting coefficient; This represents the digital twin predicted illuminance vector output by the differentiable digital twin predictor; This represents the preset target illuminance vector, which corresponds to the target illuminance value on the working surface grid. The squared Euclidean norm of a vector; This represents the control stability weighting coefficient, used to suppress dimming fluctuations; This refers to the dimming control command at the current time t; The dimming control command mentioned in the previous control moment; This represents the vector of partial derivatives of the single-step operation cost with respect to the dimming control command; This represents the power slope vector, whose elements are composed of the difference between the full-load power and the no-load power of each dimming loop; The transpose matrix of the real-time restored optical transmission coupling model is the key physical coupling term that propagates the illuminance deviation back to the control space. The Jacobian matrix representing the dimming control command with respect to the policy network parameters is obtained by the backpropagation algorithm of the reinforcement learning policy network. This indicates that the parameters update gradient; This represents the preset learning rate, used to control the step size for parameter updates.
[0068] The rolling time window length is a preset fixed time window containing several control cycles, preferably 12 control cycles, to balance short-term control effects and computational efficiency, adapting to the common setting of a 5-minute control cycle in industrial scenarios. The overall objective function is a function obtained by summing the comprehensive costs of all single-step operations within the rolling time window, used to measure the overall performance of the reinforcement learning strategy. The comprehensive cost of a single-step operation is a cost value obtained by linearly superimposing the weighted lamp operating power consumption, the weighted sum of squared illuminance deviations, and the weighted sum of squared control fluctuation amplitudes. The lamp energy consumption weighting coefficient is used to adjust the proportion of lamp operating power consumption in the single-step cost, preferably 1, to keep the energy consumption item consistent with other cost items in terms of magnitude. The total number of dimming loops is the total number of independent control units formed by aggregating electrical circuit designs. The real-time operating power is the actual operating power of the k-th dimming loop under the current dimming command, calculated using a linear affine model. The illuminance deviation weighting coefficient is used to adjust the proportion of the sum of squared illuminance deviations in the single-step cost, preferably 0.0001, to avoid the illuminance deviation item being too large and masking the influence of other cost items. The digital twin predicted illuminance vector is the estimated illuminance of the working surface at the next moment, output by the differentiable digital twin predictor. The preset target illuminance vector is the target illuminance value of each monitoring point on the pre-set working surface grid. The control stability weight coefficient is used to adjust the proportion of the sum of squares of control fluctuation amplitude in the single-step cost, preferably 10, to effectively suppress frequent fluctuations in dimming commands and avoid control oscillations. The current dimming control command is the command vector output by the reinforcement learning policy network to adjust the brightness of each dimming loop. The previous control moment dimming control command is the dimming control command vector output by the reinforcement learning policy network in the previous control cycle. The control cycle is the time interval between two dimming control command outputs, preferably 5 minutes, based on the conventional response cycle of industrial lighting control, balancing real-time performance and stability. The power slope vector is a vector composed of the difference between the full-load power and the idle power of each dimming loop, characterizing the rate at which the dimming command affects the power. The transpose matrix of the real-time restored optical transmission coupling model is the transpose matrix of the real-time restored optical transmission coupling model, used to backpropagate the illuminance deviation to the control space. Illuminance deviation sensitivity is the partial derivative vector of the comprehensive cost of a single-step operation with respect to the dimming control command, characterizing the degree to which changes in the dimming command affect the cost. Power change rate is the rate of change of the luminaire's operating power consumption with respect to the dimming control command, directly characterized by the power slope vector. Stability derivative is the derivative of the control fluctuation amplitude with respect to the dimming control command, characterizing the impact of changes in the dimming command on control stability. The policy network parameter set consists of all parameters to be updated in the reinforcement learning policy network, including the first weight matrix, the first bias vector, the second weight matrix, and the second bias vector. The parameter update gradient is the gradient vector of the overall objective function with respect to the policy network parameter set, used to guide the direction of parameter updates.The learning rate is a parameter used to control the step size of the policy network parameter updates. It is preferably 0.001 to balance the parameter update speed and convergence stability, avoiding oscillations caused by an excessively large step size or slow convergence caused by an excessively small step size. The Jacobian matrix is the derivative matrix of the dimming control command with respect to the policy network parameters, representing the influence of parameter changes on the dimming command.
[0069] The construction logic of single-step integrated cost is as follows: The three core control objectives of lamp energy consumption, illuminance deviation, and control fluctuation are integrated into single-step cost. For example, if a dimming circuit has a full-load power of 100 watts and the current dimming command is 0.8, the energy consumption item is 1×100×0.8=80; the difference between the predicted illuminance and the target illuminance is 50 lux, so the illuminance deviation item is 0.0001×50²=0.25; the difference between the current command and the previous command is 0.1, so the control fluctuation item is 10×0.1²=0.1. The single-step cost is 80+0.25+0.1=80.35, thus achieving multi-objective collaborative optimization.
[0070] The method for calculating illuminance deviation sensitivity is as follows: Utilizing the characteristics of a differentiable digital twin predictor, the difference vector between the digital twin predicted illuminance vector and the preset target illuminance vector is first calculated. For example, the difference vector is [10,-8,5]lux. Then, the difference vector is multiplied on the left by the transpose matrix of the real-time restored optical transmission coupling model to obtain the illuminance deviation sensitivity. This process establishes a direct correlation between illuminance deviation and dimming command, providing a physical basis for gradient update.
[0071] Calculation of parameter update gradient: Based on the chain rule, first calculate the partial derivative vector of the single-step cost with respect to the dimming command, then multiply it by the transpose of the Jacobian matrix of the dimming command with respect to the policy network parameters, and finally accumulate it within the rolling time window to obtain the parameter update gradient of the overall objective function. For example, with a window length of 12 cycles, the gradient contribution of 12 cycles is accumulated to ensure that the gradient can reflect the control effect over a period of time.
[0072] Policy parameter update method: Gradient descent is used to iteratively update the parameters along the opposite direction of the parameter update gradient. The learning rate of 0.001 ensures that the update step size is moderate. For example, if the parameter update gradient is [0.5, -0.3], then the parameter update amount is 0.001 × [0.5, -0.3], so as to achieve steady optimization of the policy network.
[0073] The value of the rolling time window length is determined based on the following criteria: 12 control cycles (i.e., 1 hour) is preferred. If the industrial scenario has high requirements for control response speed, 6 cycles can be used. If a smoother control effect is required, 24 cycles can be used. The value should be determined in combination with the control cycle and performance requirements of the specific scenario.
[0074] Standard for setting weighting coefficients: Lighting energy consumption weighting coefficient w E Illuminance deviation weighting coefficient w YControl stability weighting coefficient w U The settings must meet w E ×P max ≈w Y ×(ΔY max )²≈w U ×(ΔU max )², where P max For the maximum power consumption of a single loop, ΔY max ΔU is the maximum permissible illuminance deviation. max To maximize the allowable control fluctuation, for example, P max =100 watts, ΔU max =100 lux, ΔU max =0.5, then w E =1, w Y =0.0001, w U =4, ensuring that the three costs are numerically balanced.
[0075] The specific form of the linear affine power model: Real-time operating power P k(uk(t)) =P k0 +(P k1 -P k0 )×uk(t), where P k0 P represents the no-load power (typically 5-10 watts) when dimming command 0 is applied. k1 The full-load power when dimming command 1 is given (determined by the luminaire parameters).
[0076] The Jacobian matrix is solved using the backpropagation algorithm. Starting from the output layer, the derivatives of the dimming command with respect to the second bias vector, the second weight matrix, the hidden layer eigenvector, the first bias vector, and the first weight matrix are calculated sequentially to form the Jacobian matrix. The process can be achieved by using the automatic differentiation function of the deep learning framework to ensure accurate and efficient solution.
[0077] Learning rate adjustment rules: The initial learning rate is set to 0.001. Every 1000 iterations, if the overall objective function decreases by less than 1%, the learning rate is halved, and reduced to a minimum of 0.00001 to avoid oscillations in later iterations and ensure stable convergence of the network.
[0078] Stopping condition for parameter updates: When the absolute value of the change in the total objective function is less than 0.01 after 500 consecutive iterations, or when the number of iterations reaches 10,000, stop parameter updates, output the current optimal strategy parameters, and balance the optimization effect with the computational cost.
[0079] Example 2: A building energy consumption dynamic optimization system based on BIM and reinforcement learning, applied to any of the building energy consumption dynamic optimization methods based on BIM and reinforcement learning, includes: The initial modeling module is used to establish a monitoring grid and dimming loop that are aligned with the structural span modulus space of the steel portal frame based on the building information model, and to generate an initial optical transmission correlation model. The spectral domain feature extraction module is used to convert the initial optical transmission correlation model to the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermal deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the vehicle runway position, the module filters key spectral domain data and assembles them into a structured sparse spectral state set. The digital twin and state mapping module is used to establish a differentiable digital twin predictor based on the structured sparse spectral state set, and to map the spectral intensity data in the structured sparse spectral state set into a state input vector for reinforcement learning. The reinforcement learning control and update module is used to drive the reinforcement learning policy network to generate dimming control commands in response to the state input vector, and to perform parameter updates of the reinforcement learning policy network based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor.
[0080] like Figure 2 As shown, Figure 2 The paper presents the application scenarios and core component layout of a building energy consumption dynamic optimization system based on BIM and reinforcement learning. The building structure is based on a steel portal frame, with parallel runway beams on both sides for vehicles to travel back and forth. Skylights are arranged on the roof to introduce natural lighting. The interior is divided into monitoring grids aligned with the span modulus of the portal frame structure. Illuminance monitoring points are set at the grid nodes to collect illumination data of the working surface. Temperature sensors are installed at key locations on the roof to capture information related to thermal deformation. High-bay lights are suspended from the ceiling and aggregated into independent dimming circuits according to the electrical circuit design. At the same time, sky brightness data can be collected in unobstructed outdoor areas. All components work together to provide data support for the construction of the initial light transmission correlation model, the assembly of the structured sparse spectral state set, the establishment of the differentiable digital twin predictor, and the updating of the reinforcement learning strategy.
[0081] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.
Claims
1. A dynamic optimization method for building energy consumption based on BIM and reinforcement learning, characterized in that, include: Based on the building information model, a monitoring grid and dimming loop are established that are spatially aligned with the structural span modulus of the steel portal frame, and an initial optical transmission correlation model is generated. The initial optical transmission correlation model is transformed into the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermal deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the runway position, key spectral domain data are screened and assembled into a structured sparse spectral state set. A differentiable digital twin predictor is established based on the structured sparse spectral state set, and the spectral intensity data in the structured sparse spectral state set is mapped to the state input vector of reinforcement learning. The reinforcement learning policy network is driven to generate dimming control commands in response to the state input vector, and the parameters of the reinforcement learning policy network are updated based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor.
2. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1, characterized in that, Based on the Building Information Model (BIM), a monitoring grid and dimming loop are established that are spatially aligned with the structural span modulus of the steel portal frame. An initial optical transmission correlation model is then generated, including: The structural span modulus of the steel portal frame in the building information model is analyzed. The monitoring grid is divided along the length of the building using the structural span modulus as the reference alignment method. The illuminance monitoring point data on the two-dimensional monitoring grid is serialized and arranged into an illuminance observation sequence. Based on the electrical circuit design, multiple luminaires are aggregated into an independent dimming circuit, and the standard photometric distribution data file corresponding to each luminaire is read; For each illuminance monitoring point and each dimming loop, based on the light intensity distribution data recorded in the standard photometric distribution data file, combined with the inverse square relationship of the spatial distance between the luminaire and the illuminance monitoring point and the cosine relationship of the angle between the incident light and the normal of the working surface, the illuminance contribution value of a single lamp is calculated. The illuminance contribution values of all lamps belonging to the same dimming circuit to the same illuminance monitoring point are linearly superimposed to obtain the correlation coupling coefficient of the dimming circuit to the illuminance monitoring point, and the set of all correlation coupling coefficients is used as the initial optical transmission correlation model.
3. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1, characterized in that, Transforming the initial optical transmission correlation model to the spatial distribution spectral domain includes: A longitudinal spatial transformation component is constructed based on the number of nodes of the monitoring grid in the length direction of the steel portal frame, and a transverse spatial transformation component is constructed based on the number of nodes of the monitoring grid in the transverse span direction. Both the longitudinal and transverse spatial transformation components are composed of a discrete complex exponential function sequence based on the unit orthogonal normalization rule. Perform the Kronecker product operation on the horizontal spatial transformation component and the vertical spatial transformation component to generate a global two-dimensional spatial orthogonal transformation operator; The initial optical transmission correlation model is transformed by performing a linear left multiplication operation using the global two-dimensional spatial orthogonal transformation operator, thereby converting the optical transmission coupling coefficient in the physical space into a global spatial distribution spectrum model. The global spatial distribution spectrum model contains complex amplitude and phase information in the full frequency domain.
4. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 3, characterized in that, The initial optical transmission correlation model is transformed into the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermally induced deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the runway position, key spectral domain data are screened and assembled into a structured sparse spectral state set, including: The ratio of the structural span modulus of the steel portal frame to the longitudinal grid spacing is calculated to determine the discrete period, and the repetition frequency index of the frame is determined based on the ratio of the total number of longitudinal grid nodes to the discrete period. Based on the cross-frame repetition frequency index and its second harmonic index, a set of basic structure frequency points is constructed. Based on the rule of offsetting adjacent frequency points in the set of basic structure frequency points, a set of vehicle motion sideband frequency points is constructed. Combined with the set of thermal deformation distribution mode frequency points composed of high-energy characteristic frequency points determined by offline thermal deformation simulation, the union of the set of basic structure frequency points, the set of vehicle motion sideband frequency points, and the set of thermal deformation distribution mode frequency points is used as the key spectral domain screening index set. The physical coordinates of the vehicle on the runway are normalized to the relative position of the vehicle with respect to the structural span modulus, and the value of a complex exponential function with the natural logarithm base and the product of the imaginary unit, the relative position of the vehicle, and the frequency index is calculated as the displacement parameter of the vehicle motion. The preset thermal deformation distribution mode spectrum template and the vehicle motion sideband spectrum template are called. The thermal deformation disturbance term is obtained by calculating the product of the real-time thermal deformation mode coordinates and the thermal deformation distribution mode spectrum template. The vehicle motion disturbance term is obtained by calculating the product of the real part of the vehicle motion displacement parameter and the vehicle motion sideband spectrum template. The global spatial distribution spectrum model, the thermal deformation perturbation term, and the vehicle motion perturbation term are linearly superimposed, and the superposition result is extracted by a filtering operator constructed based on the key spectral domain filtering index set to obtain the structured sparse spectral state set.
5. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 4, characterized in that, A differentiable digital twin predictor is established based on the structured sparse spectral state set, including: The transpose of the filtering operator is used to perform a dimension-uppering filling operation on the structured sparse spectral state set to restore it to full-size spectral domain data. The inverse operator of the global two-dimensional space orthogonal transformation operator is used to inversely transform the full-size spectral domain data back to physical space, and its real part is taken as the real-time restoration of the optical transmission coupling model. The sky brightness weight coefficient vector is obtained by fitting outdoor sky brightness observation data, and the sky brightness weight coefficient vector is multiplied by the preset solar response basis function matrix to generate the external solar field distribution vector. The real-time restored optical transmission coupling model and the dimming control command vector are subjected to matrix multiplication, and the external sunlight field distribution vector is superimposed to output a digital twin predicted illuminance vector. The digital twin predicted illuminance vector is completely differentiable with respect to the dimming control command vector.
6. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 4, characterized in that, Mapping the spectral intensity data within the structured sparse spectral state set to the state input vector for reinforcement learning includes: The structured sparse spectral state set containing complex information is separated into real matrix components and imaginary matrix components. Column-wise expansion and serialization operations are performed on the real matrix components and the imaginary matrix components respectively to generate real feature vectors and imaginary feature vectors. The real feature vectors and the imaginary feature vectors are then concatenated end to end to generate a complex spectral feature enhancement vector. Obtain the sky brightness weighting coefficient vector used to characterize the external sunlight environment at the current moment, the thermal deformation mode coordinate amplitude vector used to characterize the thermal deformation state of the roof, and the vehicle relative position scalar used to characterize the vehicle's position. The complex spectral feature enhancement vector, the sky brightness weight coefficient vector, the thermal deformation mode coordinate amplitude vector, and the vehicle relative position scalar are concatenated in a predetermined order to construct the state input vector with determined dimensions required by the reinforcement learning policy network.
7. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1, characterized in that, Reinforcement learning policy networks, including: The hidden layer computing unit is used to receive the state input vector, perform a linear transformation operation on the state input vector using a first weight matrix and a first bias vector, and map the linear transformation result using a hyperbolic tangent nonlinear activation function to output a hidden layer feature vector. The output layer computing unit is used to receive the hidden layer feature vector, perform a linear transformation operation on the hidden layer feature vector using a second weight matrix and a second bias vector, and perform element-wise mapping on the linear transformation result using a sigmoid nonlinear activation function, and output the dimming control command whose numerical range is limited to the normalized interval of zero to one.
8. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 5, characterized in that, The reinforcement learning policy network is driven to generate dimming control commands in response to the state input vector, and the parameters of the reinforcement learning policy network are updated based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor, including: A fixed-length rolling time window containing several control cycles is set. Within the rolling time window, the comprehensive cost of single-step operation is accumulated to construct the overall objective function. The comprehensive cost of single-step operation is composed of the weighted lamp operating power consumption, the sum of squares of the illuminance deviation between the weighted digital twin predicted illuminance vector and the preset target illuminance vector, and the sum of squares of the control fluctuation amplitude between the current dimming control command and the previous dimming control command, which are linearly superimposed. The differentiability of the differentiable digital twin predictor is used to calculate the sensitivity of the single-step operation integrated cost to the illuminance deviation of the dimming control command. The calculation process includes: multiplying the difference vector between the digital twin predicted illuminance vector and the preset target illuminance vector by the transpose matrix of the real-time restored optical transmission coupling model. Based on the chain rule, and combining the illuminance deviation sensitivity, the power change rate of the lamp operating power consumption with respect to the dimming control command, and the stability derivative of the control fluctuation amplitude with respect to the dimming control command, the parameter update gradient of the total objective function with respect to the parameters of the reinforcement learning policy network is calculated. Using a preset learning rate, the parameters of the reinforcement learning policy network are iteratively updated in the opposite direction of the parameter update gradient.
9. A building energy consumption dynamic optimization system based on BIM and reinforcement learning, applied to the building energy consumption dynamic optimization method based on BIM and reinforcement learning as described in any one of claims 1-8, characterized in that, include: The initial modeling module is used to establish a monitoring grid and dimming loop that are aligned with the structural span modulus space of the steel portal frame based on the building information model, and to generate an initial optical transmission correlation model. The spectral domain feature extraction module is used to convert the initial optical transmission correlation model to the spatial distribution spectral domain. Based on the span repetition frequency index determined by the structural span modulus, the thermal deformation distribution mode determined by the roof component shading law, and the vehicle motion displacement parameters determined by the vehicle runway position, the module filters key spectral domain data and assembles them into a structured sparse spectral state set. The digital twin and state mapping module is used to establish a differentiable digital twin predictor based on the structured sparse spectral state set, and to map the spectral intensity data in the structured sparse spectral state set into a state input vector for reinforcement learning. The reinforcement learning control and update module is used to drive the reinforcement learning policy network to generate dimming control commands in response to the state input vector, and to perform parameter updates of the reinforcement learning policy network based on the energy consumption cost and illuminance deviation sensitivity fed back by the differentiable digital twin predictor.