Intelligent assessment method of campus environment integrating spatio-temporal behavior to affect individual health

By integrating spatiotemporal and environmental data into a multi-level health assessment model, and combining multimodal deep learning and dynamic graph neural networks, the limitations of traditional models in campus environment assessment are overcome. This enables accurate prediction and interpretable analysis of individual health, thereby improving the scientific rigor and precision of campus environment design.

CN121260480BActive Publication Date: 2026-03-31SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies fail to effectively combine spatiotemporal behavior and environmental data when assessing the impact of the campus environment on individual health, leading to biases in health level assessments. Furthermore, traditional models struggle to handle multimodal data, making it impossible to achieve refined design of healthy campus environments.

Method used

By combining a multi-level health assessment model with in-depth collaborative analysis of spatiotemporal and environmental data, a multimodal deep learning architecture is adopted to integrate environmental parameters, spatiotemporal behaviors and physiological and psychological indicators. The influence of environmental factors on behavioral patterns is quantified by using an environment-behavior coupled attention mechanism and a dynamic graph neural network, and the prediction results are output through a two-branch health prediction model.

Benefits of technology

This study provides a comprehensive analysis of the complex mechanisms by which the campus environment and spatiotemporal behavior affect individual health, enhancing the accuracy and operability of environmental intervention measures and providing a scientific basis for campus health design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260480B_ABST
    Figure CN121260480B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent evaluation method for the influence of a campus environment on individual health by integrating space-time behaviors, and the method comprises the following steps: inputting space static characteristics, space-time behavior characteristics, environment dynamic characteristics and physiological and psychological indexes into a multi-level health evaluation model to obtain physiological index and psychological index prediction results of an individual. The multi-level health evaluation model comprises a data preprocessing module, a feature extraction module, a multi-modal fusion module and a health prediction module. The multi-modal deep learning is used to realize the fusion of environment parameters, space-time behaviors and physiological and psychological indexes. The environment-behavior coupling attention mechanism and the dynamic graph neural network are used to quantize the real-time influence of environmental factors on behavior patterns. The health prediction results are output by the psychological index prediction branch and the physiological index prediction branch, effectively solving the pain point that multi-source heterogeneous data in a high-density urban campus environment is difficult to be collaboratively analyzed, and providing a scientific basis for campus health design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and healthy campus construction, specifically involving an intelligent assessment method for the impact of campus environment on individual health that incorporates spatiotemporal behavior. Background Technology

[0002] With urbanization, the built environment of school campuses is receiving increasing attention. Against this backdrop, creating a more supportive campus environment for student growth has become a topic worthy of in-depth research. Current research suffers from two main shortcomings: methodologically, it often assesses health levels from a static perspective, failing to couple spatiotemporal behavior with the campus environment, potentially leading to biased cognitive results; technically, traditional statistical models struggle to simultaneously process environmental and spatiotemporal data, failing to capture the interaction between student activity trajectories and campus environmental parameters. These limitations hinder the design of refined healthy campus environments, necessitating innovative technologies to achieve dynamic and collaborative analysis of multimodal data. Summary of the Invention

[0003] To address at least one of the problems existing in current technologies, this invention provides an intelligent assessment method for the impact of the campus environment on individual health, incorporating spatiotemporal behavior. Through deep collaboration between spatiotemporal and environmental data, it assesses students' health status within the campus environment to predict campus health risks and provide scientific guidance for the formulation of campus environmental health policies. First, a multimodal deep learning architecture is used to integrate environmental parameters, spatiotemporal behavior, and physiological and psychological indicators. Second, an environment-behavior coupled attention mechanism and a dynamic graph neural network are used to quantify the real-time impact of environmental factors on behavioral patterns. Finally, a dual-branch health prediction model and an interpretable module output the prediction results, providing an AI-driven scientific prediction method for high-density urban campus planning.

[0004] To achieve the objectives of this invention, this invention provides an intelligent assessment method for the impact of the campus environment on individual health, incorporating spatiotemporal behavior. The method is characterized by inputting spatial static features, spatiotemporal behavioral features, environmental dynamic features, and physiological and psychological indicators into a multi-level health assessment model to obtain predicted results for individual physiological and psychological indicators. The multi-level health assessment model includes the following modules:

[0005] The data preprocessing module is used to integrate spatial static features, spatiotemporal behavioral features, and environmental dynamic features, and to perform rasterization and data alignment.

[0006] The feature extraction module initializes spatial static features through a graph convolutional network and performs graph convolution operations to capture complex spatial relationships, thereby obtaining spatial functional distribution features; it captures dynamic behavioral patterns through a dual-pathway combination of 3D-CNN and LSTM branches, and obtains behavioral features through fusion and compression; it parses environmental dynamic features through a cross-modal Transformer encoder and outputs environmental representations; and it processes physiological signals through a temporal network and outputs physiological representations.

[0007] The multimodal fusion module takes spatial functional distribution features, behavioral features, physiological representations and environmental representations as inputs. Through the environment-behavior coupling attention mechanism, it quantifies the influence of dynamic environment and spatiotemporal behavior and outputs environment-behavior coupling features. The dynamic graph neural network takes the environment-behavior coupling features as inputs and generates spatiotemporally dependent coupling features through attention aggregation and GRU gating to represent the causal relationship between architectural design and human activities.

[0008] The health prediction module takes spatiotemporally dependent coupling features as input, outputs the probability distribution of ordered categories of psychological recovery through the psychological indicator prediction branch, and generates numerical predictions of physiological parameters through the physiological indicator prediction branch. The psychological indicator prediction branch and the physiological indicator prediction branch are jointly optimized through feature cross and dynamic weighting.

[0009] Furthermore, it also includes an interpretability module, which generates a three-dimensional heatmap based on the prediction results output by the health prediction module to locate health-sensitive areas, and combines SHAP values ​​(Shapley Additive exPlanations) to analyze the contribution of key environmental factors.

[0010] The present invention also provides a computer device.

[0011] The present invention also provides a computer-readable storage medium.

[0012] Compared with the prior art, the present invention can achieve at least the following beneficial effects:

[0013] This invention breaks through the limitations of traditional campus environment assessment methods by innovatively integrating spatiotemporal data to achieve a comprehensive analysis of the complex mechanisms by which the campus environment and spatiotemporal behavior affect individual health. Compared to existing technologies, its core advantage lies in the analysis of the correlation between environmental parameters, student behavioral dynamics, and physiological and psychological health indicators. It effectively solves the pain point of the difficulty in collaborative analysis of multi-source heterogeneous data in high-density urban campus environments, providing a feasible scientific basis for campus health design and significantly improving the accuracy and operability of environmental intervention measures. Attached Figure Description

[0014] Figure 1This is a schematic diagram illustrating the steps of an intelligent assessment method for the impact of campus environment on individual health that incorporates spatiotemporal behavior, as described in an embodiment of the present invention.

[0015] Figure 2 This is a schematic diagram of the module composition of the multi-level health assessment model in an embodiment of the present invention.

[0016] Figure 3 This is a schematic diagram of the data processing flow in the data preprocessing module of this invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] This invention provides an intelligent assessment method for the impact of the campus environment on individual health, incorporating spatiotemporal behavior. First, a hierarchical discretization technique is used to achieve precise structured representation of geographic space and behavioral trajectories, and feature alignment ensures the consistency of environmental data. Second, an environment-behavior coupled attention mechanism and a dynamic graph neural network are used to establish a quantitative relationship between environmental parameters and behavioral patterns. Finally, a dual-branch prediction architecture is used to process psychological scores and physiological time-series data separately, and a three-dimensional gradient heatmap and SHAP value analysis are combined to achieve visualized output. Specifically, an intelligent assessment method for the impact of the campus environment on individual health, incorporating spatiotemporal behavior, includes the following steps:

[0019] Step 1: Obtain spatial static features, spatiotemporal behavioral features, environmental dynamic features, and physiological and psychological indicators to construct a dataset.

[0020] In one embodiment, the data is obtained as follows:

[0021] Spatial static characteristics: Data collected from the campus and surrounding neighborhood geographic information system, including parameters such as building density (%), plot ratio (%), green coverage (%), land mixed use index (proportion of residential / commercial / educational land), population density (persons / km²), distance (m) and number of campuses from city streets, building height (m), and accessibility of public spaces (m).

[0022] Spatiotemporal behavioral characteristics: Dynamically collected through GPS positioning and Bluetooth, including activity trajectory heatmap, activity range during breaks (m²), commuting route and duration (min), time spent in hotspot areas (min), percentage of time spent on physical activity intensity (%), frequency of green space use (times / day), frequency of social interaction (times / hour), and other indicators.

[0023] Environmental dynamic characteristics: Data are collected synchronously through equipment such as panoramic cameras, sound level meters, and thermal index meters, including parameters such as green view rate (%), window-to-wall ratio (%), window-to-view natural ratio (%), class density (person / m²), sound pressure level (LAeq / dB), reverberation time (s), illuminance (lux), air temperature (°C), and relative humidity (%RH).

[0024] Physiological and psychological indicators: Physiological indicators are obtained by continuously monitoring the time-series data of EEG (alpha waves 8-12Hz, beta waves 12-30Hz), skin conductance (μS), and heart rate (bpm) through wearable devices; Psychological indicators include psychological recovery scale scores (e.g., divided into 1-9 points) and life satisfaction scores (e.g., divided into 1-5 points).

[0025] Step 2: Construct a multi-level health assessment model.

[0026] The multi-level health assessment model of this invention assesses the health status of students in a campus environment through deep collaboration of spatiotemporal data and environmental data, thereby achieving accurate prediction of campus health risks.

[0027] The multi-level health assessment model includes the following modules: data preprocessing module, feature extraction module, multimodal fusion module, health prediction module, and interpretability module. The data preprocessing module integrates spatial static features, spatiotemporal behavioral features, and environmental dynamic features. It uses GeoHash rasterization and data alignment to construct the foundation for architectural space analysis. The feature extraction module deconstructs core elements from the preprocessed data. It extracts spatial static features through a graph convolutional network and captures behavioral dynamic patterns through a dual-pathway approach using 3D-CNN and LSTM branches. Specifically, it analyzes environmental dynamic features using a cross-modal Transformer encoder and processes physiological signals through a temporal network, outputting corresponding high-order representations. The multimodal fusion module quantifies the impact of dynamic environment and spatiotemporal behavior through an environment-behavior coupled attention mechanism and drives a dynamic graph neural network to generate spatiotemporally dependent coupled features, revealing the causal relationship between architectural design and human activities. The health prediction module adopts a psychological-physiological dual-branch heterogeneous architecture. It outputs the probability distribution of ordered categories of psychological recovery through an ordered regression model and generates numerical predictions of physiological parameters through a temporal convolutional network. These two approaches are collaboratively optimized through feature cross-fertilization and dynamic weighting. The interpretability module generates a 3D heatmap based on the prediction results to locate health-sensitive areas and combines SHAP values ​​to analyze the contribution of key environmental factors, providing a causal decision-making basis for campus space renovation. Each module forms a closed loop of "spatial data structuring → environmental behavior feature deconstruction → multimodal dynamic coupling → health prediction → regional visualization", breaking through the limitations of traditional static environmental assessment.

[0028] In the data preprocessing module, spatiotemporal data encoding and environmental dynamic feature alignment are performed.

[0029] In the spatiotemporal data encoding stage, a hierarchical discretization and multi-source fusion strategy is adopted to achieve the structured transformation of the original data.

[0030] First, hierarchical discretization is performed, that is, based on the set of geographic coordinate points. A GeoHash spatial coding system is constructed, which discretizes continuous geographic space into raster units with configurable resolution through recursive spatial binary search. Among these, The total number of coordinate points (a positive integer) represents the total number of discrete location samples collected within the study area. Coordinate point index identifier (integer, range) This is used to uniquely identify each spatial location point. Let the target spatial resolution be... Encoding length From the formula Confirmed, among which For the Earth's radius, The location point is the difference in latitude and longitude in radians. The encoding process is defined as follows:

[0031] ;

[0032] ;

[0033] ;

[0034] in, The latitude binary bit value (0 or 1) represents the location point. At the level The spatial division results in latitudinal direction; For the current encoding level (integer, range) ), Location point The actual latitude coordinates; , These represent the minimum and maximum latitude boundaries of the predefined study area; This is a floor function that maps fractions to integers; Modulo-2 operation, the output result is 0 or 1, which constitutes the basic unit of binary encoding; The binary value of longitude (0 or 1) represents the location point. At the level Spatial division along the longitude direction; Location point The actual longitude coordinates; and These are the minimum and maximum longitude boundaries of the predefined study area, respectively. To make the hierarchy Shift the latitude bit value left The bit occupies the high-order bit segment of the encoding; To make the hierarchy left shift of longitude bit value The bit occupies the lower bit segment of the encoding; To sum across all levels, Morton coding space-filling curve mapping is implemented, ensuring that the encoded values ​​of neighboring spatial units remain continuous in the integer domain. Location point GeoHash integer encoding is achieved by progressively stacking latitude and longitude binary bits and assigning different weights (latitude shifted left). Position, longitude shifted to the left (bit) Generate a unique spatial identifier to achieve a hierarchical mapping from continuous geographic space to integer domain.

[0035] All spatial static features are calculated using spherical weighted aggregation within the raster cells. The aggregated spatial static features are then used to initialize node features in the feature extraction module. Spherical weighted aggregation can solve geospatial distortion problems, ensuring the spatial comparability of parameters such as building density and green coverage. It can also unify data scales, aggregating discrete geographic information point data into standardized raster cells.

[0036] In one embodiment, spatial static characteristics such as building density In grid cells Internal calculations are performed using spherical weighted aggregation:

[0037] The output of the dual-path architecture is fused and compressed.

[0038]

[0039] in, For grid cells Building density, which is the proportion of the total building area within a grid unit, is a static characteristic. For raster cell indexing; For point The building's base area at the location, Indicates membership in a grid cell Building sampling point index; For point The actual value of the latitude; For grid cells spherical surface area; and These are the maximum and minimum boundary values ​​in the longitude direction of the grid cell, respectively; and These are the maximum and minimum boundary values ​​in the latitudinal direction of the raster cell, respectively. The latitude is the center of the grid. This is corrected using spherical geometry. The parameters (items) and regional boundary parameters ensure that the building density calculation conforms to the characteristics of the real Earth surface.

[0040] For spatiotemporal behavioral feature data, the movement trajectory is processed through multi-scale spatiotemporal tensor quantization to obtain the behavioral tensor. Specifically, time window slices are defined. , For the first Time windows are divided according to behavioral patterns (e.g., break time windows). Class time window ), Indicates the time window index. This represents the total number of time windows. The generation of the behavior tensor includes:

[0041] ;

[0042] in, Indicates the first The first time window position A behavior tensor, Represents spatial raster coordinates, corresponding to the raster cell positions after GeoHash encoding; Indicates the spatiotemporal behavioral feature dimension index; Represents the original timestamp. Indicates belonging to the first The moment within a time window; For a moment The trajectory point data includes spatial coordinates and behavioral feature vectors; For the first Dynamic characteristics of the environment Weights for device sampling rate compensation; For spatial positioning functions, Represents trajectory points The GeoHash encoding result.

[0043] Finally, a static environment feature matrix is ​​obtained based on spatial static features, and a dynamic behavior feature tensor is obtained based on the behavior tensor:

[0044] ;

[0045] ;

[0046] in, This is a static environment feature matrix; Green space coverage and building density Together with other spatial static features, they constitute the input matrix; For dynamic behavior feature tensors, It is a behavior tensor.

[0047] In the environmental dynamic feature alignment stage, embodiments of the present invention address the data consistency problem across spatiotemporal dimensions for multi-source heterogeneous sensors. Environmental parameter sequence The raw data comes from sensors, which include three categories: fixed sensors, mobile devices, and weather station equipment. The fixed-location sensors (professional sensing equipment deployed at pre-set monitoring points on campus, continuously collecting environmental time-series data (air temperature, relative humidity, class density, etc.) at designated locations to generate discrete time series data) are used for these sensors. The system uses mobile monitoring devices (portable sensors) to collect dynamic environmental parameters (such as green visibility and dynamic distribution of sound environment) along with the movement of personnel, generating a set of trajectory-bound points. ; and the spatial coverage matrix provided by weather stations .in, For the first Dynamic characteristics of the environment Index for environment parameter types; This represents the total number of environmental parameter categories. Data collected from fixed-position sensors, To fix the sensor number, To fix the total number of sensors, To fix the sensor sampling time, The number of samples taken by a single fixed sensor. For the first Second sampling, For data from mobile monitoring devices, For the mobile monitoring device's serial number, The sampling time for the mobile monitoring device. The number of samples taken by a single mobile monitoring device. For the first Second sampling, The first A mobile monitoring device at the sampling time The coordinates of the time.

[0048] Furthermore, in order to realize the environmental parameter sequence and the target spatiotemporal grid For accurate mapping in the spatial dimension, by improving the Kriging interpolation model, directional range parameters are introduced to address the spatial anisotropy caused by the campus building layout. For the target grid cell At any moment The Dynamic characteristics of the environment Based on the improved spatial variability function Construct a system of Kriging equations and obtain the weight vector by solving the system of Kriging equations. .in, For sensor spacing, For a set of sensor pairs, both distance tolerance and azimuth tolerance must be met simultaneously; For sensors At any moment The Class parameter values; Number the sensor. For sensors At any moment The Class parameter value, Number the other sensor. Weight vector. Based on the improved spatial variogram, the Kriging equations are solved to obtain, where... For the first Interpolation weights for each sensor, The symbol is for vector transpose. This improved Kriging interpolation model introduces a directional range parameter. To address the anisotropic diffusion caused by the campus building complex.

[0049] In the time dimension, adaptive Gaussian interpolation is used to address the asynchronous sampling problem. For the sensor... In the absence of time interpolation estimate The calculation is as follows:

[0050] ;

[0051] in, Indicates sensor At the sensor sampling time The actual observed value; This represents any historical point in time that participates in the weighted calculation. For the summation index, and Synonyms, indicating the first Secondary sampling; Indicates the weight of temporal similarity; The confidence coefficient is... This represents the Gaussian kernel bandwidth.

[0052] Confidence coefficient The bandwidth is set according to the device type, while the Gaussian core bandwidth... Adaptive adjustment based on parameter dynamic characteristics: when parameter change rate is detected. Trigger Gaussian kernel bandwidth Adaptive compression. Among them, Indicates the first Threshold for the rate of change of dynamic characteristics of the environment.

[0053] The multi-source fusion stage (fusion of dynamic environmental features) generates the final raster values ​​through confidence-weighted integration:

[0054] ;

[0055] in, Represents grid cells At any moment , No. The final fusion value of the dynamic features of the environment; Indicates the data source type, with a set of values. , Indicates a fixed position sensor. Indicates mobile monitoring equipment, It refers to a weather station. Indicates the data source type In grid cells ,time For the first Interpolated estimates of the dynamic characteristics of the environment. Source type weights. From error variance Determined by both the effective observation density and the overall observation density This represents the preset maximum number of valid observations, used for standardizing weights. This represents the number of valid observations. Introducing spatial distribution uniformity correction:

[0056] ;

[0057] in, The variance of the sensor spacing within the neighborhood. To allow for spacing, this design avoids sampling bias caused by sensor cluster deployment. Data from the motion monitoring device undergoes additional outlier filtering based on the Grubbs test: if the observation meets... If , then replace it with the neighborhood median. Indicates data source The original number of observation points, Indicates mobile monitoring equipment The first collection The original observations of the dynamic characteristics of the environment. Indicates the number of neighbors within the neighborhood. The arithmetic mean of all observations of the dynamic characteristics of the environment. Represents the standard normal distribution Quantiles Indicates the first Sample standard deviation of neighborhood observations of environmental dynamic characteristics.

[0058] In the feature extraction module, a multimodal deep architecture is employed to extract high-order feature representations. Spatial static features are processed through a graph convolutional network. In one embodiment, the graph convolutional network employs an improved GraphSAGE convolutional network, which is built upon a campus spatial graph. Above, node Corresponding grid cell, edge The establishment follows the principle of spatial proximity (Euclidean distance). ). Represents a set of nodes. Represents a set of edges, storing the spatial connections between nodes; , Represents a node instance. , The meaning is the same as before, indicating the raster row and column index.

[0059] Graph convolutional networks are used to initialize spatial static features and perform graph convolution operations to capture complex spatial relationships. Specifically:

[0060] Node feature initialization uses , Represents all spatial static characteristics. Represents a node The initial feature vectors are then used to perform a three-layer graph convolution operation:

[0061] First, aggregate neighborhood features:

[0062] ;

[0063] in, Indicates the first Aggregation features of neighborhood nodes in a layered graph convolutional layer This represents the hierarchical index of the graph convolutional layer. This represents a node in the campus spatial map. Represents a node The adjacent nodes, For nodes The set of adjacent nodes, The weight matrix is ​​a learnable weight matrix; Indicates adjacent nodes In the The feature vector of the layer.

[0064] Next, feature fusion is performed:

[0065] ;

[0066] in, For the projection matrix, For ELU activation function, Represents a node In the The updated feature vectors from the layered graph convolutional layer are used to obtain the spatial functional distribution features as the final output, obtained by the above formula. , which is the output feature of the third layer graph convolution, encoding the spatial functional distribution characteristics.

[0067] Spatiotemporal behavioral features and environmental dynamic features are extracted collaboratively by a dual-path architecture. The dual-path architecture includes a 3D-CNN branch and an LSTM branch.

[0068] The 3D-CNN branch is used to process behavior tensors. It employs a three-layer 3D convolution to process the behavioral tensor. Spatiotemporal feature abstraction is performed. The first layer (32 channels) captures basic spatiotemporal patterns, the second layer (64 channels) extracts intermediate behavioral correlation features, and the third layer (128 channels) generates high-level semantic representations. Finally, these are compressed into spatially independent temporal feature vectors through global average pooling, transforming the original movement trajectory data into a compact feature representation that encodes the dynamic laws of behavior, providing structured input for subsequent fusion: First layer Output ,in Represents a 3D convolution operation; second layer generate Third layer Output .in, Indicates the length of the convolution kernel in the time dimension. Represents the spatial dimension convolution kernel size. Indicates the number of output channels. This represents the convolution kernel weight tensor. This represents the convolution bias vector. , , These represent the output features of the first, second, and third layers of the 3D convolution, respectively.

[0069] The LSTM branch models the temporal behavior sequence of each grid cell. It dynamically captures forward and backward dependencies by scanning forward and backward using a bidirectional LSTM, ultimately outputting a behavior evolution feature vector that fuses long-term and short-term contexts to deconstruct the temporal dynamics of behavior patterns within the grid cell. Specifically:

[0070] The LSTM branch, running parallel to the 3D-CNN branch, targets each grid cell. Processing time-series behavioral data and capturing forward / backward dependencies through a bidirectional LSTM structure:

[0071]

[0072] Indicates in Place The temporal feature vector at any given moment, i.e., the behavioral evolution feature vector that integrates long-term and short-term contexts. Indicates in Place The behavior tensor at any given moment This represents the set of learnable parameters (weight matrix and bias) of BiLSTM.

[0073] The output of the dual-path architecture is fused and compressed to obtain the fused and compressed behavioral features. :

[0074] ;

[0075] The output features of the third 3D convolution layer are processed using global average pooling (GAP). Dimensionality reduction: The fully connected layer maps the splicing layers to obtain the fused and compressed behavioral features. , This represents the weight matrix of the fused fully connected layer.

[0076] Multi-source environmental data is processed using a cross-modal Transformer encoder. Input environmental parameter sequence. First, embed as ,in Encode the parameter type. Encoding the timing position; Indicates the first Dynamic characteristics of the environment The initial embedding vector; The embedding matrix represents the projection of the original features onto the hidden space. After four Transformer encoding layers, each layer includes a multi-head self-attention mechanism and a feedforward network (FFN). Through parallel multi-set attention mechanisms, global dependencies between environmental parameters are uncovered. Based on this, the feedforward network (FFN) performs nonlinear transformations and dimensionality adjustments on the self-attention output.

[0077] ;

[0078] ;

[0079] Attention mechanisms introduce context-dependent masks:

[0080] ;

[0081] in, For the first The intermediate output of the Transformer encoding layer, For the first The final output of the Transformer encoding layer; Index for the Transformer encoding layer. For the input environmental parameter sequence Embedded representation, Indicates the number of attention heads. For the first The output projection matrix of each attention head, , , The first The projection of each attention head's query, key, and value. The dimension of the key vector, the mask matrix Generated by thresholding the correlation coefficient matrix of environmental parameters. , For mask matrix The element value, For indicator functions, when The value is 1 if the condition is met, and 0 otherwise. Environmental dynamic characteristics and The correlation coefficient.

[0082] Finally, the output of the 4th Transformer coding layer is taken. As a representation of the environment.

[0083] Physiological time-series data were processed using a 1D-ResNet and self-attention fusion architecture. Based on the original physiological data acquisition, the physiological time-series data underwent noise reduction filtering, time alignment, and normalization to obtain the physiological time-series input. Furthermore, physiological timing input The process involves five residual blocks (concatenated sequentially to form a chain-like processing method, enabling progressive deepening of multi-scale features), including temporal convolution, skip connections, and output of higher-order features. While preserving key information from the original signal, multi-scale local features are fused. Each residual block contains two temporal convolutional layers (kernel length 5, padding 2) and skip connections. ,in Indicates the first Output characteristics of layer residual blocks This represents the temporal convolution operation within the residual block. This represents the learnable parameters of a temporal convolutional layer. For the first Output characteristics of layer residual blocks This indicates that the jump connection projection matrix is ​​used to match dimensions.

[0084] The output of the residual block is fed into a multi-head self-attention layer for feature projection, global dependency modeling, and residual normalization. This fuses local and global features, outputting the self-attention output features (local features are combined through residual connections). (This is combined with the global features extracted through attention, and then normalized using LayerNorm to achieve an organic unity between local and global representations).

[0085] ;

[0086] ;

[0087] in, , , These represent the projected Query / Key / Value matrices, , , These represent the Query / Key / Value projection matrices, respectively. This indicates the output of the 5th layer residual block. This represents the self-attention output features. This represents the output projection matrix. Indicates matrix transpose. This indicates the dimension of the key vector in the computation of this self-attention layer.

[0088] Finally, physiological representations are output by adaptively max-pooling to unify the sequence length and compressing it. All feature outputs are scaled using LayerNorm and 256-dimensional linear projection before fusion, providing normalization for the multimodal fusion module.

[0089] In the multimodal fusion module, an environment-behavior coupled attention mechanism is first constructed, which aims to quantify the dynamic modulation effect of environmental parameters on behavioral patterns. By establishing an explicit mapping relationship between dynamic environment and spatiotemporal behavior, it solves the limitation of isolated modeling of environmental factors and behavioral features in traditional methods. Its core objective is to dynamically calibrate behavioral correlation through an environmental weight matrix.

[0090] The environment-behavior coupled attention mechanism is composed of three interconnected parts: a direct environment association pathway, an indirect behavior modulation pathway, and attention computation. The direct environment association pathway calculates the synergistic effect between environmental parameters using a learnable diagonal matrix, quantifying the interaction strength of environmental factors themselves. The indirect behavior modulation pathway is generated through a three-layer fully connected network, capturing the dynamic modulation effect of the environment on behavior. Attention computation then uses the output matrix of the direct environment association pathway... Spatiotemporally sensitive weights The values ​​are merged into an environment weight matrix, and the attention weights are calculated to output the attention output matrix. This process explicitly quantifies the enhancing / inhibiting effect of the environment on behavior, and obtains the environment-behavior coupling characteristics based on behavioral features and attention output matrix.

[0091] The input to the multimodal fusion module is behavioral features. Physiological characteristics and spatial environment coding Physiological characteristics are incorporated as part of behavioral characteristics to jointly generate environmental modulation weights. Among these, spatial environment coding... Based on spatial functional distribution characteristics and environmental characterization splicing and copying along the timeline, along with behavioral characteristics For alignment in the time dimension, cross-modal alignment is performed first:

[0092] ;

[0093] in, This represents the behavioral characteristics after cross-modal alignment. , It is a learnable projection matrix.

[0094] Environmental weight matrix The construction employs a dual-path adaptive mechanism, including a direct environment correlation path and an indirect behavior modulation path. The direct environment correlation path calculates the synergistic effect between environmental parameters.

[0095] ;

[0096] in, This represents the output matrix of the direct environmental association pathway. This represents a learnable diagonal matrix. for diagonal elements, It is an environmental feature embedding dimension index. Diagonal matrix. elements The interaction intensity of different environmental factors can be controlled by learnable parameters.

[0097] Indirect behavior modulation pathways are used to capture the dynamic modulation effects of the environment on behavior, generating spatiotemporally sensitive weights through a three-layer fully connected network. :

[0098] ;

[0099] in For affine transformation layer, the output is scaled to... interval; and Here are the weight matrix and bias vector of the first fully connected layer. and These are the weight matrix and bias vector for the second fully connected layer.

[0100] The final environmental weight matrix is Its element value quantifies the strength of the enhancement / inhibition of the behavioral correlation by a specific environment.

[0101] Furthermore, the attention calculation process introduces an environment-behavior coupling factor. :

[0102] ;

[0103] in Generated by projection matrices respectively. The lower triangular mask matrix ensures temporal causality. Output ( After passing through residual connections, the data is passed to the dynamic graph neural network.

[0104] ;

[0105] in, , , These represent the Query / Key / Value matrices after projection. , , These represent the Query / Key / Value projection matrices for the attention calculation process, respectively. This represents the dimension of the key vector in the self-attention layer computation. This represents the attention output matrix. This indicates the characteristics of environment-behavior coupling.

[0106] Dynamic graph neural networks feature environment-behavior coupling characteristics As input, based on dynamic topology construction, attention aggregation and GRU gating are used to finally output spatiotemporally dependent coupled features. In detail:

[0107] Firstly, consider the characteristics of environment-behavior coupling. Construct a spacetime graph based on the initial node states. Edge set Dynamically updated based on behavioral pattern similarity and spatial proximity:

[0108] ;

[0109] in, Indicates time Next node To the node edge weights, , Represents a node , eigenvectors, , Represents a node , spatial coordinates, Indicates an indicator function, when It is set to 1 if it is true, otherwise it is set to 0.

[0110] A hybrid mechanism of gated recurrent units (GRU) and attention aggregation is used for node state updates, outputting spatiotemporally dependent coupled features, where:

[0111] ;

[0112] in, Represents a node At any moment The aggregated message vector, Represents the edge feature transformation matrix. express exist Hidden state vector at time step 1, attention coefficient ,in, and These are learnable attention vectors and weight matrices, respectively, to achieve differentiated aggregation of neighbor nodes; Represents a node The set of adjacent nodes.

[0113] GRU update gating:

[0114] ;

[0115] ;

[0116] Reset door Mathematical form and update gate Perfectly symmetrical. Among them, To update the gate weight matrix, The input weight matrix, express Time Node The updated hidden state vector.

[0117] The final output shows the spatiotemporal dependent coupling characteristics. Through the hidden state vector The system integrates spatiotemporal dependencies, including environmental modulation, for use by the health prediction module.

[0118] The health prediction module employs a heterogeneous architecture with two branches, one for psychological health and the other for physiological health indicators. This architecture includes a psychological indicator prediction branch and a physiological indicator prediction branch. The input to the health prediction module is the spatiotemporally dependent coupled features output by the multimodal fusion module. The spacetime tensor formed by stacking all moments The feature distribution layer routes the data to two independent branches:

[0119] For the psychological indicator prediction branch, this branch receives the spatiotemporal tensor. By progressively compressing dimensionality through a three-layer fully connected network, higher-order psychological state representations are extracted. For discrete, ordered psychological ratings, an innovative ordered regression model is employed to model category probabilities: through a learnable threshold parameter... Continuous predicted values Mapped to cumulative probability Final output Probability distribution of ordered categories This accurately reflects the likelihood of students being at different levels of psychological state. Specifically:

[0120] In the psychological indicator prediction branch, higher-order psychological state representations are first extracted using a three-layer fully connected network:

[0121] ;

[0122] ;

[0123] ;

[0124] in, , , These are the outputs of the first, second, and third layer fully connected networks, respectively. , , , These are the weight matrices for a three-layer fully connected network. , , These are the bias vectors of a three-layer fully connected network.

[0125] Then, the continuous predicted values ​​are mapped to... using an ordinal regression model. Ordered categories:

[0126] ;

[0127] in For the sigmoid function, This represents the psychological score to be predicted. For category indexing, This represents the total number of categories in the psychological rating, and the learnable threshold parameter. Satisfying monotonicity constraints .

[0128] The final class probability is derived from the cumulative probability difference:

[0129] ;

[0130] in, Indicates that the score equals The exact probability, Indicates that the score does not exceed The cumulative probability, Indicates that the score does not exceed The cumulative probability.

[0131] For the physiological indicator prediction branch, a temporal convolutional network is used for continuous physiological parameters, with the same spatiotemporal tensor as the psychological indicator prediction branch. First, the spacetime tensor Reorganized into physiological characteristic tensors To adapt to temporal processing, an 8-layer dilated causal convolutional network (TCN) is used to capture multi-scale patterns, and residual connections are used to prevent gradient vanishing. Finally, the predicted tensor is compressed using depthwise separable convolution. This enables the prediction of physiological parameter values. Specifically:

[0132] In the physiological indicator prediction branch, the spatiotemporal features are reorganized first:

[0133]

[0134] in, This is the recombined physiological characteristic tensor.

[0135] Based on this, the physiological characteristic tensor Input dilated causal convolutional layer:

[0136] ;

[0137] in For the first The dilation factor of a layer-dilated causal convolutional layer Indicates the first Nodes in the output tensor of a layered dilated causal convolutional layer At any moment eigenvalues, For the first The kernel weights of the dilated causal convolutional layer The kernel length is 1. Indicates the first The output tensor of a layered dilated causal convolutional layer The index variable representing the convolution kernel. It is a physiological characteristic tensor The result of the layer-by-layer transformation, i.e. .

[0138] Residual connection structure is adopted:

[0139] ;

[0140] in, The jump connection weight matrix of the residual block. For a three-dimensional tensor, through all nodes and at all time points... The structure consists of multiple layers of dilated causal convolutional layers followed by separable convolutional compression of dimensions. In one embodiment, based on considerations such as receptive field requirements, feature abstraction levels, and computational efficiency constraints, it is determined to have 8 layers of dilated causal convolutional layers, i.e., 8 layers of dilated causal convolutional layers (number of channels). After that, the output layer uses separable convolutions to compress dimensions:

[0141] ;

[0142] in, This represents the final output tensor of the physiological indicator prediction branch. This represents the output tensor of the 8th layer of the TCN.

[0143] Finally, the psychological indicator prediction branch and the physiological indicator prediction branch achieve collaborative optimization through feature cross-feature injection and dynamic weighting: first, feature cross-feature injection is performed, and the output of the 4th layer of TCN is then processed. (The output layer is layer 4, balancing the receptive field and feature abstraction level.) After spatial pooling, it is combined with the psychological index prediction branch. (This is the output of the second fully connected layer, which encodes higher-order mental state representations without over-compressing information; the third layer is used for ordered regression.) The concatenation yields the feature interaction terms. :

[0144] ;

[0145] Add a shared fully connected layer to process the feature interaction term, and the output is a joint feature vector. , This is the weight matrix. This serves as the bias vector, and the joint feature vector is injected into both the psychological and physiological indicator prediction branches. Simultaneously, a cross-regularization term is added to the joint loss function. ,in, This is the embedding representation of the true label. Joint feature vector. By injecting features and adding additional loss terms, psychological and physiological branches are forced to share information, thereby optimizing prediction consistency and achieving collaborative optimization.

[0146] Next, we define the joint loss function, where the total loss is a weighted combination:

[0147] ;

[0148] in, For the total loss, For ordinal cross-entropy loss, To smooth out L1 loss, This is a balancing factor used to balance the weights. As a weighting factor, Used to control the strength of regularization This represents the set of all trainable parameters for the health prediction module, including the fully connected layer weights and threshold parameters for the psychological indicator prediction branch, and the TCN convolutional kernel weights and skip connection matrix for the physiological indicator prediction branch. In one embodiment, , .

[0149] Finally, dynamic weight adjustment is performed to balance the factors. Adaptive changes during the training process:

[0150] ;

[0151] in, For the total number of training rounds, Indicates the first The loss weighting coefficient for each round of training. This indicates the current training epoch; in one embodiment, the initial loss weights... Decay to final loss weight .

[0152] The probability distribution of the predicted branch output using psychological indicators reflects the likelihood of a student being at a certain level of mental health; for example, a high probability distribution in high score ranges (e.g., =0.7) indicates that the environment significantly promotes psychological recovery, quantifying the discrete impact of the environment on students' subjective psychological state. The physiological parameter values ​​output by the physiological indicator prediction branch reflect the predicted values ​​of physiological parameters, capturing the influence of the environment on students' objective physiological responses.

[0153] The interpretability module generates a three-dimensional gradient heatmap based on the prediction results (i.e., the outputs of the psychological indicator prediction branch and the physiological indicator prediction branch in the health prediction module). Based on the three-dimensional gradient heatmap, it locates key areas affecting health and analyzes the contribution distribution of environmental factors through SHAP values. This interpretability module transforms the black box of machine learning into a decision-making basis that can guide spatial design.

[0154] The interpretability module constructs a 3D activation heatmap based on the output of the health prediction module to locate key spatiotemporal regions. First, given the target health indicator... Calculate the feature map of the final convolutional layer. gradient:

[0155] ;

[0156] in, The spatiotemporal volume normalization factor. For feature channels gradient weights, This is the channel index of the feature map, reflecting the contribution of that feature channel to the target health indicator. For feature channels The feature map of the final convolutional layer, This represents the total number of channels in the feature map (i.e., the feature dimension). Indicates the height of the spatial grid. Indicates the width of the spatial grid. This indicates the total number of time steps.

[0157] The heatmap is generated by weighting the channels:

[0158] ;

[0159] in, Indicates position The thermal value. The ReLU function filters out the negative contribution region. To enhance spatial continuity, anisotropic Gaussian filtering is used:

[0160]

[0161] in, For separable Gaussian kernels, spatial standard deviation Grid unit, time standard deviation Hour, This represents the smoothed 3D activation heatmap. This represents the original activation heatmap. This represents a 3D convolution operation. Ultimately, the heatmap is overlaid with a geographic raster to identify key areas affecting health.

[0162] Next, an environmental parameter contribution analysis is conducted, quantifying the impact of environmental factors based on SHAP values. For any predicted sample... Environmental feature subset The contribution value is calculated as follows:

[0163] ;

[0164] in This is the complete output function for the health prediction module, covering ordered category probability distributions and numerical predictions of physiological parameters. The total number of environmental features. Representation of features SHAP value, Indicates from the complete collection Remove features To reduce computational complexity, the kernel SHAP approximation algorithm is used, first generating a background sample set. , The original data was clustered using K-means; Indicates the first One background sample, Indicates the size of the background sample set.

[0165] Next, for each predicted sample Construct perturbation samples ; express In the environmental feature subset The value on, The first in the background sample library One sample in The value on, express The supplement to .

[0166] Finally, solve the weighted linear regression:

[0167] ;

[0168] in, Indicates the baseline forecast value. This indicates that the health prediction module is effective against perturbation samples. The predicted value, Indicates perturbation sample The 3D eigenvalues, perturbation samples kernel weight Emphasizing local interpretability, the final result generates a spatial contribution distribution map.

[0169] The visualization output is achieved through the overlay of the aforementioned 3D activated heatmap and the mapping of contribution factors, including heatmap rendering and SHAP contribution spatialization, ultimately generating a heatmap of key areas affecting health and a spatial contribution distribution map.

[0170] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the methods described in the foregoing embodiments.

[0171] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described in the foregoing embodiments.

[0172] The method provided by the foregoing embodiments of the present invention has the following advantages:

[0173] (1) Environment-behavior dynamic coupling modeling mechanism: Realize environment-behavior interaction modeling based on graph neural network and adaptive attention mechanism, by dynamically constructing the correlation weight matrix between environmental factors and student behavior patterns (environment weight matrix). It can analyze the impact of environmental parameters and students' dynamic behaviors on students' personal health, breaking through the limitations of traditional models that treat the environment and behavior independently.

[0174] (2) Spatiotemporal alignment method for multi-source heterogeneous data: By hierarchical discretization coding and multi-level fusion, combined with geospatial correction algorithm and adaptive interpolation technology, the problem of matching data from different sources such as fixed sensors, mobile devices and meteorological grids in the spatiotemporal dimension is solved, and the accurate coordination of static parameters such as building density and green space distribution with dynamic streaming data such as behavioral trajectory and environmental monitoring is achieved.

[0175] (3) Framework for Co-prediction and Explanation of Physiological and Mental Health: Construct a dual-branch decoding architecture of ordered regression of psychological indicators and prediction of physiological time series. Achieve synergistic optimization of the two types of health indicators through feature cross-sharing mechanism. Integrate three-dimensional heat map positioning and parameter contribution analysis technology to reveal the path of key environmental factors on health outcomes and form a quantitative model that can guide design.

[0176] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in this invention may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for intelligent assessment of the impact of a campus environment on individual health that incorporates spatio-temporal behavior, characterized in that, The spatial static features, the space-time behavior features, the environmental dynamic features and the physiological and psychological indexes are input into a multi-level health assessment model to obtain physiological index and psychological index prediction results, wherein the multi-level health assessment model comprises the following modules: A data preprocessing module is configured to perform the following operations: A GeoHash spatial coding system is constructed based on a set of geographic coordinate points, and continuous geographic space is discretized into grid cells with a configurable resolution through recursive spatial bisection; All spatial static features are aggregated in the grid cells through spherical weighting; For space-time behavior feature data, mobile trajectories are processed through multi-scale space-time tensorization to obtain behavior tensors; A static environment feature matrix is obtained based on the spatial static features, and a dynamic behavior feature tensor is obtained based on the behavior tensors; 2.The method of claim 1, wherein, The environmental dynamic features are aligned to achieve accurate mapping of the environmental parameter sequence and the target space-time grid; The final grid value is generated through confidence weighted integration. In the feature extraction module, the graph convolution network is configured to perform the following operations: node feature initialization is performed on the spatial static features, neighborhood features are aggregated, and then feature fusion is performed to obtain the spatial function distribution features. In the multi-modal fusion module, the environment-behavior coupling attention mechanism includes a direct environment correlation path, an indirect behavior modulation path and attention calculation, the direct environment correlation path is used to calculate the synergistic effect between environmental parameters and quantify the interaction strength of environmental factors themselves, the indirect behavior modulation path is used to capture the dynamic modulation effect of the environment on the behavior and generate space-time sensitive weights, and the attention calculation is used to fuse the output matrix of the direct environment correlation path and the space-time sensitive weights into an environmental weight matrix, perform attention calculation and combine the behavior features to obtain the environment-behavior coupling features. In the health prediction module, the space-time dependent coupling features are input to output the probability distribution of the psychological recovery order category through the psychological index prediction branch, and the physiological parameter value prediction is generated through the physiological index prediction branch, and the psychological index prediction branch and the physiological index prediction branch are cooperatively optimized through feature cross and dynamic weighting. A data preprocessing module is configured to perform the following operations: A GeoHash spatial coding system is constructed based on a set of geographic coordinate points, and continuous geographic space is discretized into grid cells with a configurable resolution through recursive spatial bisection; 3.The method of claim 2, wherein, In the alignment of the environmental dynamic characteristics, in the spatial dimension, for the target grid unit at time of the first class of environmental dynamic characteristics, a Kriging equation set is constructed based on an improved spatial variation function, and a weight vector is obtained by solving the Kriging equation set, wherein a direction variation parameter is introduced into the improved spatial variation function; in the time dimension, an adaptive Gaussian process interpolation is used to solve the sampling asynchronous problem. 4.The method of claim 1, wherein, All spatial static features are aggregated in the grid cells through spherical weighting; 5.The method of claim 1, wherein, For space-time behavior feature data, mobile trajectories are processed through multi-scale space-time tensorization to obtain behavior tensors; A static environment feature matrix is obtained based on the spatial static features, and a dynamic behavior feature tensor is obtained based on the behavior tensors; The environmental dynamic features are aligned to achieve accurate mapping of the environmental parameter sequence and the target space-time grid; The final grid value is generated through confidence weighted integration. In the feature extraction module, the graph convolution network is configured to perform the following operations: node feature initialization is performed on the spatial static features, neighborhood features are aggregated, and then feature fusion is performed to obtain the spatial function distribution features. In the multi-modal fusion module, the environment-behavior coupling attention mechanism includes a direct environment correlation path, an indirect behavior modulation path and attention calculation, the direct environment correlation path is used to calculate the synergistic effect between environmental parameters and quantify the interaction strength of environmental factors themselves, the indirect behavior modulation path is used to capture the dynamic modulation effect of the environment on the behavior and generate space-time sensitive weights, and the attention calculation is used to fuse the output matrix of the direct environment correlation path and the space-time sensitive weights into an environmental weight matrix, perform attention calculation and combine the behavior features to obtain the environment-behavior coupling features. 6.The method of claim 1, wherein, The dynamic graph neural network is configured to perform the following operations: An initial node state is taken as an environment-behavior coupling feature to construct a space-time graph; A hybrid mechanism of a gated recurrent unit and attention aggregation is adopted to update the node state, and a coupling feature with space-time dependence is output. 7.The method of claim 1, wherein, An input of the health prediction module is a space-time tensor, which is obtained based on the coupling feature with space-time dependence, and the health prediction module is configured to perform the following operations: A psychological index prediction branch first extracts a high-order psychological state representation through a multi-layer fully connected network, and then outputs a probability distribution of an ordered class of psychological recovery through an ordered regression model; A physiological index prediction branch reorganizes the space-time tensor into a physiological feature tensor through a time series convolution network, and then generates a physiological parameter value through a dilated causal convolution network to capture a multi-scale law; An output of the fully connected network, in which a high-order psychological state representation is encoded and information is not excessively compressed, and an output of the physiological index prediction branch for the ordered regression are spliced to obtain a feature cross term, which is used for information sharing and loss calculation between the psychological index prediction branch and the physiological index prediction branch to realize collaborative optimization.

8. The intelligent assessment method of the impact of a campus environment on individual health integrating spatio-temporal behavior according to any one of claims 1-7, characterized in that, An explainability module is further included, which generates a three-dimensional heat map to locate a health sensitive area based on a prediction result output by the health prediction module, and analyzes a contribution degree of a key environmental factor in combination with a SHAP value.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-8.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Campus green space ownership perception evaluation method and system based on multi-modal learning

    CN119862400A

  • Old people falling risk dynamic assessment and real-time early warning system and method based on multi-modal deep learning

    CN120913834A