Embodied Perception Assessment Method and System for Campus Green Spaces Based on Multimodal Learning
Through the multimodal learning method, the environmental perception, physiological response and psychological evaluation data are integrated, and the comprehensive environmental comfort score results are deeply integrated to generate, solving the limitations of the single-modal data analysis model in the existing technology, and achieving a comprehensive and accurate assessment of the campus space environment.
Patent Information
- Application Number
- CN202510358454.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing technology has limitations in the scientific assessment of campus spatial environments, and it is impossible to comprehensively and accurately evaluate the perceived impact of the spatial environment on people.
Using a multimodal learning method, environmental perception, physiological response and psychological evaluation data are integrated, and deeply integrated through visual feature encoding network, auditory feature encoding network, physiological and psychological collaborative encoder and global fusion module to generate comprehensive environmental comfort score results and explanatory decision-making basis.
A comprehensive assessment of the embodied perception of the green space on campus has been achieved, which improves the accuracy and scientificity of the assessment, and provides decision-making support for optimizing campus space design.
Smart Images

Figure CN119862400B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent environmental assessment / environmental behavior assessment, and particularly relates to a method and system for embodied perception assessment of campus green spaces based on multimodal learning. Background Art
[0002] The improvement of the sensory environment of campus spaces has become an important way to enhance space quality and is also an important research topic in the field of environmental perception. However, the current technological development in this field has limitations, which to a certain extent affects the accuracy and effectiveness of the scientific assessment of campus space environments.
[0003] On the one hand, most existing environmental assessment systems adopt a single-modal data analysis mode. This mode either relies solely on subjective questionnaires to evaluate individuals' psychological feelings or simply uses physiological indicators such as heart rate variability to reflect environmental stress. However, this method ignores the complex relationship between audiovisual environmental characteristics and multi-dimensional psychological and physiological parameters. There is a close association between audiovisual environmental characteristics and people's psychological reactions and physiological states. Currently, the lack of comprehensive analysis of this association results in the inability of existing assessment systems to comprehensively and accurately evaluate the perceptual impact of the space environment on people.
[0004] On the other hand, traditional multimodal fusion methods have obvious deficiencies in dealing with multi-source heterogeneous data. These methods usually use simple feature concatenation or weighted averaging to fuse data of different modalities, and fail to effectively capture the non-linear association between the temporal characteristics of the audiovisual environment and the dynamic response of physiological signals. This simple fusion method ignores the complex interaction relationship between data of different modalities, resulting in the fused data being unable to accurately reflect the impact of the space environment on people's embodied perception. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for embodied perception assessment of campus green spaces based on multimodal learning, which realizes a comprehensive assessment of the embodied perception of campus green spaces, improves the accuracy and scientific nature of the assessment, and provides decision-making support for optimizing campus space design, so as to solve at least one of the above-mentioned existing technical problems.
[0006] In the first aspect, the present invention provides a method for embodied perception assessment of campus green spaces based on multimodal learning, and the method specifically includes:
[0007] Obtain an environmental perception data set, a physiological response data set, and a psychological assessment data set, where the environmental perception data set includes a visual index quantization matrix and an acoustic time series tensor;
[0008] Input the visual metric quantization matrix into a visual feature encoding network for processing, input the acoustic time series tensor into an auditory feature encoding network for processing, and input the output results of the visual feature encoding network and the auditory feature encoding network into an audiovisual modality feature fusion module for fusion processing to obtain an environmental joint representation;
[0009] Input the physiological response data set and the psychological assessment data set into a physiological-psychological collaborative encoder for processing to obtain a physiological-psychological joint representation;
[0010] Input the environmental joint representation and the physiological-psychological joint representation into a global fusion module for deep fusion to obtain a final fusion feature;
[0011] Input the final fusion feature into a multi-task prediction channel for processing, and train and optimize it in combination with a loss function to obtain a comprehensive environmental comfort score result and an interpretive decision-making basis.
[0012] In a second aspect, the present invention provides a campus green space embodied perception evaluation system based on multi-modal learning. The system specifically includes:
[0013] A data acquisition module for acquiring an environmental perception data set, a physiological response data set, and a psychological assessment data set. The environmental perception data set includes a visual metric quantization matrix and an acoustic time series tensor;
[0014] An audiovisual processing module for inputting the environmental perception data set into a visual feature encoding network and an auditory feature encoding network for processing respectively, and inputting the output results of the visual feature encoding network and the auditory feature encoding network into an audiovisual modality feature fusion module for fusion processing to obtain an environmental joint representation;
[0015] A physiological-psychological processing module for inputting the physiological response data set and the psychological assessment data set into a physiological-psychological collaborative encoder for processing to obtain a physiological-psychological joint representation;
[0016] A global fusion module for inputting the environmental joint representation and the physiological-psychological joint representation into the global fusion module for deep fusion to obtain a final fusion feature;
[0017] An evaluation output module for inputting the final fusion feature into a multi-task prediction channel for processing, and training and optimizing it in combination with a loss function to obtain a comprehensive environmental comfort score result and an interpretive decision-making basis.
[0018] In a third aspect, the present invention provides a computer device, including: a memory, a processor, and a computer program stored on the memory. When the computer program is executed on the processor, it implements the campus green space embodied perception evaluation method based on multi-modal learning as described in the above method.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the method for embodied perception evaluation of campus green space based on multimodal learning as described in the above method.
[0020] Compared with the prior art, the present invention has at least one of the following technical effects:
[0021] 1. By integrating multimodal data such as environmental perception, physiological response, and psychological evaluation, the present invention realizes a comprehensive evaluation of the embodied perception of campus green space, improves the accuracy and scientificity of the evaluation, and provides decision-making support for optimizing campus space design.
[0022] 2. The present invention innovatively constructs a multimodal dynamic perception and interpretable evaluation system, showing significant technical advantages in the field of intelligent evaluation of environmental comfort. Through an adaptive weight allocation mechanism, it realizes the dynamic coupling of environmental parameters and physiological feedback, effectively capturing the non-linear correlation characteristics of multi-source information in complex scenarios; the multi-task joint reasoning architecture breaks through the semantic decoupling bottleneck of traditional models, synchronously outputting a comprehensive evaluation index and multi-dimensional interpretable components, significantly enhancing the traceability of decision-making basis; the hierarchical normalization fusion strategy overcomes the problem of cross-modal feature alignment, ensuring high-precision mapping of physical environment quantification indicators and subjective perception evaluations in a unified semantic space. Through closed-loop feedback optimization, this technical system provides intelligent decision-making support for optimizing campus space design.
[0023] 3. Based on the real-time weight allocation mechanism of a dual-channel neural network, the present invention constructs a dynamic feature fusion framework through the collaborative analysis of environmental physical parameters and human physiological indicators. This technology breaks through the traditional static weighting mode and adopts a joint measurement mechanism of environmental perception confidence and physiological feedback sensitivity to realize the autonomous adjustment of the contribution degrees of audiovisual features and biological signals. It strengthens the dominance of physical features under extreme environmental parameter conditions, while giving priority to considering human feedback information in a significant physiological stress state, effectively balancing the interactive effects of environmental objective attributes and subjective perceptions, and significantly enhancing the evaluation robustness in complex scenarios.
[0024] 4. The present invention constructs a three-level attention-driven feature interaction system to achieve progressive fusion from micro-feature matching to macro-semantic association. It captures the local correlation rules between visual parameters and pupil dynamics through fine-grained cross-modal dot product attention, models the non-linear mapping relationship between acoustic features and psychological scores using multi-head cross attention, and finally realizes decision-level feature optimization through a gated residual network. This architecture breaks through the limitations of traditional single-scale fusion, eliminates feature conflicts between modalities while retaining the spatio-temporal characteristics of the original data, and provides a unified expression space for the deep collaboration of multi-source heterogeneous data.
[0025] 5. The present invention designs a decoupled joint training framework, which decomposes the environmental comfort evaluation into a topological structure with a main decision-making network and multi-dimensional interpretation networks in parallel. Through feature space dissociation technology, independent modeling of visual contribution, acoustic comfort, and physiological fitness is achieved, and the orthogonality of each interpretation factor is ensured by combining the mutual information constraint mechanism. While maintaining the prediction accuracy, this system provides a physically interpretable decision-making basis for intelligent environmental regulation.
[0026] 6. The present invention uses a local feature extraction layer and an LSTM time series layer to extract multi-scale acoustic features from the acoustic time series tensor, captures the time-dependent relationship, and forms an accurate acoustic embedding, providing important information for the construction of environmental joint representation.
[0027] 7. The present invention realizes the effective fusion of visual embedding and acoustic embedding through a cross-modal Transformer encoder and a dynamic weight gating unit, forms an environmental joint representation containing audio-visual environment information, and improves the comprehensive understanding of environmental features.
[0028] 8. The present invention co-encodes the psychological assessment dataset and the physiological response dataset to form a physiological and psychological joint representation containing individual psychological and physiological states, providing a key basis for evaluating the impact of the environment on human perception.
[0029] 9. The present invention extracts multi-scale physiological time series features from the physiological response dataset through a dilated time series convolutional layer and a gated attention pooling layer, calculates the importance scores of each time step, and forms an accurate physiological state embedding.
[0030] 10. The present invention realizes the effective fusion of the physiological state embedding and the psychological state embedding by dynamically weighting and correcting them, and forms a joint representation containing the comprehensive physiological and psychological states of individuals.
[0031] 11. The present invention realizes the deep fusion of the environmental joint representation and the physiological and psychological joint representation through steps such as fine-grained feature matching, semantic-level relationship modeling, and decision-level dynamic aggregation, forming a final fusion feature containing rich information. At the same time, through environmental influence evaluation and human feedback compensation, the accuracy and practicality of the fusion feature are further improved.
[0032] 12. The present invention maps the final fusion feature to the evaluation space through a three-layer fully connected network in the main prediction channel, outputs the comprehensive comfort score result, providing a direct basis for the quantitative evaluation of environmental comfort; through the auxiliary channel, the interpretability decision-making basis analysis of the final fusion feature is carried out, and multi-dimensional interpretation factors are output, providing decision-making support for the interpretation and understanding of the evaluation results. Description of the Drawings
[0033] Figure 1It is a schematic flowchart of a method for embodied perception evaluation of campus green space based on multimodal learning provided by an embodiment of the present invention;
[0034] Figure 2 It is a schematic structural diagram of a multimodal learning model of a method for embodied perception evaluation of campus green space based on multimodal learning provided by an embodiment of the present invention;
[0035] Figure 3 It is a schematic structural diagram of a system for embodied perception evaluation of campus green space based on multimodal learning provided by an embodiment of the present invention;
[0036] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0037] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0038] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0039] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0040] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0041] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0042] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear at different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0043] In the embodiments of this application, the execution subject of the process includes a terminal device. The terminal device includes, but is not limited to: devices such as servers, computers, smart phones, and tablet computers that can execute the methods disclosed in this application. Figure 1 The flowchart of the method for embodied perception evaluation of campus green space based on multimodal learning disclosed in an embodiment of the present invention is shown as follows:
[0044] S101, obtain an environmental perception data set, a physiological response data set, and a psychological evaluation data set, where the environmental perception data set includes a visual index quantization matrix and an acoustic time series tensor.
[0045] In this embodiment, the three types of core index data required include:
[0046] (1) Environmental physical indicators: covering multimodal environmental characteristics of vision (green view rate, openness, greening coverage rate, color richness, water body ratio, tree species richness, facility richness, etc., 7-dimensional quantization parameters) and acoustics (sound pressure levels of composite bird songs, human voices, mechanical sounds, MFCC spectra);
[0047] (2) Physiological response indicators: including time series signals of eye movement (gaze count, gaze time, average blink count, pupil diameter), electroencephalogram (α / β waves), electrocardiogram (heart rate HR, heart rate variability HRV), and skin conductance (skin conductance level SCL);
[0048] (3) Psychological evaluation indicators: comprehensive scoring data of visual satisfaction (naturalness, aesthetics, extensibility, peacefulness, diversity, compatibility, undulation, etc., 7 dimensions), auditory satisfaction (joyfulness, pleasantness, clarity, disorder, richness, fluency, naturalness, etc., 7 dimensions), and audiovisual coordination degree based on factor analysis, constituting the psychological representation of the human body's perception of the environment.
[0049] Based on the above core metrics, environmental perception datasets, physiological response datasets, and psychological assessment datasets are collected. Among them, the environmental perception datasets include a visual metric quantization matrix and an acoustic time series tensor. The visual metric quantization matrix is a 7-dimensional environmental feature vector composed of green view rate (0-1 ratio value), openness (three-dimensional space index), greening coverage rate (0-1 ratio value), HSV color richness (hue-saturation-brightness statistic), water body proportion (0-1 ratio value), tree species richness (discrete count), and facility richness (discrete count). The acoustic time series tensor includes the time-domain waveforms of composite bird songs, human voices, and mechanical sounds (sampling rate ≥ 44.1 kHz) and their corresponding equivalent sound levels LAeq (dB(A)) and MFCC spectral features (24 dimensions / frame).
[0050] The physiological response datasets include an eye movement trajectory matrix, an electroencephalogram feature tensor, and an electrocardiogram-electrodermal combined tensor. The eye movement trajectory matrix includes multi-channel time series of fixation coordinate point sequences (x,y), fixation duration (ms), pupil diameter change curve (mm / s), and blink interval time (ms). The electroencephalogram feature tensor includes the power spectral density matrices of alpha waves (8-13 Hz) and beta waves (14-30 Hz) (64 leads × time window). The electrocardiogram-electrodermal combined tensor includes minute-level physiological parameter sequences of heart rate variability (LF / HF ratio), standard deviation of RR interval (ms), and skin conductance level (μS).
[0051] S102, input the visual metric quantization matrix into a visual feature encoding network for processing, input the acoustic time series tensor into an auditory feature encoding network for processing, and input the output results of the visual feature encoding network and the auditory feature encoding network into an audiovisual modality feature fusion module for fusion processing to obtain an environmental joint representation.
[0052] In this embodiment, the visual metric quantization matrix is input into the visual feature encoding network. This network includes structures such as an input layer, a convolutional layer, a pooling layer, and a fully connected layer, and can automatically learn and extract key features in the image. Through the layer-by-layer processing of the network, a visual embedding, that is, a low-dimensional vector representation containing image information, is finally output.
[0053] The acoustic time series tensor is input into the auditory feature encoding network. This network includes structures such as an input layer, an LSTM (long short-term memory network) layer, and a fully connected layer, and can process time series data and extract the time-dependent relationships therein. Through the layer-by-layer processing of the network, an acoustic embedding, that is, a low-dimensional vector representation containing sound information, is finally output.
[0054] Input the visual embedding and the acoustic embedding into the fusion module. This module can adopt methods such as attention mechanism, concatenation, weighted summation, etc. to fuse the two embeddings. Through the fusion process, finally output the environmental joint representation, that is, the comprehensive representation containing visual and auditory information.
[0055] In this embodiment, by separately processing visual and auditory data and fusing them into the environmental joint representation, the environmental characteristics of the green space can be more comprehensively reflected, which helps to improve the accuracy of the evaluation and makes the evaluation results more in line with the actual situation. Since visual and auditory data are separately processed and fused into the environmental joint representation, it is easier to analyze the influence of different senses on environmental comfort, which helps to enhance the interpretability of the evaluation results and makes the evaluation process more transparent and understandable.
[0056] S103, input the physiological response data set and the psychological evaluation data set into the physiological-psychological collaborative encoder for processing to obtain the physiological-psychological joint representation.
[0057] In this embodiment, the physiological-psychological joint representation integrates physiological and psychological information, providing researchers with a more comprehensive perspective to understand an individual's response under specific environmental stimuli, which helps to reveal the potential associations and interaction mechanisms between physiology and psychology. By training and optimizing the physiological-psychological collaborative encoder, the model can learn the complex relationship between physiological and psychological responses, which makes the model have higher accuracy in predicting an individual's future physiological or psychological responses.
[0058] S104, input the environmental joint representation and the physiological-psychological joint representation into the global fusion module for deep fusion to obtain the final fusion feature.
[0059] In this embodiment, the final fusion feature integrates environmental, physiological and psychological information, providing researchers with a more comprehensive perspective to understand an individual's overall response under specific environmental stimuli, which helps to reveal the potential associations and interaction mechanisms between the environment, physiology and psychology. By learning the deep associations between different representations, the global fusion module can more accurately predict an individual's overall response under future environmental stimuli, which is of great significance for formulating personalized intervention measures, improving an individual's adaptability and quality of life.
[0060] S105, input the final fusion feature into the multi-task prediction channel for processing, and train and optimize it in combination with the loss function to obtain the comprehensive evaluation result of environmental comfort and the interpretive decision-making basis.
[0061] In this embodiment, the fused features are input into a multi-task prediction channel. This channel includes two sub-tasks: one is to predict the comprehensive score of environmental comfort, and the other is to generate an explanatory decision basis. To handle these two tasks simultaneously, a multi-task learning model can be designed, which includes a shared hidden layer and two output layers corresponding to different tasks respectively. In the task of predicting the comprehensive score of environmental comfort, the output layer can be a regression model for outputting a continuous score value. In the task of generating an explanatory decision basis, the output layer can be a classification model or a sequence generation model for outputting discrete decision labels or descriptive texts.
[0062] To train the multi-task learning model, a comprehensive loss function needs to be designed.
[0063] For the main task loss , the mean squared error (MSE) is adopted for the comfort score regression task, and the formula is as follows:
[0064] ;
[0065] where the symbols and represent the true comfort score and the model-predicted comfort score of the th sample respectively, and the predicted score is calculated through the main prediction channel.
[0066] For the auxiliary task loss , the Smooth L1 loss is adopted for the explanatory factor regression task:
[0067] ;
[0068] where the symbols and represent the true sub-dimension score and the model-predicted sub-dimension score of the th explanatory factor respectively.
[0069] Based on the above, a logical consistency regularization is adopted to calculate the regularization term , and the regularization term constrains the mutual information between explanatory factors to prevent dimensional redundancy.
[0070] ;
[0071] Finally, a multi-objective optimization framework is adopted to balance the main and auxiliary tasks, and the total loss adopts weighted summation:
[0072] ;
[0073] Among them, when there is a significant deviation between the auxiliary score and the main evaluation result (the difference > 2 standard deviations), the model self-check mechanism is triggered to recalibrate the feature weights.
[0074] In terms of the optimizer, AdamW is used, with the initial learning rate and weight decay. For the learning rate scheduling, a cosine annealing strategy is adopted, with a period of 10 epochs and the minimum learning rate. During the training process, we use the backpropagation algorithm to update the weight parameters of the model to minimize the comprehensive loss function. Through continuous iterative training, the model will gradually learn how to predict the comprehensive score of environmental comfort and generate explanatory decision-making bases based on the input features. Finally, the trained model is applied to new data inputs to predict the comprehensive score of environmental comfort and generate explanatory decision-making bases. These results can help users better understand the comfort status of the indoor environment and make corresponding adjustments or improvements.
[0075] In this embodiment, through the multi-task learning model, the performance of two tasks can be optimized simultaneously, thereby improving the prediction accuracy of the comprehensive score of environmental comfort. In addition, the feature fusion technology also helps to comprehensively utilize information at different levels, further improving the prediction ability of the model. The task of generating explanatory decision-making bases enables the model to not only predict the results but also provide decision-making support, helping users better understand the prediction results of the model and make corresponding decisions. The multi-task prediction channel optimized by combining the loss function training can provide users with accurate and interpretable environmental comfort assessment results, thus enhancing the user experience and satisfaction.
[0076] In some embodiments, referring to Figure 2 , the audiovisual modality feature fusion module includes a cross-modal Transformer encoder and a dynamic weight gating unit; in the above step S102, the specific process of inputting the output results of the visual feature encoding network and the auditory feature encoding network into the audiovisual modality feature fusion module for fusion processing to obtain the environmental joint representation includes:
[0077] Input the visual index quantization matrix into the first input layer, and perform min-max standardization processing on the visual index quantization matrix to obtain the standardized visual index quantization matrix;
[0078] Use the first fully connected layer to map the standardized visual index quantization matrix to a high-dimensional space to obtain the original projection features, and the first fully connected layer satisfies:
[0079] ;
[0080] Among them, represents the first visual feature mapping result, represents the second visual feature mapping result, and The weight matrix representing the visual feature mapping, and respectively represent the bias vectors of each fully connected layer of the first fully connected layer, represents the GELU activation function;
[0081] The self-attention layer is adopted to capture the interaction relationship between each index of the normalized visual metric quantization matrix, forming attention features. The self-attention layer satisfies:
[0082] ;
[0083] ;
[0084] ;
[0085] wherein, 、 and respectively represent query, key, and value, 、 and respectively represent 、 and 's projection matrix, represents the single-head attention calculation result, represents the scaling factor, represents the attention feature, represents the concatenation operation, represents that the first-head attention is used to focus on low-order projection features, represents that the second-head attention is used to focus on high-order semantic features, represents that the third-head attention is used to focus on residual features and cross-head interactions, represents the fusion matrix of the multi-head attention output, represents the Softmax function;
[0086] The original projection features and the attention features are input into the first output layer for residual connection, and the visual embedding is output. The visual embedding is , wherein, represents the visual embedding, represents the layer normalization operation.
[0087] In this embodiment, the processing steps of the visual feature encoding network include an input layer, a fully connected projection, and a self-attention mechanism.
[0088] Input layer: 7-dimensional visual metric vector , corresponding to the green view rate 、openness Equal quantization parameters, processed by min-max normalization:
[0089] ;
[0090] Among them, is the original measurement value of the th visual index, and correspond to the theoretical or statistical maximum / minimum values of each index respectively, and are used for normalization, is the output result after normalization of the th index.
[0091] Fully connected projection: Map sparse features to a high-dimensional space through two layers of MLP, and the activation function is GELU:
[0092] ;
[0093] Among them, , both are weight matrices for visual feature mapping, mapping the input dimension 7 to a 128-dimensional hidden space respectively, and then expanding to 256 dimensions. , correspond to the bias vectors of the two layers of fully connected layers respectively, and are added to the result of multiplying with the weight matrix through the broadcasting mechanism. is the operation result of the 7→128-dimensional projection + GELU, and its function is to decouple low-order features and separate basic physical attributes, while is the operation result of the 128→256-dimensional projection + GELU, and its function is to perform high-order semantic abstraction, extract composite environmental features, and use them for non-linear enhancement; hierarchical mapping gradually separates index coupling through the intermediate 128-dimensional hidden space.
[0094] For the discrete visual parameters unique to this field (such as the green view rate being discrete sampling values), a phased progressive mapping strategy is proposed to solve the over-smoothing problem of traditional single-layer MLP on low-dimensional sparse data. Map the sparse 7-dimensional visual indices (such as green view rate, openness, etc.) to a 256-dimensional hidden space through two layers of MLP to achieve feature decoupling and information concentration. The first layer (7→128 dimensions) breaks through the representation bottleneck of traditional single-layer projection and separates low-order visual attributes; the second layer (128→256 dimensions) uses the GELU activation function to enhance the non-linear representation ability and extract high-order semantic features.
[0095] Self-attention mechanism: Capture the interaction relationship between indices through 3-head attention to generate attention features , the 256-dimensional features are evenly divided into 3 groups (85 dimensions per group), calculate attention in independent subspaces respectively, and then the attention outputs of each group are restored to 256 dimensions through concatenation (concat) and linear projection:
[0096] ;
[0097] ;
[0098] ;
[0099] in, , are the projection matrices of query, key, and value, respectively, mapping hidden features to the attention calculation space, and is a scaling factor, which is determined by the number of heads (3) and the feature dimension (256) of the multi-head attention mechanism and is used to stabilize the distribution of attention weights; is the result of single-head attention calculation, and is the fusion matrix of the multi-head attention output. The 3-head attention captures the interaction relationship of different types of indicators through parallel subspace modeling. Focusing on low-order projection features, it is used to quantify the linear relationship of indicators, such as calculating the complementarity between green view rate and openness; Focus on high-order semantic features to achieve perceptual semantic coupling, such as the synergistic effect of openness and color richness; The focus is on residual features and cross-head interactions to achieve dynamic environmental adaptation, that is, adjusting weights according to different situations, such as strengthening the weight of openness in densely built-up areas and strengthening the weight of green view rate in natural landscape areas.
[0100] Construct a 3-head self-attention layer in the 256-dimensional latent space through a dynamic weight matrix , quantify the potential correlation between visual indicators, and explore the implicit coupling rules between multiple indicators. Through the multi-head mechanism, local details (single-head focusing on color-space correlation) and global semantics (multi-head joint modeling landscape coordination) are separated.
[0101] Based on the above, the self-attention features and the original projection features are combined to output the residual connection. :
[0102] ;
[0103] Fully connected projection and self-attention form a complementary mechanism - the former solves feature sparsity, and the latter eliminates the empirical bias of manually preset weights. The two jointly realize end-to-end mapping from physical parameters to psychological semantics. The fully connected projection module and self-attention mechanism of this design form a specific encoding paradigm for environmental perception data through progressive feature decoupling and dynamic indicator relationship modeling.
[0104] In some embodiments, reference Figure 2, the auditory feature encoding network includes a second input layer, a local feature extraction layer, an LSTM time series layer, and a second output layer; in the above step S102, the process of inputting the acoustic time series tensor into the auditory feature encoding network specifically includes:
[0105] Input the sound pressure level time series and MFCC spectrum of the acoustic time series tensor into the second input layer, and splice them into joint acoustic features through a sliding window.
[0106] Extract multi-scale acoustic features from the joint acoustic features by using multi-layer dilated convolution of the local feature extraction layer, and the local feature extraction layer satisfies , where represents the output of the th layer of dilated convolution, represents the output of the th layer of dilated convolution, represents the convolution kernel, represents the number of channels of the convolutional layer, and represents a one-dimensional convolution operation, represents the ReLU activation function.
[0107] Capture the time dependence of the multi-scale acoustic features through the LSTM time series layer to form an acoustic embedding, and the LSTM time series layer satisfies:
[0108] ;
[0109] ;
[0110] where represents the concatenation result of the forward states at the th time step, represents the concatenation result of the backward states at the th time step, represents the concatenation result of the forward states at the th time step, represents the concatenation result of the backward states at the th time step, represents the output of the 3rd layer of dilated convolution;
[0111] The acoustic embedding is , where represents the acoustic embedding, represents the total number of windows generated after sliding window segmentation, represents the final output sequence of the LSTM time series layer;
[0112] Output the acoustic embedding through the second output layer.
[0113] In this embodiment, the processing steps of the auditory feature encoding network include input preprocessing, 1D-CNN local feature extraction, and LSTM temporal modeling.
[0114] Input preprocessing: Sound pressure level time series and MFCC spectra , which are segmented by a sliding window (window length: 2 s, step size: 1 s) and then concatenated into joint features . Among them is the total duration (number of time steps) of the acoustic signal, is the total number of windows generated after sliding window segmentation. The number of sampling points within a 2-second window length is ( is the sampling rate), the original signal is segmented into windows, is the number of step points; is the feature concatenation operation, which combines the sound pressure level (1D) and MFCC (20D) along the feature dimension into a 21D joint feature.
[0115] 1D-CNN local feature extraction: Three layers of dilated convolution (dilation rate = 1, 2, 4) are used to extract multi-scale acoustic features. The dilation rate increases (1 → 2 → 4) to gradually expand the receptive field, capturing instantaneous events, short-term patterns, and long-term regularities respectively. The output of each layer is as follows:
[0116] ;
[0117] Among them , the output is the output of the th layer of dilated convolution. The receptive field expands exponentially with the dilation rate (layer 1: 5 points, layer 2: 9 points, layer 3: 17 points), ; is the convolution kernel, is the number of channels in the convolutional layer, and 64 feature maps are output for each layer.
[0118] LSTM temporal modeling: Bidirectional LSTM captures long-term dependencies. The dimension of the single hidden state is 128, and the modeling formula is as follows,
[0119] ;
[0120] ;
[0121] Based on the above, bidirectional concatenation is performed to obtain the hidden state of the bidirectional LSTM at time step . After concatenating the forward and backward states, the dimension is 256, and the final output sequence is ; Gradually compress the time dimension (4 times) through convolutional strides to balance detail retention and computational efficiency, and improve the training speed.
[0122] Final acoustic embedding is obtained through temporal average pooling operation, which compresses the temporal features into global statistics, retaining the steady-state characteristics and dynamic patterns of the acoustic environment:
[0123] ;
[0124] The encoding network achieves fine-grained analysis of acoustic comfort through multi-scale time-frequency joint modeling. Among them, the dilated convolutional layer of 1D-CNN gradually expands the receptive field (5→9→17 time points) to capture multi-granularity features from transient noise to periodic sound events; bidirectional LSTM models the long-term evolution law of acoustic parameters to capture time dependencies, and temporal average pooling then condenses the temporal features into semantically clear acoustic embeddings to improve the robustness to sudden noises. At the same time, the 256-dimensional embedding is aligned with the visual embedding in dimension, providing compatibility guarantee for subsequent cross-modal fusion.
[0125] In some embodiments, referring to Figure 2 , the audiovisual modality feature fusion module includes a cross-modal Transformer encoder and a dynamic weight gating unit; in the above step S102, the specific process of inputting the output results of the visual feature encoding network and the auditory feature encoding network into the audiovisual modality feature fusion module for fusion processing to obtain an environmental joint representation includes:
[0126] Performing cross-modal attention calculation on the visual embedding and the acoustic embedding using a cross-modal Transformer encoder, and the cross-modal Transformer encoder satisfies:
[0127] ;
[0128] ;
[0129] ;
[0130] Among them, , and represent query, key, and value respectively, , and represent the projection matrices of , and respectively, represents the cross-modal attention weight matrix, Represents the fused feature result, represents the scaling factor;
[0131] The dynamic weight gating unit is used to adjust the modal contribution degrees of the visual embedding and the acoustic embedding to obtain an environmental joint representation, and the environmental joint representation satisfies:
[0132] ;
[0133] ;
[0134] where, represents the gating coefficient, represents the parameter matrix for vector compression of the visual embedding and the acoustic embedding, represents the Sigmoid activation function, represents the environmental joint representation.
[0135] In this embodiment, the processing steps of the audiovisual modal feature fusion module include a spatial alignment strategy and dynamic weight gating.
[0136] Spatial alignment strategy: Design a cross-modal Transformer encoder with the visual embedding as the Query and the acoustic embedding as the Key-Value:
[0137] ;
[0138] ;
[0139] ;
[0140] where , are the projection matrices of the query (Query), key (Key), and value (Value) of this part respectively, and the scaling factor is used to stabilize the cross-modal attention calculation; is the cross-modal attention weight matrix, quantifying the retrieval intensity of visual features on acoustic features; is the value projection of the acoustic feature, encoding the semantic expression of auditory information. This strategy breaks through the traditional symmetric cross-attention and designs a retrieval mode dominated by visual Query to improve the evaluation accuracy of landscape-dominated environments (such as campus environments). is the fused feature result, and the structure it adopts enables the residual features to be retained, injecting cross-modal correction amounts while retaining the visual basic features to avoid information loss.
[0141] The difference from the visual feature encoder is that the Query source here is the visual embedding , and the Key-Value source is the acoustic embedding ; The projection matrix different from the visual feature encoder is used for intra-modal feature recombination, and the function of the projection matrix corresponding to the audio-visual modal feature fusion module is to achieve cross-modal feature alignment and establish audio-visual semantic associations. The input sources are inconsistent (intra-modal feature interaction and inter-modal feature retrieval), and the weight learning objectives are also inconsistent (visual self-attention and cross-modal attention).
[0142] Compared with the traditional symmetric cross-attention method, the present invention adopts visual-dominated cross-attention, with only vision as the Query and audition as the Key-Value, reducing the computational complexity through one-way retrieval, being more suitable for campus scenes dominated by landscapes, and at the same time using visual semantic constraints to weight distribution to capture key audio-visual associations.
[0143] Dynamic weight gating: Adjust the modal contribution degree through the gating coefficient adaptive to environmental features, through the Sigmoid activation function , and output the gating coefficient :
[0144] ;
[0145] ;
[0146] Parameter matrix Compresses the concatenated 512-dimensional vector into a 1-dimensional scalar to achieve modal-level global weight allocation; the gating coefficient is dynamically generated by the environmental joint features instead of fixed weights, globally adjusting the audio-visual modal contribution degree to reduce the evaluation error of the model in the case of audio-visual conflicts, and finally outputting the environmental joint representation .
[0147] This part realizes the environment context adaptive fusion of cross-modal features through the spatial alignment strategy and dynamic weight gating. The cross-modal Transformer uses the visual embedding as the query vector to retrieve semantically related components from the acoustic embedding (such as strengthening the negative impact of low-frequency noise when the visual openness is high); the dynamic gating generates modal weights through environmental features and automatically adjusts the dominant modality in the audio-visual conflict scenario (such as landscape accompanied by traffic noise). The visual encoder (decoupled self-attention) and the auditory encoder (hierarchical temporal modeling) complement each other, achieve semantic alignment through the cross-modal fusion module, and construct a joint representation system for complex perceptual environments.
[0148] In some embodiments, referring to Figure 2 , the physiological and psychological collaborative encoder includes a psychological feature projection layer, a physiological temporal coding network, and a collaborative attention fusion layer; in the above step S103, the inputting of the physiological response data set and the psychological evaluation data set into the physiological and psychological collaborative encoder for processing to obtain a physiological and psychological joint representation specifically includes:
[0149] Input the psychological assessment data set into a psychological feature projection layer for processing to obtain a psychological state embedding, where the psychological feature projection layer satisfies:
[0150] ;
[0151] ;
[0152] ;
[0153] Among them, represents the first layer MLP of the psychological feature projection layer and is used to achieve primary feature decoupling, represents the second layer MLP of the psychological feature projection layer and is used to achieve high-order semantic abstraction, represents the psychological state embedding, , and respectively represent the weight matrices of each layer MLP of the psychological feature projection layer, , and respectively represent the bias vectors of each layer MLP of the psychological feature projection layer, represents the psychological index vector, represents the tanh activation function, represents the real matrix space of represents the real matrix space of represents the real matrix space of;
[0154] Input the physiological response data set into a physiological time series encoding network for processing to obtain a physiological state embedding;
[0155] Use a collaborative attention fusion layer to fuse the psychological state embedding and the physiological state embedding to obtain a physiological-psychological joint representation.
[0156] In this embodiment, the physiological-psychological collaborative encoder realizes the joint modeling of the human physiological response and the subjective psychological state through a time series feature extraction and dynamic weight allocation mechanism.
[0157] In the psychological feature projection layer, the psychological index vector includes visual satisfaction , auditory satisfaction , and audio-visual coordination . After Z-score standardization, it is input into a three-layer MLP:
[0158] ;
[0159] ;
[0160] ;
[0161] of the first layer realizes primary feature decoupling, maps 3D psychological indicators to a 16D space, and initially separates the coupling of subjective perceptions; of the second layer realizes high-order semantic abstraction, further extracts 32D features, and captures the potential interactions of psychological indicators. The third layer outputs the psychological state embedding , is the hyperbolic tangent activation function, which constrains the output to the interval to enhance feature stability. The psychological feature projection layer maps discrete psychological indicators (3D) to a continuous embedding space. Through a hierarchical mapping of 3→16→32→64 dimensions, it gradually separates the coupling effect of perceptual dimensions, and the nonlinear transformation decouples the implicit dimensions of subjective perception to improve semantic discrimination. The ReLU activation function enhances sparsity, and the tanh output layer suppresses the feature amplitude to avoid modal imbalance in the subsequent fusion stage.
[0162] In some embodiments, referring to Figure 2 , the physiological time series encoding network includes a third input layer, a dilated time series convolutional layer, and a gated attention pooling layer; further, inputting the physiological response data set into the physiological time series encoding network for processing to obtain a physiological state embedding specifically includes:
[0163] Input the physiological response data set into the third input layer and align it into a physiological time series tensor according to the sampling rate;
[0164] Use the dilated time series convolutional layer to extract multi-scale physiological time series features from the physiological time series tensor, and the dilated time series convolutional layer satisfies:
[0165] ;
[0166] where represents the output feature vector of the th dilated convolution of the dilated time series convolutional layer, represents the output feature vector of the th dilated convolution of the dilated time series convolutional layer, represents the output feature vector of the th dilated convolution of the dilated time series convolutional layer at time step , represents the th weight matrix of the th convolutional kernel, represents the dilation coefficient, represents the integer index, Denote the bias vector of the dilated convolution of the th layer,
[0167] and denote the convolutional kernel;
[0168] ;
[0169] ;
[0170] wherein, denotes the physiological state embedding, denotes the time step 's contribution weight to the physiological state embedding , denotes a fixed time step, denotes a variable time step taking values from 1 to , denotes the output feature vector of the 4th dilated convolution of the dilated temporal convolutional layer at time step , denote denotes the output feature vector of the 4th dilated convolution of the dilated temporal convolutional layer at time step , denotes the time step size, denotes a learnable weight, denotes 's transpose, denotes the 64-dimensional real matrix space.
[0171] In this embodiment, the processing steps of the physiological temporal encoding network include input structuring, dilated temporal convolution (TCN), and gated attention pooling.
[0172] Input structuring: Align multi-source physiological signals into a temporal tensor , where (4D eye movement + 2D EEG + 2D ECG + 2D GSR), and the time step size (corresponding to 2 minutes of data, sampled at 1 Hz).
[0173] Dilated temporal convolution (TCN): Use 4 layers of causal dilated convolutions to extract multi-scale features, and the dilation coefficient of each layer is in turn, the convolutional kernel size , and the number of channels is 64:
[0174] ;
[0175] wherein, is the output feature vector of the -th dilated convolution at time step . is the -th weight matrix of the convolution kernel of the -th layer, is an integer index for causal dilation constraint, which only depends on the current and historical time steps. Then is the bias vector of the -th dilated convolution, providing a baseline offset for each output channel and enhancing the robustness of the model to input distribution shifts. Finally, the output feature
[0176] is obtained, retaining the original temporal length. Gated attention pooling: Calculate the importance scores of each time step through learnable weights
[0177] ;
[0178] ;
[0179] is the output feature vector of the 4th layer of TCN at time step . is the contribution weight of time step to the final physiological embedding . The physiological temporal encoding network extracts temporal dynamic features from multi-modal physiological signals (eye movement, EEG, ECG, skin conductance). Causal dilated convolutions capture multi-granularity responses from transient events (such as pupil constriction, time scale 0.1 - 0.5 seconds) to long-period patterns (such as skin conductance baseline drift, time scale > 30 seconds); gated attention pooling adaptively focuses on key physiological events (such as eye movement saccades corresponding to attention shifts). This part is based on hybrid temporal modeling, integrating the local feature extraction ability of TCN and the global weight allocation of the attention mechanism to reduce the detection latency of sudden physiological events (such as sudden increase in heart rate); the physiological signal alignment strategy designs a unified sampling rate (1Hz) and temporal length (T = 120), and reduces the sampling rate difference of multiple devices through interpolation algorithms.
[0180] Furthermore, the fusion processing of the mental state embedding and the physiological state embedding to obtain a physiological-psychological joint representation specifically includes:
[0181] Dynamically weighting the physiological state embedding to obtain enhanced physiological features, and the enhanced physiological features satisfy:
[0182] ;
[0183] ;
[0184] Among them, represents the gating coefficient, represents the trainable parameter matrix, represents physiological feature enhancement, represents the element-wise multiplication operator, represents the Sigmoid activation function;
[0185] The mental state embedding is corrected to obtain a mental feature correction, and the mental feature correction satisfies:
[0186] ;
[0187] ;
[0188] Among them, represents the cross-modal attention weight matrix from physiological signals to mental features, represents the cross-modal correlation matrix, represents the transpose of the mental state embedding, represents the mental feature correction;
[0189] The physiological feature enhancement and the mental feature correction are fused to obtain a physiological-mental joint representation, and the physiological-mental joint representation is , where represents the physiological-mental joint representation, represents the layer normalization operation.
[0190] In this embodiment, the dynamic alignment between the mental cognitive state and the physiological response signal is achieved through a bidirectional interactive attention mechanism. The specific steps include:
[0191] Mental-guided physiological feature enhancement (forward path):
[0192] Construct a bidirectional gating unit, and splice the mental embedding vector with the physiological embedding vector and generate a gating coefficient matrix through a fully connected layer, and use the Sigmoid function to constrain the gating value in the interval [0,1] to form a feature selection mask; the dynamic weighting of physiological features is realized through the Hadamard product (element-wise multiplication). The following formulas are the generation of the gating coefficient and the physiological feature enhancement :
[0193] ;
[0194] ;
[0195] Among them, is a trainable parameter matrix that realizes the projection of a 128-dimensional concatenated vector into a 64-dimensional gating space, generates a feature-level local mask, and can perform fine-grained enhancement on specific frequency bands or time points of physiological signals; is an element-wise multiplication operator; is the Sigmoid activation function. The gating coefficient of this part functions to locally enhance the sensitive areas of physiological features and enhance the features related to mental states in physiological signals.
[0196] Physiological feedback mental feature correction (reverse path):
[0197] Construct a cross-modal attention matrix, calculate the soft alignment weights of physiological signals to mental metrics, and recalibrate the original mental embedding vector through the attention weights. The calculation of the attention weights is as follows, is the mental feature correction:
[0198] ;
[0199] ;
[0200] Among them, is the cross-modal correlation matrix that learns the potential mapping relationship between physiological and mental features; Softmax is the normalized exponential function to ensure that the sum of weights in each row is 1; represents the cross-modal attention weight matrix from physiological signals to mental features, which is used to model the dynamic correction relationship of physiological features to mental states, and finally realizes the output of mental feature correction .
[0201] Joint representation generation:
[0202] This part breaks through the traditional one-way fusion paradigm (physiology → psychology or psychology → physiology), fuses the outputs of the two-way paths, retains the original features through residual connections, and layer normalization suppresses the dimensional differences between modalities to generate a joint feature representation with cross-modal consistency . First, perform residual connections on the output vectors of the two paths and stabilize the training process through layer normalization (LayerNorm).
[0203] ;
[0204] The physiological and psychological collaborative encoder decouples through a bidirectional dynamic interaction mechanism and hierarchical feature decoupling, achieving a fine-grained modeling of the human body's multi-modal responses. Through a three-level fusion architecture consisting of a psychological projection layer (static semantics), a physiological TCN (dynamic time series), and collaborative attention (cross-modal interaction), multi-granularity spatio-temporal alignment is achieved. In response to the non-stationary characteristics of physiological signals (such as baseline drift in electrocardiogram signals), layer normalization replaces BatchNorm to improve training stability in few-shot scenarios. Its innovation is reflected in the cross-modal closed-loop correction strategy and the time series-semantic joint representation framework.
[0205] In some embodiments, referring to Figure 2 , the global fusion module includes a fine-grained feature matching unit, a semantic-level relationship modeling unit, a decision-level dynamic aggregation unit, an environmental influence evaluation unit, and a human body feedback compensation unit; in the above step S104, the inputting of the environmental joint representation and the physiological and psychological joint representation into the global fusion module for deep fusion to obtain the final fusion feature specifically includes:
[0206] Dimensionality increase of the physiological and psychological joint representation to obtain the dimensionality-increased physiological and psychological joint representation , where represents the physiological and psychological joint representation, represents the dimensionality-increased physiological and psychological joint representation, represents the dimensionality increase projection matrix, represents the GELU activation function;
[0207] After decomposing the environmental joint representation into a visual subspace and an auditory subspace using the fine-grained feature matching unit, respectively perform fine-grained feature matching with the dimensionality-increased physiological and psychological joint representation ;
[0208] Use the bidirectional cross-attention mechanism of the semantic-level relationship modeling unit to establish a semantic connection between the environmental joint representation and the physiological and psychological joint representation;
[0209] Use the gated residual network of the decision-level dynamic aggregation unit to adjust the contribution ratios of the environmental joint representation and the physiological and psychological joint representation respectively, and fuse the environmental joint representation and the physiological and psychological joint representation according to the contribution ratios, where the gated residual network satisfies:
[0210] ;
[0211] ;
[0212] where represents the real-time adjustment of the contribution ratios of the environment and human body features, represents the dynamic gating weight matrix, represents the environmental joint representation, Indicates the fusion feature of the th iteration, indicating the fusion feature of the th iteration;
[0213] The dual-channel neural network of the environmental influence evaluation unit is used to analyze the decision-making contribution degree of the environmental joint representation and the physiological and psychological joint representation, and obtain the environmental dominance coefficient. The dual-channel neural network satisfies:
[0214] ;
[0215] Among them, represents the environmental influence evaluation network in the dual-channel neural network, represents the human response evaluation network in the dual-channel neural network, represents the environmental dominance coefficient;
[0216] The human feedback compensation unit is used to correct the environmental dominance coefficient to obtain the human correction coefficient. The human correction coefficient satisfies where, represents the human correction coefficient, represents the transpose of the physiological and psychological joint representation after dimensionality increase, represents the ReLU activation function;
[0217] According to the environmental dominance coefficient and the human correction coefficient, dynamic weighted fusion is performed to obtain the final fusion feature. The final fusion feature satisfies:
[0218] ;
[0219] Among them, represents the final fusion feature, represents the fusion feature finally output by the decision-level dynamic aggregation unit, represents the layer normalization operation.
[0220] In this embodiment, this module realizes the deep fusion of the audio-visual environment features and the human physiological and psychological responses through a multi-level feature interaction and dynamic weight regulation mechanism, and finally outputs the comprehensive evaluation result of the campus environmental comfort. The specific technical implementation includes feature space alignment and calibration, a hierarchical feature interaction mechanism, and adaptive weight allocation.
[0221] To address the problem of cross-modal feature scale differences, a bidirectional projection and distribution constraint strategy is adopted. First, the physiological and psychological features are dimensionally elevated. The 64-dimensional collaborative representation is extended to 256 dimensions through non-linear mapping to match the dimensions of the audiovisual environment features. The GELU activation function is used to enhance the expressive ability:
[0222] ;
[0223] Among them is the high-dimensional collaborative representation, which maps the 64-dimensional joint features to the 256-dimensional space to enhance the cross-modal interaction ability; is the dimensional elevation projection matrix, which learns the non-linear mapping rules of physiological and psychological features to the high-dimensional space.
[0224] Secondly, environmental feature dimensionality reduction and screening are carried out: The key components in the audiovisual encoding are retained through an interpretable gating mechanism to generate a 128-dimensional refined representation. The gating mask is dynamically generated by sorting the feature importance, and the top 50% of the high-contribution features are retained.
[0225] The hierarchical feature interaction mechanism realizes multi-granularity information fusion by constructing a three-level attention network, including fine-grained feature matching, semantic-level relationship modeling, and decision-level dynamic aggregation.
[0226] Fine-grained feature matching: The audiovisual environment features are decomposed into visual and auditory subspaces, and dot-product attention calculations are performed with the physiological and psychological features respectively. This layer focuses on capturing cross-modal local correlations, such as the association pattern between the green view rate and heart rate changes.
[0227] Semantic-level relationship modeling: Through a bidirectional cross-attention mechanism, semantic connections between environmental physical parameters and human subjective perceptions are established. A multi-head design (4 heads) is adopted to calculate multiple groups of interaction relationships in parallel to improve the parsing ability of complex scenes.
[0228] Decision-level dynamic aggregation: Feature iterative optimization is achieved by designing a gated residual network, and the formula is as follows:
[0229] ;
[0230] ;
[0231] Among them, represents the dynamic gating weight matrix, which is used to learn the joint mapping relationship between environmental features and physiological and psychological collaborative representations and generate the gating coefficient , adjusts the contribution ratio of the environment and human body features in real time, and finally realizes the fusion feature of the th iteration, and gradually optimizes the cross-modal representation through residual connection; is the fused feature of the previous iteration: preserving the historical fusion state to achieve the continuity of feature evolution, and the number of iterations .
[0232] The adaptive weight allocation mechanism realizes the intelligent adjustment of the fusion weight by dynamically evaluating the interaction between environmental features and human responses. It specifically includes the following content:
[0233] Design a dual-channel neural network to analyze the decision-making contribution degrees of audiovisual environmental features and physiological and psychological responses respectively. Among them, the environmental evaluation network focuses on the physical significance of acoustic indicators (sound pressure level, spectral complexity) and visual parameters (green view rate, spatial openness). Analyze the decision-making influence of environmental features through a dual-channel neural network. When the sound pressure level > 65 dB or the green view rate < 25%, the system automatically increases the weight coefficient of environmental features to reflect the physical dominant effect under extreme environmental conditions. is the environmental dominance coefficient, dynamically calculated by the dual-channel neural network:
[0234] ;
[0235] Among them, represents the environmental influence evaluation network, quantifying the decision-making contribution degree of the audiovisual environment to comfort; represents the human response evaluation network, quantifying the correction ability of physiological and psychological characteristics to comfort.
[0236] Introduce a feature correlation compensation term, which is corrected based on the fluctuation range of physiological indicators to enhance the sensitivity of the model to the human stress state. When the coefficient of variation of physiological indicators > 0.3 (reflecting a significant stress state), activate the feedback compensation path. is the human correction coefficient:
[0237] ;
[0238] Among them is the adaptive gain parameter, and the initial value is obtained through learning historical data. This design ensures that when the user is in a high physiological stress state, the human response characteristics obtain additional weight compensation, avoiding evaluation biases caused by solely relying on environmental parameters.
[0239] The final fused feature is calculated by the following formula to achieve the dynamic weighted fusion of environmental features and human responses:
[0240] ;
[0241] Based on the weighted sum of environmental interaction features and human response features, the fused 256-dimensional features are normalized. This operation can stabilize the training process and make the model insensitive to feature scale changes.
[0242] In some embodiments, referring to Figure 2 , the multi-task prediction channel includes a main prediction channel and an auxiliary channel; in the above step S105, the process of inputting the final fusion feature into the multi-task prediction channel specifically includes:
[0243] Using the three-layer fully connected network of the main prediction channel to map the final fusion feature to the evaluation space, and outputting the comprehensive comfort score result, the comprehensive comfort score result satisfies:
[0244] ;
[0245] ;
[0246] ;
[0247] Wherein, represents the output of the first-layer fully connected network of the main prediction channel, represents the output of the second-layer fully connected network of the main prediction channel, represents the comprehensive comfort score result, , and respectively represent the weight matrices of each layer of the fully connected network, , and respectively represent the bias vectors of each layer of the fully connected network, represents the GELU activation function, represents the Dropout regularization function, and represents the Sigmoid function;
[0248] Using the auxiliary channel to perform an analysis of the interpretability decision basis for the final fusion feature, and outputting a multi-dimensional interpretation factor, the multi-dimensional interpretation factor satisfies:
[0249] ;
[0250] ;
[0251] ;
[0252] Wherein, represents the visual contribution degree analysis, represents the acoustic comfort analysis, represents the physiological adaptability analysis, , and respectively represent the weight matrices of the visual contribution degree analysis, the acoustic comfort analysis, and the physiological adaptability analysis, represents the ReLU activation function, represents the tanh activation function, represents the real matrix space of
[0253] In this embodiment, the main prediction channel maps the fused features to the evaluation space through a three-layer fully connected network, and outputs the comprehensive comfort score result , and the three-layer structure is as follows:
[0254] The first layer (256→128 dimensions): the high-order feature extraction layer, responsible for extracting high-order interaction features, using the GELU activation function to enhance the non-linear expression ability, and the 128-dimensional hidden layer features retain the key decision-making information:
[0255] ;
[0256] The second layer (128→64 dimensions): the robustness enhancement layer, performs feature dimensionality reduction, introduces Dropout (p = 0.3) to suppress overfitting and improve the generalization ability of the model:
[0257] ;
[0258] The output layer (64→1 dimension): the bisection mapping layer, uses the Sigmoid function ( ) to constrain the result within the range of [0, 1], and linearly extends it to a 0-10 score system.
[0259] ;
[0260] In the above formula, , , respectively represent the weight matrices of each layer, , and respectively correspond to the bias vectors of the three-layer fully connected layers. This part first maps the global fused features to a 128-dimensional hidden space, further reduces the dimension to a 64-dimensional robust feature space, and finally the third layer outputs the predicted comfort score value.
[0261] Then, a parallel branch network is constructed to output multi-dimensional interpretation factors to quantify the contributions of each dimension, and the auxiliary channel realizes index decoupling through feature decomposition technology.
[0262] Visual contribution analysis (0-5 points): Focus on analyzing the weighted impacts of indicators such as the green view rate and spatial layout, , and the calculation method is as follows. The special network focuses on the spatial frequency features (0.5-4 Hz) and the color histogram distribution, and the weight matrix is processed by L1 regularization:
[0263] ;
[0264] Acoustic comfort analysis (0 - 3 points): Quantify the negative effects of parameters such as background noise and sound field uniformity, , and the calculation method is as follows. Integrate global features and acoustic features :
[0265] ;
[0266] Physiological adaptability analysis (0 - 2 points): Reflect the matching degree between the human biological rhythm and environmental parameters, , and the calculation method is as follows. Based on the physiological and psychological features after dimensionality increase , constrain the output range through the hyperbolic tangent function:
[0267] ;
[0268] Referring to Figure 3 , an embodiment of the present invention provides a multi - modal learning - based embodied perception evaluation system 3 for campus green spaces. The system specifically includes:
[0269] A data acquisition module 301, configured to obtain an environmental perception data set, a physiological response data set, and a psychological evaluation data set. The environmental perception data set includes a visual index quantization matrix and an acoustic time - series tensor;
[0270] An audiovisual processing module 302, configured to input the environmental perception data set into a visual feature encoding network and an auditory feature encoding network for processing respectively, and input the output results of the visual feature encoding network and the auditory feature encoding network into an audiovisual modal feature fusion module for fusion processing to obtain an environmental joint representation;
[0271] A physiological and psychological processing module 303, configured to input the physiological response data set and the psychological evaluation data set into a physiological and psychological collaborative encoder for processing to obtain a physiological and psychological joint representation;
[0272] A global fusion module 304, configured to input the environmental joint representation and the physiological and psychological joint representation into the global fusion module for deep fusion to obtain a final fusion feature;
[0273] An evaluation output module 305, configured to input the final fusion feature into a multi - task prediction channel for processing, and train and optimize in combination with a loss function to obtain an environmental comfort comprehensive scoring result and an interpretive decision - making basis.
[0274] It can be understood that, as Figure 1The content in the embodiment of the method for embodied perception assessment of campus green space based on multimodal learning shown above is applicable to the embodiment of the system for embodied perception assessment of campus green space based on multimodal learning. The specific functions implemented by the embodiment of the system for embodied perception assessment of campus green space based on multimodal learning are the same as those in the embodiment of the method for embodied perception assessment of campus green space based on multimodal learning shown in Figure 1 and the beneficial effects achieved are also the same as those in the embodiment of the method for embodied perception assessment of campus green space based on multimodal learning shown in Figure 1 .
[0275] It should be noted that for the information interaction, execution process, etc. between the above systems, since they are based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, reference can be made to the method embodiment part, which will not be elaborated here.
[0276] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, which will not be elaborated here.
[0277] Referring to Figure 4 , the embodiment of the present invention also provides a computer device 4, including: a memory 402, a processor 401, and a computer program 403 stored on the memory 402. When the computer program 403 is executed on the processor 401, it implements the method for embodied perception assessment of campus green space based on multimodal learning as described in any one of the above methods.
[0278] The computer device 4 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art can understand that Figure 4The computer device 4 is merely an example and does not limit the computer device 4. It may include more or fewer components than those shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0279] The so-called processor 401 may be a central processing unit (CPU). The processor 401 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0280] In some embodiments, the memory 402 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In some other embodiments, the memory 402 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 4. Further, the memory 402 may also include both the internal storage unit and the external storage device of the computer device 4. The memory 402 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program, etc. The memory 402 may also be used to temporarily store the data that has been output or will be output.
[0281] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the method for embodied perception evaluation of campus green space based on multi-modal learning as described in any one of the above methods.
[0282] In this embodiment, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0283] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0284] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0285] In the embodiments disclosed in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0286] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
Claims
1. A campus green space embodied perception evaluation method based on multimodal learning, characterized by: The method specifically comprises: Acquire an environmental perception dataset, a physiological response dataset, and a psychological assessment dataset, wherein the environmental perception dataset includes a visual indicator quantization matrix and an acoustic time series tensor; Inputting the visual index quantization matrix into a visual feature coding network for processing, inputting the acoustic time series tensor into an auditory feature coding network for processing, and inputting the output results of the visual feature coding network and the auditory feature coding network into an audiovisual modality feature fusion module for fusion processing to obtain a joint representation of the environment; The audiovisual modality feature fusion module includes a cross-modal Transformer encoder and a dynamic weight gating unit; the output results of the visual feature encoding network and the auditory feature encoding network are input into the audiovisual modality feature fusion module for fusion processing to obtain a joint representation of the environment, specifically including: A cross-modal Transformer encoder is used to perform cross-modal attention calculation on visual embedding and acoustic embedding; Using a dynamic weight gating unit to adjust the modal contribution of the visual embedding and the acoustic embedding to obtain a joint representation of the environment; Inputting the physiological response data set and the psychological assessment data set into a physiological and psychological collaborative encoder for processing to obtain a physiological and psychological joint representation; the physiological and psychological collaborative encoder includes a psychological feature projection layer, a physiological temporal coding network and a collaborative attention fusion layer; Inputting the environmental joint representation and the physiological and psychological joint representation into a global fusion module for deep fusion to obtain a final fusion feature; The global fusion module includes a decision-level dynamic aggregation unit, an environmental impact assessment unit, and a human feedback compensation unit; the inputting of the environmental joint representation and the physiological and psychological joint representation into the global fusion module for deep fusion to obtain the final fusion feature specifically includes: The physiological and psychological joint representation is dimensionally upgraded, and fine-grained feature matching and semantic connection establishment are performed with the environmental joint representation; The gated residual network of the decision-level dynamic aggregation unit is used to adjust the contribution ratio of the environmental joint representation and the physiological and psychological joint representation, and the environmental joint representation and the physiological and psychological joint representation are fused according to the contribution ratio. The gated residual network satisfies: ; ; in, Indicates the contribution ratio of real-time adjustment of environment and human characteristics, represents the dynamic gating weight matrix, represents the joint representation of the environment, It represents the joint physiological and psychological representation after dimensionality increase. represents the fusion feature of the tth iteration, Indicates The fusion features of the iteration, represents a multi-layer perceptron, Represents the Sigmoid activation function; The dual-channel neural network of the environmental impact assessment unit is used to analyze the decision-making contribution of the environmental joint representation and the physiological and psychological joint representation to obtain the environmental dominance coefficient; The human body feedback compensation unit is used to correct the environment dominant coefficient to obtain the human body correction coefficient; Perform dynamic weighted fusion according to the environment dominant coefficient and the human body correction coefficient to obtain a final fusion feature; The final fusion features are input into the multi-task prediction channel for processing, and combined with loss function training optimization to obtain the comprehensive scoring result of environmental comfort and explanatory decision-making basis.
2. The method according to claim 1, characterized in that The visual feature encoding network includes a first input layer, a first fully connected layer, a self-attention layer and a first output layer; the step of inputting the visual indicator quantization matrix into the visual feature encoding network for processing specifically includes: Inputting the visual index quantization matrix into the first input layer, performing maximum and minimum normalization processing on the visual index quantization matrix, and obtaining a normalized visual index quantization matrix; Using a first fully connected layer to map the standardized visual index quantization matrix to a high-dimensional space to obtain original projection features; A self-attention layer is used to capture the interaction between each indicator of the standardized visual indicator quantization matrix to form an attention feature; The original projection features and the attention features are input into the first output layer for residual connection, and the visual embedding is output.
3. The method according to claim 2, characterized in that The auditory feature encoding network includes a second input layer, a local feature extraction layer, an LSTM timing layer, and a second output layer; the step of inputting the acoustic timing tensor into the auditory feature encoding network for processing specifically includes: Inputting the sound pressure level time series and MFCC spectrum of the acoustic time series tensor into the second input layer, segmenting and splicing into joint acoustic features through a sliding window; Extracting multi-scale acoustic features from the joint acoustic features using multi-layer dilated convolution of a local feature extraction layer; The temporal dependency of the multi-scale acoustic features is captured by the LSTM timing layer to form an acoustic embedding; The acoustic embedding is output through the second output layer.
4. The method according to claim 1, characterized in that: The step of inputting the physiological response data set and the psychological assessment data set into a physiological and psychological collaborative encoder for processing to obtain a physiological and psychological joint representation specifically includes: Inputting the psychological assessment data set into the psychological feature projection layer for processing to obtain a psychological state embedding; Inputting the physiological response data set into a physiological temporal coding network for processing to obtain a physiological state embedding; The psychological state embedding and the physiological state embedding are fused by using a collaborative attention fusion layer to obtain a joint physiological and psychological representation.
5. The method according to claim 4, characterized in that The physiological temporal coding network includes a third input layer, an expanded temporal convolution layer, and a gated attention pooling layer; The step of inputting the physiological response data set into a physiological temporal coding network for processing to obtain a physiological state embedding specifically includes: Inputting the physiological response data set into the third input layer and aligning it into a physiological time series tensor according to the sampling rate; Extracting multi-scale physiological timing features from the physiological timing tensor using an expanded timing convolutional layer; The importance scores of multi-scale physiological timing features at each time step are calculated through a gated attention pooling layer to obtain the physiological state embedding.
6. The method according to claim 5, characterized in that The fusing of the psychological state embedding and the physiological state embedding to obtain a physiological and psychological joint representation specifically includes: Dynamically weighting the physiological state embedding to obtain physiological feature enhancement; Modifying the psychological state embedding to obtain a psychological feature modification; The physiological feature enhancement and the psychological feature correction are fused to obtain a physiological and psychological joint representation.
7. The method according to claim 1, characterized in that The multi-task prediction channel includes a main prediction channel and an auxiliary channel; the inputting the final fusion feature into the multi-task prediction channel for processing specifically includes: A three-layer fully connected network of the main prediction channel is used to map the final fusion features to the evaluation space, and output a comprehensive comfort score result; An auxiliary channel is used to perform interpretable decision basis analysis on the final fusion features and output a multi-dimensional explanatory factor.
8. A campus green space embodied perception evaluation system based on multimodal learning that implements the method described in any one of claims 1 to 7, characterized in that: The system specifically comprises: A data acquisition module, used to acquire an environmental perception data set, a physiological response data set, and a psychological assessment data set, wherein the environmental perception data set includes a visual indicator quantization matrix and an acoustic time series tensor; An audio-visual processing module, used to input the environment perception data set into a visual feature coding network and an auditory feature coding network for processing, and input the output results of the visual feature coding network and the auditory feature coding network into an audio-visual modality feature fusion module for fusion processing to obtain a joint representation of the environment; A physiological and psychological processing module, used for inputting the physiological response data set and the psychological assessment data set into a physiological and psychological collaborative encoder for processing to obtain a physiological and psychological joint representation; A global fusion module, used for inputting the environmental joint representation and the physiological and psychological joint representation into the global fusion module for deep fusion to obtain a final fusion feature; The evaluation output module is used to input the final fusion features into the multi-task prediction channel for processing, and combine the loss function training optimization to obtain the comprehensive scoring result of environmental comfort and the explanatory decision basis.
Citation Information
Patent Citations
Cognitive state interpretable method based on brain-language-vision large model
CN119227819A
Autoencoder assisted radar for target identification
US20190383904A1