Intelligent construction and optimization method and system for text travel illumination scene based on self cognition
By analyzing tourist behavior and spatial structure through embodied cognitive models, lighting response zones are delineated and lighting parameters are optimized, solving the problem of insufficient dynamic response in the lighting design of cultural and tourism spaces, and achieving more precise lighting control and enhanced experience.
Patent Information
- Application Number
- CN202610031218.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-02-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing lighting designs for cultural and tourism spaces lack a dynamic response mechanism to meet the actual needs of tourists. They cannot adaptively adjust to the functional characteristics of different spaces and the behavioral characteristics of tourists, resulting in a discrepancy between the lighting effects and the cognitive needs of tourists, which affects the quality of the experience.
Based on the embodied cognition model, by acquiring spatial structure data and tourist behavior trajectory data, we extract motion dynamics features, calculate cognitive load levels, divide lighting response areas, establish lighting demand mapping relationships, generate lighting parameter control instructions, and perform iterative optimization in combination with feedback data.
It enables refined management and intelligent evolution of lighting scenarios, improving the accuracy of lighting design in responding to tourists' cognitive needs and the precision of system adaptation.
Smart Images

Figure CN121502233A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to artificial intelligence technology, and in particular to a method and system for intelligent construction and optimization of cultural and tourism lighting scenes based on embodied cognition. Background Technology
[0002] Currently, lighting design in cultural and tourism spaces mainly relies on designers' experience and static lighting standards, lacking a dynamic response mechanism to the actual needs of tourists. It cannot adaptively adjust according to the functional characteristics of different spatial areas and the behavioral characteristics of tourists, resulting in a discrepancy between the lighting effect and tourists' cognitive needs, which affects the overall experience quality of cultural and tourism spaces.
[0003] Existing smart lighting technologies primarily focus on the automated control and energy consumption optimization of lighting equipment, neglecting the intrinsic connections between tourists' physical movements, sensory interactions, and cognitive experiences within a space. These methods lack a deep understanding of the cognitive mechanisms behind tourists' behavioral trajectories, making it difficult to capture the changing patterns of tourists' cognitive load in different spatial areas. This results in lighting control lacking specificity and precision, and failing to effectively support the construction of immersive experiences in cultural and tourism spaces.
[0004] Existing lighting optimization methods lack an iterative mechanism based on actual usage feedback, making it difficult to continuously optimize lighting parameters based on visitor behavior data once they are set. Summary of the Invention
[0005] This invention provides a method and system for intelligent construction and optimization of cultural and tourism lighting scenes based on embodied cognition, which can solve the problems in the prior art.
[0006] A first aspect of this invention provides a method for intelligent construction and optimization of cultural and tourism lighting scenes based on embodied cognition, comprising: Acquire spatial structure data and tourist behavior trajectory data of the target cultural and tourism space; Based on the embodied cognition model, the spatial structure data and the tourist behavior trajectory data are jointly analyzed. Motion dynamics features are extracted from the tourist behavior trajectory data, and the cognitive load level of tourists is calculated by combining the spatial structure data to obtain cognitive feature representation results. The cognitive feature representation results include spatial perception response patterns and cognitive load level distribution. Based on the spatial perception response mode, the target cultural and tourism space is divided into multiple lighting response areas, and a lighting demand descriptor associated with the cognitive load level is established for each lighting response area to obtain a lighting demand mapping relationship. Based on the lighting demand mapping relationship and lighting constraints, lighting parameter control instructions are generated for each lighting response area. The lighting equipment is controlled based on the lighting parameter control instructions, and the embodied cognition model is iteratively optimized based on the tourist behavior feedback data collected after the lighting control.
[0007] The steps for jointly analyzing the spatial structure data and the tourist behavior trajectory data based on the embodied cognitive model to obtain cognitive feature representation results include: Motion dynamics features are extracted from the tourist behavior trajectory data. These motion dynamics features include motion rate parameters, motion direction change parameters, and pause pattern parameters. The pause pattern parameters are obtained by calculating the spatiotemporal clustering density of trajectory points. A decoding mapping relationship is established from motion dynamics features to embodied cognitive state. The decoding mapping relationship maps the motion dynamics features into multi-dimensional cognitive features including body perception, motion interaction and context embedding. The motion interaction dimension characterizes the interaction intensity between the body and space by analyzing the cooperative pattern of the rate of change of motion speed parameter and the change of motion direction parameter. Based on the multidimensional cognitive features and the spatial complexity in the spatial structure data, the cognitive load level of tourists is calculated. The spatial complexity includes the functional zoning density and visual information density within a unit space. The cognitive load level is obtained by fusing motion complexity, dwell frequency, and spatial complexity. The motion complexity is determined based on the variance of the motion rate parameter and the motion direction change parameter.
[0008] The embodied cognitive model includes: A spatial-motion joint coding layer is established through a bidirectional attention mechanism to create an interactive representation of the motion dynamics features and the spatial structure data. The attention weight of the motion dynamics features on the spatial region is used to identify the spatial focus of tourists, and the influence weight of spatial complexity on the motion dynamics features is used to predict environmentally induced motion changes. A multi-dimensional cognitive load decoding module decomposes the cognitive load level of tourists into visual cognitive load, navigation cognitive load and emotional cognitive load based on the multi-dimensional cognitive features. The visual cognitive load is calculated based on visual information density and gaze duration, the navigation cognitive load is calculated based on path complexity and turning frequency, and the emotional cognitive load is calculated based on pause patterns and movement smoothness. A lighting feedback-driven online learning module calculates an effectiveness score for the lighting adjustment based on visitor behavior feedback data collected after the lighting adjustment, and updates the parameters of the embodied cognition model based on the feedback.
[0009] The steps to establish a decoding mapping relationship between motion dynamics characteristics and embodied cognitive states include: A temporal encoding representation of motion dynamics features is constructed, which is obtained by performing a temporal convolution operation on the motion dynamics features. The temporal convolution operation extracts the change patterns of motion rate parameters and motion direction change parameters in the time dimension. Based on the temporal encoding representation, a body perception dimension feature vector is calculated, which represents the perceptual correlation strength between the tourist's body motion state and the spatial environment. Calculate the mutual information between the rate of change of motion speed parameter and the rate of change of motion direction parameter, and generate a motion interaction dimension feature vector based on the mutual information. The mutual information represents the degree of coordination between the change of motion speed and the change of direction. Based on the pause mode parameters and the spatial semantic information in the spatial structure data, a context embedding dimension feature vector is generated, wherein the spatial semantic information includes the functional attributes and cultural attributes of the space. The body perception dimension feature vector, the motion interaction dimension feature vector, and the context embedding dimension feature vector are concatenated to form a multidimensional cognitive feature.
[0010] Based on the spatial perception response pattern, the target cultural and tourism space is divided into multiple lighting response zones, and a lighting demand descriptor associated with the cognitive load level is established for each lighting response zone to obtain the lighting demand mapping relationship. The steps include: Based on the distribution of tourists' spatial focus in the spatial perception response mode, a perception activity index for the spatial area is extracted. The perception activity index represents the comprehensive level of tourists' perception interaction frequency and perception duration within a unit space. Based on the perceived activity index and the cognitive load level, the lighting response area is divided by a spatial clustering method. The spatial clustering method determines the area boundary based on the cognitive load similarity and spatial adjacency of adjacent spatial units. For each lighting response area, a lighting demand descriptor is constructed that includes a reference lighting parameter set and a cognitive load adjustment coefficient. The reference lighting parameter set is determined based on the average cognitive load level of the lighting response area, and the cognitive load adjustment coefficient is determined based on the standard deviation of the cognitive load level in the lighting response area. Establish a mapping relationship between lighting response areas and lighting demand descriptors to form a lighting demand mapping relationship.
[0011] The steps for generating lighting parameter control instructions for each lighting response area based on the lighting demand mapping relationship and lighting constraints include: Based on the spatial adjacency relationship of each lighting response region in the lighting demand mapping relationship, a spatial association graph between lighting response regions is constructed. The nodes of the spatial association graph represent lighting response regions, and the edge weights represent the cognitive load gradient between adjacent lighting response regions. The cognitive load gradient is obtained by calculating the ratio of the difference in the average cognitive load level of adjacent lighting response regions to the spatial distance. The lighting constraints include power constraints, color temperature adjustment range, and brightness adjustment range of the lighting equipment; For each lighting response region, initial target lighting parameters are calculated based on the reference lighting parameter set, and cross-regional coordinated adjustment is performed according to the edge weight distribution of the lighting response region in the spatial association graph. The cross-regional coordinated adjustment is achieved by minimizing the weighted sum of lighting parameter gradients between the lighting response region and its neighboring lighting response regions. The weight coefficient of the weighted sum of lighting parameter gradients is determined jointly based on the edge weights and the cognitive load adjustment coefficient of the lighting response region. The adjusted target lighting parameters are verified based on the lighting constraints, and lighting parameter control instructions are generated.
[0012] The steps for iteratively optimizing the embodied cognition model based on visitor behavior feedback data collected after lighting control include: Collect visitor behavior feedback data after executing lighting parameter control commands, extract motion dynamics characteristics after lighting adjustment, and generate embodied cognitive states of response after lighting adjustment; Obtain the baseline embodied cognitive state before lighting adjustment, construct a comparison sequence of embodied cognitive states before and after lighting adjustment, calculate the state trajectory offset vector of the baseline embodied cognitive state and the response embodied cognitive state in the multidimensional cognitive load space, and generate cognitive load response features; A model prediction deviation quantification index is constructed based on the cognitive load response characteristics. The model prediction deviation quantification index is determined by calculating the directional consistency and magnitude deviation between the cognitive load change trend predicted by the embodied cognitive model and the actual cognitive load change trend in the cognitive load response characteristics. The update gradient of the attention weight parameters of the spatial-motor joint coding layer in the embodied cognition model is determined based on the directional consistency, and the update gradient of the decoding coefficients of the multi-dimensional cognitive load decoding module in the embodied cognition model is determined based on the amplitude deviation.
[0013] A second aspect of this invention provides a system for intelligent construction and optimization of cultural and tourism lighting scenes based on embodied cognition, comprising: The first unit is used to acquire spatial structure data and tourist behavior trajectory data of the target cultural and tourism space; The second unit is used to jointly analyze the spatial structure data and the tourist behavior trajectory data based on the embodied cognitive model, extract motion dynamics features from the tourist behavior trajectory data, calculate the tourist's cognitive load level in combination with the spatial structure data, and obtain cognitive feature representation results. The cognitive feature representation results include spatial perception response patterns and cognitive load level distribution. The third unit is used to divide the target cultural and tourism space into multiple lighting response areas according to the spatial perception response mode, and to establish a lighting demand descriptor associated with the cognitive load level for each lighting response area to obtain a lighting demand mapping relationship. The fourth unit is used to generate lighting parameter control instructions for each lighting response area based on the lighting demand mapping relationship and lighting constraints, control the lighting equipment based on the lighting parameter control instructions, and iteratively optimize the embodied cognition model based on the tourist behavior feedback data collected after lighting control.
[0014] A third aspect of the present invention, An electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0015] Fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0016] This invention introduces embodied cognition theory into the intelligent construction process of cultural and tourism lighting scenarios, enabling an understanding of lighting needs from the perspective of tourists' physical perception and spatial interaction. This breaks through the limitations of traditional lighting design, which only focuses on physical light environment parameters. This method can deeply explore tourists' cognitive processing in specific spaces, transforming abstract psychological perceptions into quantifiable cognitive load indicators. It provides a theoretical basis for lighting design that is more in line with the laws of human perception, allowing lighting solutions to more accurately meet the actual experience needs of tourists.
[0017] This invention establishes a dynamic mapping relationship between spatial structure, visitor behavior and lighting needs. By extracting motion dynamics features and combining them with spatial perception response modes, it intelligently divides the lighting response area, thereby achieving refined management of lighting scenes.
[0018] This invention constructs a closed-loop feedback optimization mechanism to achieve intelligent evolution of lighting scenarios. Over long-term use, the system's adaptation accuracy and service quality will continuously improve. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the intelligent construction and optimization method for cultural and tourism lighting scenes based on embodied cognition, as described in an embodiment of the present invention. Figure 2 A flowchart for calculating cognitive feature representation results based on spatial structure data and tourist behavior trajectory data. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0022] Figure 1 This is a flowchart illustrating the intelligent construction and optimization method for cultural tourism lighting scenes based on embodied cognition, as described in an embodiment of the present invention. Figure 1 As shown, the method includes: Acquire spatial structure data and tourist behavior trajectory data of the target cultural and tourism space; Based on the embodied cognition model, the spatial structure data and the tourist behavior trajectory data are jointly analyzed. Motion dynamics features are extracted from the tourist behavior trajectory data, and the cognitive load level of tourists is calculated by combining the spatial structure data to obtain cognitive feature representation results. The cognitive feature representation results include spatial perception response patterns and cognitive load level distribution. Based on the spatial perception response mode, the target cultural and tourism space is divided into multiple lighting response areas, and a lighting demand descriptor associated with the cognitive load level is established for each lighting response area to obtain a lighting demand mapping relationship. Based on the lighting demand mapping relationship and lighting constraints, lighting parameter control instructions are generated for each lighting response area. The lighting equipment is controlled based on the lighting parameter control instructions, and the embodied cognition model is iteratively optimized based on the tourist behavior feedback data collected after the lighting control.
[0023] In one optional implementation, the step of jointly analyzing the spatial structure data and the tourist behavior trajectory data based on an embodied cognitive model to obtain cognitive feature representation results includes: Motion dynamics features are extracted from the tourist behavior trajectory data. These motion dynamics features include motion rate parameters, motion direction change parameters, and pause pattern parameters. The pause pattern parameters are obtained by calculating the spatiotemporal clustering density of trajectory points. A decoding mapping relationship is established from motion dynamics features to embodied cognitive state. The decoding mapping relationship maps the motion dynamics features into multi-dimensional cognitive features including body perception, motion interaction and context embedding. The motion interaction dimension characterizes the interaction intensity between the body and space by analyzing the cooperative pattern of the rate of change of motion speed parameter and the change of motion direction parameter. Based on the multidimensional cognitive features and the spatial complexity in the spatial structure data, the cognitive load level of tourists is calculated. The spatial complexity includes the functional zoning density and visual information density within a unit space. The cognitive load level is obtained by fusing motion complexity, dwell frequency, and spatial complexity. The motion complexity is determined based on the variance of the motion rate parameter and the motion direction change parameter.
[0024] Combination Figure 2 The flowchart illustrating the calculation of cognitive feature representation results based on spatial structure data and visitor behavior trajectory data explains that visitor behavior trajectory data is acquired through a depth camera array deployed at the top of the space, with a sampling frequency set to 10 frames per second. Each trajectory point record includes a timestamp, two-dimensional coordinate position, and movement direction angle. The coordinate system is established with the space entrance as the origin, achieving an accuracy of 0.1 meters. Spatial structure data is collected using a 3D laser scanner, including the geometric boundaries of the space, wall positions, functional zoning information, and the spatial distribution of visual information elements. Functional zoning information labels the attribute type of each area, such as exhibition area, rest area, and passageway area. Visual information elements include the spatial coordinates and size parameters of exhibit positions, signage positions, and decorative elements. The collected trajectory data undergoes noise filtering to remove invalid trajectory points with a dwell time of less than 3 seconds, and linear interpolation is used to complete trajectory breaks to ensure the continuity and integrity of the trajectory.
[0025] Motion dynamics features were extracted. The motion speed parameter was obtained by dividing the Euclidean distance between adjacent trajectory points by the time interval, measured in meters per second. A 5-second moving average was used as the time window to eliminate instantaneous fluctuations. The motion direction change parameter was obtained by calculating the angle between the motion direction vectors at adjacent time points, ranging from 0 to 180 degrees. An angle exceeding 30 degrees was marked as a significant direction change event. The extraction of pause pattern parameters relied on spatiotemporal clustering analysis of the trajectory points. A density-based spatial clustering algorithm was used to cluster the trajectory points, with a cluster radius of 1.5 meters, a minimum number of cluster points of 5, and a time window of 10 seconds. For each cluster, its spatiotemporal cluster density was calculated as a pause pattern parameter. The density value was obtained by dividing the number of trajectory points within the cluster by the product of the spatial area covered by the cluster and the time span. A higher density value indicated a more significant pause behavior at that location. The extracted motion dynamics features are organized in the form of a time series. Each time step corresponds to a set of motion rate parameter values, motion direction change parameter values, and pause mode parameter values, forming a feature sequence with a length equal to the time span divided by the sampling interval.
[0026] A multi-layer neural network structure is constructed as the decoder. The input layer receives time-series data of motion dynamics features, and the hidden layer contains three parallel branches to process feature extraction in three dimensions: body perception, motion interaction, and contextual embedding. The body perception dimension feature vector is generated by extracting motion rhythm patterns through convolution operations on the motion rate parameter time series. The convolution kernel size is set to 5 time steps, the stride is 1 time step, and the output feature vector dimension is 64. The motion interaction dimension feature vector is generated by first calculating the rate of change of the motion rate parameter, obtained by dividing the difference in motion rate parameters between adjacent time steps by the time interval. Then, the mutual information between the motion rate parameter rate of change sequence and the motion direction change parameter sequence is calculated. The mutual information is obtained by calculating the information entropy difference after estimating the joint probability distribution and marginal probability distribution of the two sequences. A larger mutual information indicates a higher degree of coordination between rate change and direction change. Based on the mutual information value, the motion rate parameter rate of change and motion direction change parameter are weighted and fused to generate the motion interaction dimension feature vector. The weight coefficients are determined after normalization based on the mutual information, and the feature vector dimension is 64. The generation of context embedding dimension feature vectors requires combining pause pattern parameters and spatial semantic information. For each pause position, the density of the corresponding functional attribute and surrounding visual information elements in the spatial structure data is queried. The pause pattern parameter value, functional attribute encoding vector, and visual information density value are concatenated and mapped through a fully connected layer to form a context embedding dimension feature vector with a dimension of 64. The feature vectors of the three dimensions are concatenated to form a multidimensional cognitive feature vector with a dimension of 192, which serves as a representation of the embodied cognitive state.
[0027] Spatial complexity indices are extracted from spatial structure data. The calculation of functional zoning density involves dividing the spatial structure data into a 10-meter square grid. For each grid cell, the geometric center points of all functional zones in the spatial structure data are traversed, and it is determined whether each center point falls within the current grid cell. The number of regions with different functional attributes within the grid is counted, with the count incremented by 1 for each region with a different functional attribute. Functional attributes include categories such as display areas, rest areas, passageways, service areas, and interactive areas. The functional zoning density value equals the number of regions with different functional attributes counted within that grid cell, ranging from 0 to 5. When there are no functional zones within the grid, the density value is 0; when there are 5 or more regions with different functional attributes within the grid, the density value is capped at 5. Visual information density is obtained by statistically analyzing the number and area proportion of visual information elements within a unit of space, including exhibits, signage, and decorative elements. The density value is calculated by multiplying the number of elements by their area proportion using a weighted sum. The weighting coefficient is set according to the visual attractiveness of the element type: exhibits have a weight of 0.5, signage has a weight of 0.3, and decorative elements have a weight of 0.2. Motion complexity is determined based on the variances of motion rate and direction change parameters. The variances of motion rate and direction change parameters are calculated separately within a 5-second time window, and the average of these variances after normalization is taken as the motion complexity value. Dwell frequency is obtained by counting the number of pause events per unit time, with the unit time set to 1 minute. A pause event is defined as the moment when the pause mode parameter value exceeds a preset threshold, which is set to 0.5. Cognitive load level is obtained by weighted fusion of motion complexity, dwell frequency, and spatial complexity, with weighting coefficients set to 0.4, 0.3, and 0.3, respectively. The fused cognitive load level ranges from 0 to 1, with higher values indicating higher cognitive load.
[0028] This invention establishes a systematic decoding mapping relationship between motion dynamics characteristics and embodied cognitive states, transforming tourists' physical movement behavior into quantifiable cognitive load levels. This provides a scientific basis for the precise control of lighting systems in cultural and tourism spaces and improves the accuracy of lighting design in responding to tourists' cognitive needs.
[0029] In one alternative implementation, the embodied cognitive model includes: A spatial-motion joint coding layer is established through a bidirectional attention mechanism to create an interactive representation of the motion dynamics features and the spatial structure data. The attention weight of the motion dynamics features on the spatial region is used to identify the spatial focus of tourists, and the influence weight of spatial complexity on the motion dynamics features is used to predict environmentally induced motion changes. A multi-dimensional cognitive load decoding module decomposes the cognitive load level of tourists into visual cognitive load, navigation cognitive load and emotional cognitive load based on the multi-dimensional cognitive features. The visual cognitive load is calculated based on visual information density and gaze duration, the navigation cognitive load is calculated based on path complexity and turning frequency, and the emotional cognitive load is calculated based on pause patterns and movement smoothness. A lighting feedback-driven online learning module calculates an effectiveness score for the lighting adjustment based on visitor behavior feedback data collected after the lighting adjustment, and updates the parameters of the embodied cognition model based on the feedback.
[0030] For example, the space-motion joint coding layer of the embodied cognition model receives motion dynamics features and spatial structure data as input, and establishes an interactive representation of the two through a bidirectional attention mechanism. Motion dynamics features include time series of motion rate parameters, motion direction change parameters, and pause pattern parameters, with a feature vector dimension of 128 at each time step. Spatial structure data includes the geometric location, functional attribute encoding, and spatial complexity value of spatial regions, with each spatial region represented as a feature vector of dimension 64. The bidirectional attention mechanism includes two branches: motion-to-space attention calculation and space-to-motion attention calculation. The attention weight of motion on spatial regions is calculated by performing a dot product similarity calculation between the motion dynamics feature vector of the current time step and the feature vectors of all spatial regions. The similarity value is converted into an attention weight through a softmax normalization function, with weight values ranging from 0 to 1, and the sum of the weights of all spatial regions being 1. Spatial regions with an attention weight greater than 0.3 are identified as the spatial focus of the visitor's attention, and the functional attributes and location information of this region are recorded. The weighting of the influence of spatial complexity on motion dynamics features is calculated by weighting the spatial complexity value of each spatial region with the motion dynamics feature vector at the current time step. The weighting matrix has a dimension of 64×128 and is implemented through a fully connected layer. The influence weights are used to predict environment-induced motion changes. When the spatial complexity value is greater than 2.0, the influence weight is multiplied by a magnification factor of 1.5, representing the amplifying effect of the complex environment on the motion pattern. The output of the bidirectional attention mechanism is a fused interaction representation vector with a dimension of 256, containing the correlation information between motion features and spatial features.
[0031] The multidimensional cognitive load decoding module receives a multidimensional cognitive feature vector as input. This vector has 192 dimensions and includes features across three dimensions: body perception, motion interaction, and contextual embedding. The decoding module calculates visual cognitive load, navigational cognitive load, and emotional cognitive load through three parallel fully connected layer branches. Visual cognitive load calculation relies on two input parameters: visual information density and gaze duration. Visual information density is extracted from spatial structure data, specifically obtained by statistically analyzing the number and area proportion of visual information elements within a 5-meter radius of the visitor's current location, including exhibits, signs, and decorative elements. The density value is calculated by weighted summation of the element quantity multiplied by the area proportion, with exhibits having a weight of 0.5, signs 0.3, and decorative elements 0.2. The density value ranges from 0 to 5. Gaze duration is obtained by analyzing pause pattern parameters. A valid gaze event is defined as a pause pattern parameter value exceeding 0.5 and a duration exceeding 5 seconds. Gaze duration equals the cumulative time of continuous gaze events. Visual cognitive load is calculated by multiplying visual information density by a normalized fixation duration value, with a normalization range of 0 to 1. When visual information density exceeds 4.0 or fixation duration exceeds 30 seconds, the visual cognitive load value is set to the upper limit of 1. Navigation cognitive load is calculated based on two input parameters: path complexity and turning frequency. Path complexity is calculated by dividing the number of switches between different spatial regions along the trajectory path by the total duration, measured in times per minute. Turning frequency is calculated by dividing the number of events where the direction of motion changes by more than 30 degrees by the total duration, also measured in times per minute. The navigation cognitive load value is calculated as a weighted sum of path complexity and turning frequency, with weighting coefficients of 0.6 and 0.4, respectively, and the result is normalized to the range of 0 to 1. The calculation of emotional cognitive load relies on two input parameters: pause pattern and movement fluency. The pause pattern parameter is directly obtained from motion dynamics characteristics. This parameter is calculated through density-based spatial clustering analysis of trajectory points, with a cluster radius set to 1.5 meters, a minimum number of cluster points set to 5, and a time window set to 10 seconds. The pause pattern parameter value = number of trajectory points within a cluster / (spatial area covered by the cluster × time span). A higher density value indicates a more significant pause. Movement fluency is represented by the reciprocal of the standard deviation of the movement rate parameter. A smaller standard deviation indicates smoother movement and a higher fluency value. The emotional cognitive load value is calculated by multiplying the pause pattern parameter value by 0.5 and adding the reciprocal of the movement fluency value by 0.5, then normalizing to the range of 0 to 1. When the pause pattern parameter exceeds 0.8 or the movement fluency value is below 0.2, an emotional cognitive load value exceeding 0.7 indicates a high emotional load state.
[0032] The lighting feedback-driven online learning module initiates a data acquisition process after the lighting system completes parameter adjustments. It collects visitor behavior feedback data within a 30-120 second time window following the adjustment. This feedback data includes adjusted movement rate parameters, pause pattern parameters, and spatial focus distribution. Spatial focus distribution is obtained by statistically analyzing the attention weight distribution of visitors to each spatial area within the adjusted time window. For each spatial area, the average attention weight is calculated across all time steps. Areas with an average weight exceeding 0.3 are marked as focus areas. The spatial focus distribution records the location identifiers and corresponding average attention weight values of all focus areas. The effectiveness score is calculated by comparing the change in cognitive load levels before and after the adjustment. The change equals the adjusted cognitive load level minus the unadjusted cognitive load level. A negative change with an absolute value exceeding 0.1 is defined as effective improvement. The effectiveness score is equal to the absolute value of the change multiplied by 10, ranging from 0 to 10. An effectiveness score below 3.0 is considered an ineffective or negative adjustment, triggering a parameter rollback mechanism. The online learning module updates the parameters of the embodied cognition model based on performance scores, employing a gradient descent optimization algorithm with a learning rate set to 0.001. When the performance score exceeds 7.0, the learning rate is increased to 0.002 to accelerate parameter optimization. Model parameter updates include the attention weight matrix of the spatial-motor joint encoding layer and the weight matrix of the fully connected layer of the multi-dimensional cognitive load decoding module. Parameter updates are performed in batches of 10 every 10 effective improvement samples. The model parameters are saved as versioned files with incrementing version numbers. The current version of the parameters is backed up before each update. If the performance score decreases after three consecutive updates, a version rollback is performed, restoring the parameters to the historical version with the highest performance score.
[0033] The training process of the embodied cognition model uses historically collected tourist behavior trajectory data and corresponding lighting adjustment records as training samples, with a minimum of 5000 samples. The training data is divided into training and validation sets in an 8:2 ratio. Supervised learning is employed, with the loss function being the mean squared error between the predicted cognitive load level and the actual labeled cognitive load level. The optimizer uses the Adam algorithm, with an initial learning rate of 0.001, a batch size of 32, and 100 training epochs. Model performance is evaluated on the validation set after each training epoch. Early stopping is triggered when the validation set loss fails to decrease for five consecutive epochs, terminating the training. Model initialization parameters use the Xavier initialization method, with the initial values of the weight matrix randomly sampled from a normal distribution with a mean of 0 and a standard deviation of 0.1. The trained model parameters serve as the initial parameters for the online learning module, supporting subsequent incremental updates based on feedback data.
[0034] This invention achieves deep interactive representation of space and motion through a bidirectional attention mechanism, decomposes cognitive load into three quantifiable dimensions: vision, navigation, and emotion, and continuously optimizes model parameters through an online learning module, thereby improving the lighting system's perception accuracy and responsiveness to tourists' cognitive states.
[0035] In one alternative implementation, the step of establishing the decoding mapping relationship from motion dynamics features to embodied cognitive states includes: A temporal encoding representation of motion dynamics features is constructed, which is obtained by performing a temporal convolution operation on the motion dynamics features. The temporal convolution operation extracts the change patterns of motion rate parameters and motion direction change parameters in the time dimension. Based on the temporal encoding representation, a body perception dimension feature vector is calculated, which represents the perceptual correlation strength between the tourist's body motion state and the spatial environment. Calculate the mutual information between the rate of change of motion speed parameter and the rate of change of motion direction parameter, and generate a motion interaction dimension feature vector based on the mutual information. The mutual information represents the degree of coordination between the change of motion speed and the change of direction. Based on the pause mode parameters and the spatial semantic information in the spatial structure data, a context embedding dimension feature vector is generated, wherein the spatial semantic information includes the functional attributes and cultural attributes of the space. The body perception dimension feature vector, the motion interaction dimension feature vector, and the context embedding dimension feature vector are concatenated to form a multidimensional cognitive feature.
[0036] For example, the temporal encoding representation of motion dynamics features is obtained by performing a temporal convolution operation on motion rate parameters and motion direction change parameters. The motion rate parameters and motion direction change parameters are organized in time series form, with a sampling frequency of 10 frames per second. Each time step records a set of motion rate parameter values and motion direction change parameter values. The temporal convolution operation uses a one-dimensional convolution kernel to extract features from the time series. The kernel size is set to 5 time steps, the stride is 1 time step, and the number of convolution kernels is 32. The convolution kernel slides along the temporal dimension, performing a weighted summation of the motion rate parameter values and motion direction change parameter values from 5 consecutive time steps within each time window. The weight coefficients are determined by the convolution kernel parameters. The output of the convolution operation is a 32-channel feature sequence, with each channel corresponding to a feature pattern extracted by one convolution kernel. The sequence length is equal to the original time series length minus the convolution kernel size plus 1. The temporal encoding representation is processed by a non-linear activation function, using the ReLU function, which maps negative values to 0 while keeping positive values unchanged. The temporal encoding represents a dimension of 32 times the time series length, extracting the change patterns of motion rate parameters and motion direction change parameters in the time dimension, including acceleration and deceleration patterns, periodic fluctuation patterns, and abrupt change patterns.
[0037] The body perception dimension feature vector is calculated based on the temporal coding representation, which is aggregated along the time dimension through a global average pooling operation. Global average pooling calculates the average value of the feature sequence for each channel, compressing the feature sequence of length equal to the length of the time series into a single value; 32 channels produce 32 average values. The body perception dimension feature vector has a dimension of 64; the first 32 elements are the average pooling results for the motion rate parameter channel, and the last 32 elements are the average pooling results for the motion direction change parameter channel. The body perception dimension feature vector represents the perceptual correlation strength between the visitor's body motion state and the spatial environment; elements with larger values in the vector correspond to feature dimensions with stronger interaction between motion patterns and the spatial environment. When the motion rate parameter exhibits periodic fluctuations in the time series, the average pooling value of the corresponding channel reflects the fluctuation amplitude; the larger the fluctuation amplitude, the stronger the perceptual correlation between the visitor and the spatial environment.
[0038] The mutual information between the rate of change of motion speed parameters and the rate of change of motion direction parameters is used to generate the motion interaction dimension feature vector. The rate of change of motion speed parameters is obtained by dividing the difference between the motion speed parameters at adjacent time points by the time interval, which is 0.1 seconds. The length of the rate of change sequence is equal to the length of the motion speed parameter sequence minus 1. The mutual information is obtained by calculating the joint probability distribution and marginal probability distribution after discretizing the sequence of the rate of change of motion speed parameters and the sequence of the rate of change of motion direction parameters. Discretization maps continuous values to 10 discrete intervals. The discrete interval boundaries for the rate of change of motion speed parameters are uniformly divided from -2.0 m / s to +2.0 m / s, and the discrete interval boundaries for the rate of change of motion direction parameters are uniformly divided from 0 degrees to 180 degrees. The joint probability distribution is obtained by dividing the frequency of the combination of the rate of change of motion speed parameters and the rate of change of motion direction parameters falling in each discrete interval by the total number of samples. The marginal probability distribution is obtained by dividing the frequency of a single sequence falling in each discrete interval by the total number of samples. Mutual information is equal to the logarithm of the ratio of the product of the joint probability distribution and the marginal probability distribution, multiplied by the joint probability distribution, and summed over all discrete interval combinations. The mutual information value ranges from 0 to 3; a higher value indicates a higher degree of coordination between changes in motion rate and direction. The motion interaction dimension feature vector has a dimension of 64. It is generated by inputting the sequence of changes in motion rate parameters and the sequence of changes in motion direction parameters into a fully connected layer to generate 32-dimensional sub-vectors. These two sub-vectors are then weighted and fused using the normalized value of mutual information as weights, and then concatenated to form a 64-dimensional vector. The weight coefficients are equal to the mutual information divided by 3 for normalization, and the normalized weights range from 0 to 1.
[0039] The context embedding dimension feature vector is generated based on pause pattern parameters and spatial semantic information from the spatial structure data. The pause pattern parameters are calculated through density-based spatial clustering analysis of trajectory points, with a cluster radius of 1.5 meters, a minimum number of cluster points of 5, and a time window of 10 seconds. The pause pattern parameter value is calculated as: number of trajectory points within a cluster / (spatial area covered by the cluster × time span). Spatial semantic information includes the functional and cultural attributes of the space. Functional attributes include exhibition areas, rest areas, passageways, and interactive areas; cultural attributes include historical themes, art themes, and science and technology themes. For each pause location, the corresponding functional and cultural attributes are queried from the spatial structure data. Functional attributes are encoded as 4-dimensional one-hot vectors, and cultural attributes as 3-dimensional one-hot vectors. The context embedding dimension feature vector has a dimension of 64 and is obtained by concatenating the pause pattern parameter value, functional attribute encoded vector, and cultural attribute encoded vector and then inputting it into a fully connected layer for mapping. The fully connected layer has an input dimension of 8, including one pause pattern parameter value, four functional attribute encoding elements, and three cultural attribute encoding elements, and an output dimension of 64. The weight matrix of the fully connected layer has a dimension of 8 x 64, and the bias vector has a dimension of 64. The output feature vector is calculated through matrix multiplication and bias addition. When there are multiple pause positions in the trajectory, a context embedding dimension feature vector is generated for each pause position, and the average value is taken as the final feature vector.
[0040] Multidimensional cognitive features are formed by concatenating feature vectors from the body perception dimension, motion interaction dimension, and context embedding dimension. The body perception dimension feature vector has a dimension of 64, the motion interaction dimension feature vector has a dimension of 64, and the context embedding dimension feature vector has a dimension of 64, resulting in a concatenated multidimensional cognitive feature dimension of 192. The concatenation operation arranges the elements of each dimension sequentially according to the feature vector order, forming a one-dimensional vector of length 192. The first 64 elements of the multidimensional cognitive feature represent the body perception dimension, the middle 64 elements represent the motion interaction dimension, and the last 64 elements represent the context embedding dimension. These multidimensional cognitive features, as a representation of embodied cognitive states, are input into the subsequent cognitive load decoding module for processing.
[0041] This invention transforms motion dynamics features into multidimensional cognitive feature representations, improving the accuracy and interpretability of embodied cognitive state decoding and providing a structured feature foundation for subsequent cognitive load assessment.
[0042] In one optional implementation, the step of dividing the target cultural and tourism space into multiple lighting response zones based on the spatial perception response pattern, and establishing a lighting demand descriptor associated with the cognitive load level for each lighting response zone to obtain the lighting demand mapping relationship includes: Based on the distribution of tourists' spatial focus in the spatial perception response mode, a perception activity index for the spatial area is extracted. The perception activity index represents the comprehensive level of tourists' perception interaction frequency and perception duration within a unit space. Based on the perceived activity index and the cognitive load level, the lighting response area is divided by a spatial clustering method. The spatial clustering method determines the area boundary based on the cognitive load similarity and spatial adjacency of adjacent spatial units. For each lighting response area, a lighting demand descriptor is constructed that includes a reference lighting parameter set and a cognitive load adjustment coefficient. The reference lighting parameter set is determined based on the average cognitive load level of the lighting response area, and the cognitive load adjustment coefficient is determined based on the standard deviation of the cognitive load level in the lighting response area. Establish a mapping relationship between lighting response areas and lighting demand descriptors to form a lighting demand mapping relationship.
[0043] For example, the spatial focus distribution is calculated by analyzing the distribution of attention weights for each spatial area within a statistical time window. For each spatial area, the average attention weight obtained at all time steps is calculated, and areas with an average weight exceeding 0.3 are marked as focus areas. The target cultural and tourism space is divided into spatial units using a 3-meter side square grid. Each spatial unit records its coordinates and functional attributes. The perceptual activity index includes two sub-indicators: perceptual interaction frequency and perceptual duration. Perceptual interaction frequency is obtained by counting the number of times the spatial unit is marked as a focus area per unit of time, with the unit time set to 1 minute and the frequency ranging from 0 to 20 times per minute. Perceptual duration is obtained by accumulating the length of time the spatial unit maintains a focus state. When a spatial unit maintains an average attention weight exceeding 0.3 for more than 3 consecutive seconds, it is defined as a sustained attention event. The duration is equal to the sum of the durations of all sustained attention events, with a value ranging from 0 to 120 seconds. The perceptual activity index = normalized value of perceptual interaction frequency × 0.6 + normalized value of perceptual duration × 0.4. Normalization maps the frequency to the range of 0 to 1 using the maximum value normalization method, and maps the duration to the range of 0 to 1 using the linear normalization method of dividing by 120 seconds. The perceptual activity index value ranges from 0 to 1. The larger the value, the higher the overall level of visitor perceptual interaction frequency and perceptual duration in that spatial unit.
[0044] The lighting response area is divided using a spatial clustering method based on perceptual activity indicators and cognitive load levels. The spatial clustering method employs a density-based clustering algorithm to cluster all spatial units. Clustering is based on two dimensions: cognitive load similarity and spatial adjacency. Cognitive load similarity is obtained by calculating the absolute value of the difference in cognitive load levels between adjacent spatial units. A difference less than 0.15 is defined as high similarity, a difference between 0.15 and 0.3 as medium similarity, and a difference greater than 0.3 as low similarity. Spatial adjacency is determined by whether two spatial units share a boundary. Spatial units sharing a boundary are defined as spatially adjacent, while those not sharing a boundary are defined as spatially non-adjacent. The core spatial unit criterion for the clustering algorithm is that the number of adjacent spatial units with high cognitive load similarity is not less than three. Spatial units meeting this condition are marked as core spatial units. The cluster expansion process starts from the core spatial unit and adds all its spatial units with high cognitive load similarity and spatial adjacency to the same cluster. This expansion operation is recursively executed until no new spatial units can be added. After processing all core spatial units, the clustering algorithm performs a merging operation on isolated spatial units that have not been assigned to any cluster. Isolated spatial units are assigned to the cluster with the highest similarity to their cognitive load and are spatially adjacent. If no spatially adjacent cluster exists, the isolated spatial unit forms its own cluster. Each cluster is defined as a lighting response region, and the region boundary is determined by the outer contours of all spatial units contained within the cluster. Contour extraction is performed by connecting the outer boundary segments of all boundary spatial units to form a closed polygon.
[0045] The lighting demand descriptor is constructed for each lighting response area and consists of two components: a reference lighting parameter set and a cognitive load adjustment coefficient. The reference lighting parameter set includes three parameters: luminance, color temperature, and color rendering index (CRI). Luminance is measured in lux, color temperature in Kelvin, and the CRI is dimensionless and ranges from 0 to 100. The reference lighting parameter set is determined based on the average cognitive load level of the lighting response area, which is obtained by averaging the cognitive load levels of all spatial units within that area. When the average cognitive load level is in the range of 0 to 0.3, the reference luminance is set to 300 lux, the reference color temperature to 4000 Kelvin, and the reference CRI to 80. When the average cognitive load level is in the range of 0.3 to 0.6, the reference luminance is set to 400 lux, the reference color temperature to 4500 Kelvin, and the reference CRI to 85. When the average cognitive load level is in the range of 0.6 to 1.0, the reference lighting luminance is set to 500 lux, the reference color temperature to 5000 Kelvin, and the reference color rendering index to 90. The cognitive load adjustment coefficient is determined based on the standard deviation of the cognitive load level within the lighting response area. The cognitive load adjustment coefficient = standard deviation × 2, with a range between 0 and 1. When the standard deviation exceeds 0.5, the adjustment coefficient is capped at 1. The cognitive load adjustment coefficient is used to control the response amplitude during subsequent dynamic adjustments of lighting parameters. A larger coefficient indicates higher spatial variability of the cognitive load level within the area, requiring a larger parameter variation range to accommodate local differences during lighting adjustments.
[0046] The lighting demand mapping is stored in a data table. Each row records a mapping item for a lighting response area. Fields include an area identifier, area boundary coordinate sequence, reference lighting luminance value, reference color temperature value, reference color rendering index value, and cognitive load adjustment coefficient value. The area identifier uses integer numbers, starting from 1 and incrementing. The area boundary coordinate sequence records the coordinates of all vertices of the area's outline polygon. The coordinate format is a pair of x and y coordinates in a two-dimensional Cartesian coordinate system, with a precision of 0.1 meters. The mapping supports dynamic query operations. Inputting any spatial coordinates determines the lighting response area by checking if the coordinates are inside a certain area boundary polygon. The query algorithm uses a ray-mapping method to determine the positional relationship between a point and the polygon. Rays are emitted from the query point in any direction, and the number of intersections between the ray and the polygon boundary is counted. An odd number of intersections indicates the point is inside the polygon, and an even number indicates the point is outside. The query result returns the corresponding lighting demand descriptor, containing the complete set of reference lighting parameters and the cognitive load adjustment coefficient for that area.
[0047] This invention enables automatic division of lighting response areas based on cognitive load, establishes a precise mapping relationship between areas and lighting parameters, and improves the spatial targeting and cognitive adaptation accuracy of lighting control.
[0048] In one optional implementation, the step of generating lighting parameter control instructions for each lighting response area based on the lighting demand mapping relationship and lighting constraints includes: Based on the spatial adjacency relationship of each lighting response region in the lighting demand mapping relationship, a spatial association graph between lighting response regions is constructed. The nodes of the spatial association graph represent lighting response regions, and the edge weights represent the cognitive load gradient between adjacent lighting response regions. The cognitive load gradient is obtained by calculating the ratio of the difference in the average cognitive load level of adjacent lighting response regions to the spatial distance. The lighting constraints include power constraints, color temperature adjustment range, and brightness adjustment range of the lighting equipment; For each lighting response region, initial target lighting parameters are calculated based on the reference lighting parameter set, and cross-regional coordinated adjustment is performed according to the edge weight distribution of the lighting response region in the spatial association graph. The cross-regional coordinated adjustment is achieved by minimizing the weighted sum of lighting parameter gradients between the lighting response region and its neighboring lighting response regions. The weight coefficient of the weighted sum of lighting parameter gradients is determined jointly based on the edge weights and the cognitive load adjustment coefficient of the lighting response region. The adjusted target lighting parameters are verified based on the lighting constraints, and lighting parameter control instructions are generated.
[0049] For example, the spatial association graph between lighting response areas is constructed based on the spatial adjacency relationships of each lighting response area in the lighting demand mapping relationship. Spatial adjacency is determined by whether there is a shared line segment at the boundary between two lighting response areas. Areas with a shared line segment longer than 1 meter are defined as spatially adjacent; those without a shared line segment or with a shared line segment shorter than 1 meter are defined as spatially non-adjacent. The spatial association graph is represented using an undirected graph data structure. The node set contains all lighting response areas, and each node stores an area identifier, an area boundary coordinate sequence, an average cognitive load level, and a cognitive load adjustment coefficient. The edge set contains all spatially adjacent pairs of lighting response areas, and each edge stores the identifiers of the two endpoint nodes and the edge weight. The edge weight represents the cognitive load gradient between adjacent lighting response areas, obtained by calculating the ratio of the difference in average cognitive load levels between adjacent lighting response areas to the spatial distance. The difference in average cognitive load levels is equal to the absolute value of the difference in average cognitive load levels between the two areas. The spatial distance is defined as the Euclidean distance between the centroids of the two areas, and the centroid coordinates are obtained by averaging the coordinates of all vertices of the region boundary polygon. Cognitive load gradient = average cognitive load level difference / spatial distance, in meters, with a value ranging from 0 to 0.5 per meter. A larger value indicates a more drastic spatial variation in cognitive load between adjacent areas. After the spatial association graph is constructed, it is stored in the form of an adjacency matrix. The matrix dimension is equal to the number of lighting response areas, and the matrix elements are edge weights. Elements with no edges have a value of 0.
[0050] Lighting constraints include three types: power constraints, color temperature adjustment range, and brightness adjustment range. Power constraints limit the maximum power consumption of a single lighting device, with an upper limit of 150 watts. When multiple lighting devices simultaneously serve a lighting response area, the total power consumption shall not exceed the area area × 5 watts per square meter. Color temperature adjustment range limits the range of color temperatures that lighting devices can output, with a lower limit of 2700 Kelvin and an upper limit of 6500 Kelvin, and a color temperature adjustment accuracy of 100 Kelvin. Brightness adjustment range limits the range of lighting brightness that lighting devices can output, with a lower limit of 100 lux and an upper limit of 1000 lux, and a brightness adjustment accuracy of 10 lux. Lighting constraints are stored in configuration files. Each type of constraint includes a constraint type identifier, constraint parameter name, and constraint parameter value. The configuration files support dynamic loading and updating at runtime, and changes to constraint parameters take effect immediately without requiring a system restart.
[0051] The initial target lighting parameters are calculated for each lighting response region based on a reference lighting parameter set. The reference lighting parameter set includes reference lighting luminance, reference color temperature, and reference color rendering index. The initial target lighting parameters directly copy the values from the reference lighting parameter set as initial values. The initial target lighting luminance is equal to the reference lighting luminance, the initial target color temperature is equal to the reference color temperature, and the initial target color rendering index is equal to the reference color rendering index. These initial target lighting parameters serve as the starting point for cross-regional coordinated adjustments, and subsequent optimization adjustments are made based on the edge weight distribution in the spatial correlation graph.
[0052] Cross-regional coordinated adjustment is achieved by minimizing the weighted sum of lighting parameter gradients between the current lighting response region and its neighboring lighting response regions. The lighting parameter gradients are calculated separately for both luminance and color temperature. The luminance gradient equals the absolute value of the difference between the target luminance of the current lighting response region and the target luminance of its neighboring regions; the color temperature gradient equals the absolute value of the difference between the target color temperature of the current lighting response region and the target color temperature of its neighboring regions. The weighted sum of lighting parameter gradients is obtained by traversing all neighboring regions of the current lighting response region, calculating the luminance gradient × corresponding edge weight + color temperature gradient × corresponding edge weight for each neighboring region, and summing the results for all neighboring regions. The weight coefficient is determined jointly based on the edge weight and the cognitive load adjustment coefficient of the lighting response region. Specifically, it is calculated as edge weight × the reciprocal of the cognitive load adjustment coefficient. When the cognitive load adjustment coefficient is less than 0.1, 0.1 is used as the denominator to avoid numerical overflow. Minimizing the weighted sum of lighting parameter gradients employs an iterative optimization algorithm. In each iteration, the target lighting parameters of the current lighting response region are fine-tuned. The fine-tuning amount is set to the current parameter value × 0.05, and the adjustment direction is to decrease the weighted sum of the lighting parameter gradients. The iteration terminates when the change in the weighted sum of the gradients of the lighting parameters is less than 0.01 or the number of iterations reaches 50. The target lighting parameters after the iteration is completed are used as the adjusted target lighting parameters.
[0053] The lighting constraint verification checks the compliance of the adjusted target lighting parameters. The brightness adjustment range verification determines whether the target lighting brightness is within the range of 100 lux to 1000 lux. If it exceeds the lower limit, the target lighting brightness is corrected to 100 lux; if it exceeds the upper limit, the target lighting brightness is corrected to 1000 lux. The color temperature adjustment range verification determines whether the target color temperature is within the range of 2700 Kelvin to 6500 Kelvin. If it exceeds the lower limit, the target color temperature is corrected to 2700 Kelvin; if it exceeds the upper limit, the target color temperature is corrected to 6500 Kelvin. The power constraint verification calculates the number of lighting devices required for the current lighting response area. The number of devices = area area / coverage area of a single lighting device. The coverage area is assumed to be 9 square meters, and the number of devices is rounded up. Total power consumption = number of devices × power consumption of a single device at the target lighting brightness. Power consumption is directly proportional to the lighting brightness, with a proportionality factor of 0.2 watts per lux. The total power consumption is checked to determine whether it exceeds the upper limit of area area × 5 watts per square meter. If it exceeds the upper limit, the power is limited by reducing the target lighting brightness. The reduction range is determined by the ratio of total power consumption to the power upper limit. The target lighting brightness is corrected to the original value / this ratio.
[0054] The lighting parameter control commands are generated based on the verified target lighting parameters. Each control command contains four fields: a lighting response area identifier, a target lighting brightness value, a target color temperature value, and a target color rendering index value. The control command format is a key-value pair structure, serialized using JSON encoding. Control commands are sent to the execution module of the lighting control system via a message queue. The message queue uses a first-in, first-out (FIFO) scheduling strategy. Each control command carries a timestamp and a priority identifier. The timestamp records the time the command was generated. The priority is determined based on the average cognitive load level of the lighting response area: areas with an average cognitive load level exceeding 0.7 have a high priority; areas with an average cognitive load level between 0.3 and 0.7 have a medium priority; and areas with an average cognitive load level below 0.3 have a low priority. After receiving the control command, the execution module of the lighting control system queries the list of lighting devices included in the lighting response area based on the lighting response area identifier. It then sends a parameter setting command to each lighting device, which includes the target lighting brightness, target color temperature, and target color rendering index. Upon receiving the command, the lighting device performs the parameter adjustment operation, with the adjustment time not exceeding 2 seconds.
[0055] This invention realizes cross-regional lighting parameter coordination optimization based on cognitive load gradient, and ensures the device executability of control commands through constraint condition verification, thereby improving the spatial continuity and system stability of lighting regulation.
[0056] In one optional implementation, the step of iteratively optimizing the embodied cognition model based on visitor behavior feedback data collected after lighting control includes: Collect visitor behavior feedback data after executing lighting parameter control commands, extract motion dynamics characteristics after lighting adjustment, and generate embodied cognitive states of response after lighting adjustment; Obtain the baseline embodied cognitive state before lighting adjustment, construct a comparison sequence of embodied cognitive states before and after lighting adjustment, calculate the state trajectory offset vector of the baseline embodied cognitive state and the response embodied cognitive state in the multidimensional cognitive load space, and generate cognitive load response features; A model prediction deviation quantification index is constructed based on the cognitive load response characteristics. The model prediction deviation quantification index is determined by calculating the directional consistency and magnitude deviation between the cognitive load change trend predicted by the embodied cognitive model and the actual cognitive load change trend in the cognitive load response characteristics. The update gradient of the attention weight parameters of the spatial-motor joint coding layer in the embodied cognition model is determined based on the directional consistency, and the update gradient of the decoding coefficients of the multi-dimensional cognitive load decoding module in the embodied cognition model is determined based on the amplitude deviation.
[0057] For example, visitor behavior feedback data after lighting adjustment is collected starting after the lighting parameter control command is executed, with the collection time window set to 30 to 120 seconds after the lighting adjustment is completed. The behavior feedback data includes a sequence of visitor trajectory points, a sequence of motion rate parameters, and a sequence of motion direction change parameters. The sampling frequency is 10 frames per second, and the data format is consistent with the behavior trajectory data before lighting adjustment. The extraction of motion dynamics features after lighting adjustment uses the same processing flow as before the adjustment. Temporal encoding representations are obtained by performing temporal convolution operations on the motion rate parameters and motion direction change parameters. The convolution kernel size is 5 time steps, the stride is 1 time step, and the number of convolution kernels is 32. The pause mode parameter is calculated through density-based spatial clustering analysis, with a cluster radius of 1.5 meters, a minimum cluster size of 5 points, and a time window of 10 seconds. The pause mode parameter value = number of trajectory points within the cluster / (spatial area covered by the cluster × time span). The motion dynamics features after lighting adjustment are input into the embodied cognitive model. After processing by the space-motion joint encoding layer and the multi-dimensional cognitive load decoding module, the embodied cognitive state of response after lighting adjustment is generated. The embodied cognitive state of response includes numerical values of three dimensions: visual cognitive load, navigational cognitive load, and emotional cognitive load, with each dimension ranging from 0 to 1.
[0058] The baseline embodied cognitive state is obtained from historical data prior to lighting adjustments, specifically generated from visitor behavior data collected 60 to 30 seconds before the execution of lighting parameter control commands. The baseline embodied cognitive state also includes values for three dimensions: visual cognitive load, navigational cognitive load, and emotional cognitive load. The comparison sequence of embodied cognitive states before and after lighting adjustments is formed by arranging the baseline and response embodied cognitive states in chronological order. The comparison sequence has a length of 2, with the first element being the baseline embodied cognitive state and the second element being the response embodied cognitive state. The state trajectory offset vector is calculated in a multidimensional cognitive load space, a 3-dimensional space where the three coordinate axes correspond to visual cognitive load, navigational cognitive load, and emotional cognitive load, respectively. The state trajectory offset vector = response embodied cognitive state vector - baseline embodied cognitive state vector. The three components of the vector represent the changes in visual cognitive load, navigational cognitive load, and emotional cognitive load, respectively. The component values range from -1 to +1, with positive values indicating an increase in cognitive load and negative values indicating a decrease in cognitive load. The cognitive load response characteristics consist of the magnitude and orientation angle of the state trajectory offset vector. The magnitude is the square root of (the square of the change in visual cognitive load + the square of the change in navigational cognitive load + the square of the change in emotional cognitive load), representing the overall magnitude of the cognitive load change. The orientation angle is determined by the direction cosines of the state trajectory offset vector in the multidimensional cognitive load space. The three direction cosines are each equal to the three components divided by the magnitude, and the orientation angle represents the trend direction of the cognitive load change.
[0059] The model prediction deviation quantification index is constructed based on cognitive load response characteristics. It is determined by calculating the directional consistency and magnitude deviation between the cognitive load change trend predicted by the embodied cognitive model and the actual cognitive load change trend in the cognitive load response characteristics. The cognitive load change trend predicted by the embodied cognitive model is calculated simultaneously when the lighting parameter control command is generated. The prediction is based on the reference lighting parameter set and cognitive load adjustment coefficient in the lighting demand descriptor. The expected cognitive load state after lighting adjustment is estimated through the lighting response prediction module. The difference between the expected cognitive load state and the reference embodied cognitive state forms the predicted state trajectory offset vector, which is also represented in the multidimensional cognitive load space. The directional consistency is obtained by calculating the cosine value of the angle between the predicted state trajectory offset vector and the actual state trajectory offset vector. The cosine value = (sum of the three components of the predicted vector multiplied by the three components of the actual vector) / (predicted vector magnitude × actual vector magnitude). The cosine value ranges from -1 to +1. The closer the value is to 1, the higher the directional consistency. The closer the value is to -1, the opposite the direction. Amplitude deviation = Actual state trajectory offset vector magnitude - Predicted state trajectory offset vector magnitude. The amplitude deviation ranges from -2 to +2. A positive value indicates that the actual change is greater than the predicted change, and a negative value indicates that the actual change is less than the predicted change. Model prediction deviation quantification index = (1 - Directional consistency) × 0.5 + Absolute value of amplitude deviation × 0.5. The index ranges from 0 to 2. The larger the value, the greater the model prediction deviation.
[0060] The update gradient of the attention weight parameters in the spatial-motion joint encoding layer is determined based on directional consistency. The attention weight parameters are the weight matrix used to calculate the dot product similarity between the motion dynamics feature vector and the spatial region feature vector in the spatial-motion joint encoding layer; the matrix dimension is 128×64. The update gradient is calculated using the backpropagation algorithm. The loss function for gradient calculation is the directional inconsistency loss, where the loss value = 1 - directional consistency, and the loss value ranges from 0 to 2. The gradient calculation process starts from the loss value and propagates forward layer by layer to the attention weight parameters. The gradient of each layer is equal to the gradient of the next layer multiplied by the derivative of the activation function of that layer. When the directional consistency is below 0.5, it is considered a directional inconsistency scenario, and the update gradient is multiplied by a factor of 2 to accelerate parameter adjustment. When the directional consistency is above 0.8, it is considered a highly directional consistency scenario, and the update gradient is multiplied by a factor of 0.5 to maintain parameter stability. The numerical range of the update gradient is limited to between -0.1 and +0.1 through gradient clipping to avoid gradient explosion leading to training instability.
[0061] The update gradient of the decoding coefficients in the multi-dimensional cognitive load decoding module is determined based on the amplitude deviation. The decoding coefficients are the weight matrices and bias vectors of the three parallel fully connected layer branches in the multi-dimensional cognitive load decoding module. The weight matrix dimensions for the visual cognitive load branch are 192×64, the navigation cognitive load branch's weight matrix dimensions are 192×64, and the emotion cognitive load branch's weight matrix dimensions are 192×64. The update gradient is calculated using the backpropagation algorithm. The loss function for gradient calculation is the amplitude deviation loss, where the loss value equals the square of the amplitude deviation, and the loss value ranges from 0 to 4. The gradient calculation process starts from the loss value and propagates to the decoding coefficients of the three branches. The gradient of each branch is related to the contribution of the change in the corresponding cognitive load dimension. When the absolute value of the amplitude deviation exceeds 0.3, it is considered a scenario with excessive amplitude deviation, and the update gradient is multiplied by a magnification factor of 1.5. When the absolute value of the amplitude deviation is less than 0.1, it is considered a scenario with small amplitude deviation, and the update gradient is multiplied by a reduction factor of 0.8. The range of updated gradient values is limited to between -0.05 and +0.05 by gradient clipping.
[0062] Parameter updates are performed using a stochastic gradient descent optimizer with a learning rate of 0.001 and a momentum coefficient of 0.9. The update formula for attention weights is: new parameter value = current parameter value - learning rate × update gradient; the update formula for decoding coefficients is: new parameter value = current parameter value - learning rate × update gradient. Parameter updates are performed in batches of 10 accumulated illumination adjustment feedback samples, with a batch size of 10. Updated parameters are saved as a new version, with the version number incremented. The current version parameters are backed up before each update. Model performance is evaluated by calculating the average of the model prediction bias quantification index on the validation set. When the average value increases after three consecutive updates, a version rollback is performed, restoring the parameters to the historical version with the lowest average value.
[0063] This invention quantifies the prediction deviation of the embodied cognition model by comparing the embodied cognitive state before and after lighting adjustment, and optimizes the attention weight parameters and decoding coefficients based on directional consistency and amplitude deviation, respectively, thereby realizing closed-loop iterative optimization of the embodied cognitive model and improving the model's prediction accuracy and adaptability to lighting response.
[0064] A second aspect of this invention provides a system for intelligent construction and optimization of cultural and tourism lighting scenes based on embodied cognition, comprising: The first unit is used to acquire spatial structure data and tourist behavior trajectory data of the target cultural and tourism space; The second unit is used to jointly analyze the spatial structure data and the tourist behavior trajectory data based on the embodied cognitive model, extract motion dynamics features from the tourist behavior trajectory data, calculate the tourist's cognitive load level in combination with the spatial structure data, and obtain cognitive feature representation results. The cognitive feature representation results include spatial perception response patterns and cognitive load level distribution. The third unit is used to divide the target cultural and tourism space into multiple lighting response areas according to the spatial perception response mode, and to establish a lighting demand descriptor associated with the cognitive load level for each lighting response area to obtain a lighting demand mapping relationship. The fourth unit is used to generate lighting parameter control instructions for each lighting response area based on the lighting demand mapping relationship and lighting constraints, control the lighting equipment based on the lighting parameter control instructions, and iteratively optimize the embodied cognition model based on the tourist behavior feedback data collected after lighting control.
[0065] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0066] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0067] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
Claims
1. A method for intelligent construction and optimization of cultural tourism lighting scenes based on embodied cognition, characterized in that, include: Acquire spatial structure data and tourist behavior trajectory data of the target cultural and tourism space; Based on the embodied cognition model, the spatial structure data and the tourist behavior trajectory data are jointly analyzed. Motion dynamics features are extracted from the tourist behavior trajectory data, and the cognitive load level of tourists is calculated by combining the spatial structure data to obtain cognitive feature representation results. The cognitive feature representation results include spatial perception response patterns and cognitive load level distribution. Based on the spatial perception response mode, the target cultural and tourism space is divided into multiple lighting response areas, and a lighting demand descriptor associated with the cognitive load level is established for each lighting response area to obtain a lighting demand mapping relationship. Based on the lighting demand mapping relationship and lighting constraints, lighting parameter control instructions are generated for each lighting response area. The lighting equipment is controlled based on the lighting parameter control instructions, and the embodied cognition model is iteratively optimized based on the tourist behavior feedback data collected after the lighting control.
2. The method according to claim 1, characterized in that, The steps for jointly analyzing the spatial structure data and the tourist behavior trajectory data based on the embodied cognitive model to obtain cognitive feature representation results include: Motion dynamics features are extracted from the tourist behavior trajectory data. These motion dynamics features include motion rate parameters, motion direction change parameters, and pause pattern parameters. The pause pattern parameters are obtained by calculating the spatiotemporal clustering density of trajectory points. A decoding mapping relationship is established from motion dynamics features to embodied cognitive state. The decoding mapping relationship maps the motion dynamics features into multi-dimensional cognitive features including body perception, motion interaction and context embedding. The motion interaction dimension characterizes the interaction intensity between the body and space by analyzing the cooperative pattern of the rate of change of motion speed parameter and the change of motion direction parameter. Based on the multidimensional cognitive features and the spatial complexity in the spatial structure data, the cognitive load level of tourists is calculated. The spatial complexity includes the functional zoning density and visual information density within a unit space. The cognitive load level is obtained by fusing motion complexity, dwell frequency, and spatial complexity. The motion complexity is determined based on the variance of the motion rate parameter and the motion direction change parameter.
3. The method according to claim 2, characterized in that, The embodied cognitive model includes: A spatial-motion joint coding layer is established through a bidirectional attention mechanism to create an interactive representation of the motion dynamics features and the spatial structure data. The attention weight of the motion dynamics features on the spatial region is used to identify the spatial focus of tourists, and the influence weight of spatial complexity on the motion dynamics features is used to predict environmentally induced motion changes. A multi-dimensional cognitive load decoding module decomposes the cognitive load level of tourists into visual cognitive load, navigation cognitive load and emotional cognitive load based on the multi-dimensional cognitive features. The visual cognitive load is calculated based on visual information density and gaze duration, the navigation cognitive load is calculated based on path complexity and turning frequency, and the emotional cognitive load is calculated based on pause patterns and movement smoothness. A lighting feedback-driven online learning module calculates an effectiveness score for the lighting adjustment based on visitor behavior feedback data collected after the lighting adjustment, and updates the parameters of the embodied cognition model based on the feedback.
4. The method according to claim 2, characterized in that, The steps to establish a decoding mapping relationship between motion dynamics characteristics and embodied cognitive states include: A temporal encoding representation of motion dynamics features is constructed. The temporal encoding representation is obtained by performing a temporal convolution operation on the motion dynamics features. The temporal convolution operation extracts the change patterns of motion rate parameters and motion direction change parameters in the time dimension. The body perception dimension feature vector is calculated based on the time-series coding representation, and the body perception dimension feature vector represents the strength of the perceptual association between the tourist's body movement state and the spatial environment. Calculate the mutual information between the rate of change of motion speed parameter and the rate of change of motion direction parameter, and generate a motion interaction dimension feature vector based on the mutual information. The mutual information represents the degree of coordination between the change of motion speed and the change of direction. Based on the pause mode parameters and the spatial semantic information in the spatial structure data, a context embedding dimension feature vector is generated, wherein the spatial semantic information includes the functional attributes and cultural attributes of the space. The body perception dimension feature vector, the motion interaction dimension feature vector, and the context embedding dimension feature vector are concatenated to form a multidimensional cognitive feature.
5. The method according to claim 1, characterized in that, Based on the spatial perception response pattern, the target cultural and tourism space is divided into multiple lighting response zones, and a lighting demand descriptor associated with the cognitive load level is established for each lighting response zone to obtain the lighting demand mapping relationship. The steps include: Based on the distribution of tourists' spatial focus in the spatial perception response mode, a perception activity index for the spatial area is extracted. The perception activity index represents the comprehensive level of tourists' perception interaction frequency and perception duration within a unit space. Based on the perceived activity index and the cognitive load level, the lighting response area is divided by a spatial clustering method. The spatial clustering method determines the area boundary based on the cognitive load similarity and spatial adjacency of adjacent spatial units. For each lighting response area, a lighting demand descriptor is constructed that includes a reference lighting parameter set and a cognitive load adjustment coefficient. The reference lighting parameter set is determined based on the average cognitive load level of the lighting response area, and the cognitive load adjustment coefficient is determined based on the standard deviation of the cognitive load level in the lighting response area. Establish a mapping relationship between lighting response areas and lighting demand descriptors to form a lighting demand mapping relationship.
6. The method according to claim 5, characterized in that, The steps for generating lighting parameter control instructions for each lighting response area based on the lighting demand mapping relationship and lighting constraints include: Based on the spatial adjacency relationship of each lighting response region in the lighting demand mapping relationship, a spatial association graph between lighting response regions is constructed. The nodes of the spatial association graph represent lighting response regions, and the edge weights represent the cognitive load gradient between adjacent lighting response regions. The cognitive load gradient is obtained by calculating the ratio of the difference in the average cognitive load level of adjacent lighting response regions to the spatial distance. The lighting constraints include power constraints, color temperature adjustment range, and brightness adjustment range of the lighting equipment; For each lighting response region, initial target lighting parameters are calculated based on the reference lighting parameter set, and cross-regional coordinated adjustment is performed according to the edge weight distribution of the lighting response region in the spatial association graph. The cross-regional coordinated adjustment is achieved by minimizing the weighted sum of lighting parameter gradients between the lighting response region and its neighboring lighting response regions. The weight coefficient of the weighted sum of lighting parameter gradients is determined jointly based on the edge weights and the cognitive load adjustment coefficient of the lighting response region. The adjusted target lighting parameters are verified based on the lighting constraints, and lighting parameter control instructions are generated.
7. The method according to claim 1, characterized in that, The steps for iteratively optimizing the embodied cognition model based on visitor behavior feedback data collected after lighting control include: Collect visitor behavior feedback data after executing lighting parameter control commands, extract motion dynamics characteristics after lighting adjustment, and generate embodied cognitive states of response after lighting adjustment; Obtain the baseline embodied cognitive state before lighting adjustment, construct a comparison sequence of embodied cognitive states before and after lighting adjustment, calculate the state trajectory offset vector of the baseline embodied cognitive state and the response embodied cognitive state in the multidimensional cognitive load space, and generate cognitive load response features; A model prediction deviation quantification index is constructed based on the cognitive load response characteristics. The model prediction deviation quantification index is determined by calculating the directional consistency and magnitude deviation between the cognitive load change trend predicted by the embodied cognitive model and the actual cognitive load change trend in the cognitive load response characteristics. The update gradient of the attention weight parameters of the spatial-motor joint coding layer in the embodied cognition model is determined based on the directional consistency, and the update gradient of the decoding coefficients of the multi-dimensional cognitive load decoding module in the embodied cognition model is determined based on the amplitude deviation.
8. A system for intelligent construction and optimization of cultural and tourism lighting scenes based on embodied cognition, used to implement the method of any one of claims 1-7, characterized in that, include: The first unit is used to acquire spatial structure data and tourist behavior trajectory data of the target cultural and tourism space; The second unit is used to jointly analyze the spatial structure data and the tourist behavior trajectory data based on the embodied cognitive model, extract motion dynamics features from the tourist behavior trajectory data, calculate the tourist's cognitive load level in combination with the spatial structure data, and obtain cognitive feature representation results. The cognitive feature representation results include spatial perception response patterns and cognitive load level distribution. The third unit is used to divide the target cultural and tourism space into multiple lighting response areas according to the spatial perception response mode, and to establish a lighting demand descriptor associated with the cognitive load level for each lighting response area to obtain a lighting demand mapping relationship. The fourth unit is used to generate lighting parameter control instructions for each lighting response area based on the lighting demand mapping relationship and lighting constraints, control the lighting equipment based on the lighting parameter control instructions, and iteratively optimize the embodied cognition model based on the tourist behavior feedback data collected after lighting control.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.