Intelligent public space interactive design method and system based on space perception
By fusing multi-source sensor data and using deep learning models, the system can perceive the public space environment and user behavior in real time, generate adaptive interaction strategies, solve the problem of insufficient personalization and intelligence in the interactive experience in traditional design, and improve the utilization efficiency and user satisfaction of public spaces.
Patent Information
- Application Number
- CN202511303528.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Traditional public space interaction design lacks the ability to perceive and dynamically respond to the spatial environment and user behavior in real time. It cannot adapt its interaction design according to the number of people, their location, their behavior, and changes in the environment. As a result, the interaction experience is not personalized or intelligent enough, making it difficult to meet the diverse needs of users and failing to fully utilize the functionality and efficiency of the space.
Data is collected in real time by multiple sensors, fused using the LIO-SAM algorithm, and a three-dimensional dynamic spatial model is constructed using the MaskR-CNN model for semantic segmentation. User behavior is predicted based on the Social-STGCNN-Transformer network, generating intelligent interaction strategies, which are then executed through interactive devices in public spaces. The strategies are evaluated and optimized using a multi-objective Bayesian optimization algorithm.
It enables real-time perception and dynamic interactive adjustment of public spaces, enhances personalized interactive experience and space utilization efficiency, optimizes facility usage, alleviates congestion, and improves the intelligence and energy efficiency of public spaces.
Smart Images

Figure CN120803279A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent interaction design, and particularly relates to an intelligent public space interaction design method and system based on space perception. BACKGROUND
[0002] With the acceleration of urbanization and the increasing demand for public space experience, intelligent public space interaction design is attracting more and more attention. Traditional public space interaction design often lacks real-time perception and dynamic response capability for space environment and user behavior. Existing interaction design is mostly based on preset rules and fixed modes, and cannot adaptively adjust the interaction design according to the number, position, behavior state of personnel in the space and the changes of space environment such as light, temperature, air quality and other factors. This limitation leads to insufficient individualization and intelligence of the interaction experience of public space, and it is difficult to meet the increasingly diverse needs of users, and also cannot fully play the functionality and use efficiency of public space. SUMMARY
[0003] The present application relates to the technical field of intelligent interaction design, and particularly relates to an intelligent public space interaction design method and system based on space perception.
[0004] The first aspect of the present application provides an intelligent public space interaction design method based on space perception, which comprises the following steps: Real-time data acquisition in the public space through multi-source sensors, and fusion using LIO-SAM algorithm to obtain multi-source fusion data; Semantic segmentation of multi-source fusion data using MaskR-CNN model, and construction of a three-dimensional dynamic space model; Construction of a behavior prediction model based on Social-STGCNN-Transformer network, prediction of future behavior trends of personnel based on multi-source fusion data, and obtaining of user behavior analysis results; Generating intelligent interaction strategies according to the three-dimensional dynamic space model and the user behavior analysis results, and executing the intelligent interaction strategies through interaction devices in the public space; Obtaining feedback data, and evaluating and optimizing the intelligent interaction strategies based on the feedback data using a multi-objective Bayesian optimization algorithm.
[0005] Optionally, in the first implementation manner of the first aspect of the present application, the real-time data acquisition in the public space through multi-source sensors, and the fusion using LIO-SAM algorithm to obtain multi-source fusion data, comprises: Real-time data acquisition in the public space through multi-source sensors, and time stamp alignment preprocessing of the collected data; The feature points extracted by the laser radar are matched with a map to obtain constraint information of the laser radar pose, and a laser radar factor is formed; According to the pose information calculated from the IMU data and the relationship between the poses at adjacent moments, an IMU factor is constructed. In each iteration, the errors of the laser radar factor and the IMU factor are recalculated according to the current pose estimation value, and the pose estimation value of the node in the factor graph is continuously adjusted to minimize the laser point cloud registration error and the IMU inertial measurement error, thereby obtaining the multi-source fusion data.
[0006] Optionally, in the second implementation manner of the first aspect of the present application, the LIO-SAM algorithm is based on a factor graph optimization framework, wherein the factor graph is composed of nodes and edges, the nodes represent variables containing laser radar poses and IMU poses, and the edges represent constraint relationships between the variables.
[0007] Optionally, in the third implementation manner of the first aspect of the present application, the multi-source fusion data is subjected to semantic segmentation by using a MaskR-CNN model, and a three-dimensional dynamic space model is constructed, including: The multi-source fusion data is input into the MaskR-CNN model, feature extraction is performed on the multi-source fusion data by using a convolutional neural network, and a region proposal network is used to generate candidate regions that may contain targets; Feature alignment is performed on the candidate regions by using an ROIAlign layer to perform classification, bounding box regression and mask prediction on each target, and a semantic segmentation result is obtained; Based on the semantic segmentation result, the three-dimensional coordinate information of the laser point cloud data is combined, and a three-dimensional dynamic space model is constructed by using three-dimensional modeling technology.
[0008] Optionally, in the fourth implementation manner of the first aspect of the present application, a behavior prediction model is constructed based on a Social-STGCNN-Transformer network, a future behavior trend of a person is predicted based on the multi-source fusion data, and a user behavior analysis result is obtained, including: The multi-source fusion data is preprocessed, motion features of the person and relative position relationships between the person and others are extracted, and a spatio-temporal graph is constructed; The spatio-temporal graph is input into the behavior prediction model, a spatio-temporal graph convolutional layer extracts features of the spatio-temporal graph, and local features of the behavior of the person in the spatial and temporal dimensions are captured; A Transformer layer models long-distance dependency relationships between different persons by using a multi-head attention mechanism, and learns social interaction patterns between the persons; The behavior prediction model is used to predict a behavior trend of the person within 10-30 seconds in the future, and a user behavior analysis result is output, wherein the behavior trend includes a moving direction, a staying probability and a possibility of interaction with others.
[0009] Optionally, in the fifth implementation form of the first aspect of the present application, the generating the intelligent interaction strategy according to the three-dimensional dynamic space model and the user behavior analysis result, and executing the intelligent interaction strategy through the interaction device in the public space comprises: setting different interaction objects in the public space as a plurality of agents, and constructing a state space for each agent based on the three-dimensional dynamic space model and the user behavior analysis result; constructing a strategy network based on the multi-layer perception mechanism, the strategy network taking the state space of the agent as input and outputting a probability distribution of each possible action in the current state; the agent selecting an action to execute by using a greedy selection according to the action probability distribution output by the strategy network, the joint action being constituted when all the agents select and execute the action, and the public space generating a corresponding reward signal; optimizing the strategy network of the agent by using a proximal policy optimization based on the obtained reward signal and state information, and integrating the action selection strategy output by the strategy network of each agent to obtain the intelligent interaction strategy.
[0010] Optionally, in the sixth implementation form of the first aspect of the present application, the obtaining the feedback data and evaluating and optimizing the intelligent interaction strategy by using a multi-objective Bayesian optimization algorithm based on the feedback data comprises: obtaining the feedback data through the interaction device, constructing a probability model of the objective function, and gradually finding the optimal strategy combination by continuously sampling and evaluating a new intelligent interaction strategy, so that a plurality of optimization objectives are balanced, wherein the optimization objectives at least include user satisfaction, space use efficiency and interaction response speed.
[0011] The second aspect of the present application provides an intelligent public space interaction design system based on space perception, which comprises: a collection module configured to collect data in the public space through a plurality of source sensors in real time, and fuse the data by using a LIO-SAM algorithm to obtain multi-source fusion data; a construction module configured to perform semantic segmentation on the multi-source fusion data by using a MaskR-CNN model, and construct a three-dimensional dynamic space model; a prediction module configured to construct a behavior prediction model based on a Social-STGCNN-Transformer network, predict a future behavior trend of a person based on the multi-source fusion data, and obtain a user behavior analysis result; an execution module configured to generate an intelligent interaction strategy according to the three-dimensional dynamic space model and the user behavior analysis result, and execute the intelligent interaction strategy through an interaction device in the public space; An optimization module is configured to acquire feedback data, and evaluate and optimize the intelligent interaction strategy based on the feedback data using a multi-objective Bayesian optimization algorithm.
[0012] A third aspect of the present application provides a spatial perception-based intelligent public space interaction design device, comprising a memory and at least one processor, wherein the memory stores instructions; and the at least one processor invokes the instructions in the memory to enable the spatial perception-based intelligent public space interaction design device to perform the steps of the spatial perception-based intelligent public space interaction design method according to any one of the preceding aspects.
[0013] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and the instructions are executed by a processor to implement the steps of the spatial perception-based intelligent public space interaction design method according to any one of the preceding aspects.
[0014] In the technical solution provided by the present application, data is collected in real time by a multi-source sensor in a public space, and is fused using an LIO-SAM algorithm to obtain multi-source fusion data; the multi-source fusion data is subjected to semantic segmentation using a MaskR-CNN model, and a three-dimensional dynamic space model is constructed; a behavior prediction model is constructed based on a Social-STGCNN-Transformer network, the future behavior trend of a person is predicted based on the multi-source fusion data, and a user behavior analysis result is obtained; an intelligent interaction strategy is generated based on the three-dimensional dynamic space model and the user behavior analysis result, and the intelligent interaction strategy is executed by an interaction device in the public space; feedback data is acquired, and the intelligent interaction strategy is evaluated and optimized based on the feedback data using a multi-objective Bayesian optimization algorithm; the present application can perceive the environmental changes and user behaviors in a public space in real time, dynamically generate an interaction strategy, and realize adaptive adjustment of the interaction design of the public space, thereby greatly improving the personalized interaction experience of a user in the public space; the use efficiency of the space is optimized: through analysis and prediction of user behaviors and guidance of the intelligent interaction strategy, the distribution of people flow can be reasonably adjusted, the use of facilities in the public space can be optimized, the overall use efficiency of the public space can be improved, and problems such as space congestion can be alleviated, thereby realizing dynamic interaction optimization of the public space and improving the intelligent, personalized and energy-saving level of the public space. BRIEF DESCRIPTION OF DRAWINGS
[0015] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments with reference made to the accompanying drawings. The drawings are for purposes of illustration only and are not considered as limiting the application.
[0016] Figure 1 A flowchart of the intelligent public space interaction design method based on spatial perception provided by the embodiment of the present application is shown in FIG. 1. Figure 2 A structural schematic diagram of the intelligent public space interaction design system based on spatial perception provided by the embodiment of the present application is shown in FIG. 2. Figure 3 A structural schematic diagram of the intelligent public space interaction design device based on spatial perception provided by the embodiment of the present application is shown in FIG. 3. DETAILED DESCRIPTION
[0017] The terms "first", "second", "third", "fourth" and the like in the description, claims, and drawings of the present application (if any) are used for distinguishing between similar objects, not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of these terms herein is to be construed as interchangeable in order to distinguish between the similar objects. It is also to be understood that the terms "comprising", "having", "including" and "containing" are to be construed as open-ended terms (i.e., meaning "including, but not limited to") unless otherwise noted. It is intended that the application encompass the following embodiments and any and all equivalents thereof.
[0018] For the purpose of facilitating understanding, the specific flow of the embodiment of the present application is described below, and reference is made to FIG. 1. Figure 1 The flowchart of the intelligent public space interaction design method based on spatial perception provided by the embodiment of the present application includes the following steps: Step 101, real-time collection of data in the public space by a multi-source sensor and fusion by a LIO-SAM algorithm to obtain multi-source fusion data. In the embodiment, in the public space, a laser radar, a visual camera, an inertial measurement unit (IMU), a microphone array, an environmental sensor and the like are deployed as multi-source sensors. The sensors collect data in real time at different frequencies, wherein the laser radar collects space point cloud data at a frequency of 10Hz-20Hz, the visual camera collects image data at a frame rate of 30fps, and the IMU obtains motion attitude data at a frequency of 100Hz.
[0019] In this embodiment, data is collected in real time by multi-source sensors in a public space, and the collected data is preprocessed by timestamp alignment; the feature points extracted by the laser radar are matched with the map to obtain the constraint information of the laser radar pose and form the laser radar factor; the pose information calculated according to the IMU data and the relationship between adjacent poses are used to construct the IMU factor; in each iteration, the error of the laser radar factor and the IMU factor is recalculated according to the current pose estimation value, the pose estimation value of the node in the factor graph is continuously adjusted, the laser point cloud registration error and the IMU inertial measurement error are minimized, and the multi-source fusion data is obtained.
[0020] In this embodiment, the point cloud data collected by the laser radar will undergo preprocessing, and the algorithm will remove abnormal points in the point cloud data due to factors such as occlusion and noise. At the same time, according to the measurement principle of the laser radar, the data is converted in coordinates and unified to the global coordinate system of the public space. Then, the algorithm extracts features from the processed point cloud data, mainly extracting plane feature points and edge feature points. Plane feature points come from relatively flat areas in the public space, such as the ground and the wall surface; edge feature points come from the boundaries and corners of objects, which can effectively represent the geometric structure information of the space. For the data collected by the inertial measurement unit, the algorithm first performs filtering processing to eliminate noise caused by sensor errors and external interference. IMU data contains the measurement values of the accelerometer and the gyroscope, which reflect the acceleration and angular velocity of the object respectively. Through integration operation, the changes of acceleration and angular velocity with time are converted into position, velocity and attitude information of the object. However, due to the drift problem of IMU, the calculated position information will have a large error over time, so it needs to be fused and corrected with laser radar data later. When fusing laser radar and IMU data, the LIO-SAM algorithm works based on the factor graph optimization framework. The factor graph is composed of nodes and edges, where the nodes represent the pose of the laser radar, the pose of the IMU and other variables, and the edges represent the constraint relationship between the variables. The algorithm matches the feature points extracted by the laser radar with the constructed map to obtain the constraint information of the laser radar pose and form the laser radar factor. At the same time, the pose information calculated according to the IMU data and the relationship between adjacent poses are used to construct the IMU factor. The LIO-SAM algorithm continuously adjusts the pose estimation value of the node in the factor graph through iterative optimization, and the goal is to minimize the laser radar point cloud registration error and the IMU inertial measurement error. In each iteration, the error of the laser radar factor and the IMU factor is recalculated according to the current pose estimation value, and then the pose is adjusted by using an optimization algorithm to continuously reduce the overall error. After multiple iterations, a high-precision laser radar and IMU fused pose estimation is obtained, and then high-precision multi-source fusion data is obtained, so as to realize accurate perception of the public space environment and personnel position and motion information.
[0021] Step 102, performing semantic segmentation on the multi-source fusion data by using a MaskR-CNN model, and constructing a three-dimensional dynamic space model; In this embodiment, the multi-source fusion data is input into the MaskR-CNN model, the convolutional neural network is used to extract features of the multi-source fusion data, and the region proposal network is used to generate candidate regions that may contain targets; the ROIAlign layer is used to align the features of the candidate regions, so as to classify, regress the boundary box and predict the mask of each target, and obtain the semantic segmentation result; based on the semantic segmentation result, the three-dimensional coordinate information of the laser radar point cloud data is combined, and the three-dimensional modeling technology is used to construct a three-dimensional dynamic space model.
[0022] In this embodiment, after the multi-source fusion data is input into the MaskR-CNN model, the convolutional neural network (CNN) begins to play a key role. The CNN is composed of multiple convolutional layers, pooling layers and activation function layers. It will gradually extract features of the multi-source fusion data according to the hierarchical structure. First, the data enters the first convolutional layer, and the convolution kernel slides on the data to extract basic features such as edges and textures through convolution operation. With the continuous transmission of data in the network, the subsequent convolutional layers will further extract more abstract and representative high-level semantic features based on the basic features extracted in the previous step. For example, in a public space scene, the features of chairs, tables, pedestrians and other objects are gradually recognized from the initial recognition of lines and small areas. In this process, the pooling layer reduces the data dimension through downsampling operation, reduces the calculation amount, and at the same time preserves important features. The activation function gives the network nonlinear ability, so that the network can learn complex feature relationships. Finally, a feature map containing rich spatial semantic information is obtained. A region proposal network (RPN) is used to generate candidate regions that may contain targets based on the feature maps extracted by the CNN. The RPN slides a small network over the feature map, generates a plurality of anchor boxes of different scales and aspect ratios centered at each sliding position, and these anchor boxes are like preset search boxes in space, covering various possible target shapes and sizes. Then, the RPN evaluates each anchor box to determine whether it contains a target and predicts the offset between the anchor box and the real target. By setting a suitable score threshold, the anchor boxes with higher scores are selected as candidate regions. For example, in a public space, the RPN can quickly find regions where pedestrians and facilities may exist. After obtaining the candidate regions, the ROIAlign layer is used to align the features of the candidate regions to realize classification, bounding box regression and mask prediction for each target, thereby obtaining the semantic segmentation result. The main function of the ROIAlign layer is to accurately align and sample the features of the candidate regions to a fixed size. It first determines the corresponding feature region of the candidate region according to its position in the original feature map, then divides the region into a plurality of small units, and accurately extracts the features in each small unit by bilinear interpolation and other methods, thereby avoiding the feature deviation problem caused by quantization operation in the traditional ROIpooling. After feature alignment, the obtained features are input into the classification branch, the bounding box regression branch and the mask prediction branch. The classification branch determines which category the target belongs to, such as pedestrian, seat, trash can, etc. The bounding box regression branch further adjusts the position and size of the candidate region to make it more consistent with the real target. The mask prediction branch generates a pixel-level mask for each target to accurately outline the contour of the target, and finally realizes accurate semantic segmentation of various targets in the public space to clearly distinguish different objects and personnel. Based on the semantic segmentation result and the three-dimensional coordinate information of the laser radar point cloud data, a three-dimensional dynamic space model is constructed using three-dimensional modeling technology. The semantic segmentation result determines the category and two-dimensional contour of each target in the space, and the laser radar point cloud data provides accurate three-dimensional spatial position information. By combining the two, the corresponding three-dimensional coordinates are first assigned to each target after semantic segmentation to determine its actual position and shape in the space. Then, a suitable three-dimensional modeling algorithm, such as the triangular meshing algorithm, is used to convert the discrete point cloud data and semantic information into a continuous three-dimensional geometric model. During modeling, the dynamic characteristics of the public space are considered, and the model is updated in real time. For example, when personnel move or object positions change, the newly collected multi-source data will be reprocessed through the above process to timely correct and update the three-dimensional model, so that it always accurately reflects the real-time state of the public space.
[0023] Step 103, constructing a behavior prediction model based on the Social-STGCNN-Transformer network, predicting the future behavior trend of the personnel based on the multi-source fusion data, and obtaining the user behavior analysis result; In this embodiment, the multi-source fusion data is preprocessed, the motion features of the personnel and the relative position relationship between the personnel are extracted, and a space-time graph is constructed. The space-time graph is input into the behavior prediction model, the space-time graph convolution layer extracts the features of the space-time graph, and captures the local features of the personnel behavior in the spatial and temporal dimensions. The Transformer layer models the long-distance dependency relationship between different personnel through the multi-head attention mechanism, and learns the social interaction mode between the personnel. The behavior prediction model is used to predict the behavior trend of the personnel within 10-30 seconds in the future, and the user behavior analysis result is output, wherein the behavior trend includes the moving direction, the staying probability and the interaction possibility with others.
[0024] Step 104, generating an intelligent interaction strategy according to the three-dimensional dynamic space model and the user behavior analysis result, and executing the intelligent interaction strategy through the interaction device in the public space; In this embodiment, different interaction objects in the public space are set as multiple agents, a state space is constructed for each agent based on the three-dimensional dynamic space model and the user behavior analysis result. A strategy network is constructed based on the multilayer perception mechanism, the strategy network takes the state space of the agent as the input, and outputs the probability distribution of each possible action in the current state. The agent selects an action to execute according to the action probability distribution output by the strategy network, and when all the agents select and execute the action, a joint action is formed, and the public space generates a corresponding reward signal. Based on the obtained reward signal and state information, the strategy network of the agent is optimized by using the proximal policy optimization, and the action selection strategy output by the strategy network of each agent is integrated to obtain the intelligent interaction strategy.
[0025] In this embodiment, different interaction objects or interaction targets in the public space are set as multiple agents, for example, different area display screens, voice broadcasting systems, intelligent lighting devices, etc. are respectively regarded as independent agents, each agent has autonomous decision-making ability, a state space is constructed for each agent based on the three-dimensional dynamic space model and the user behavior analysis result. The state space includes the spatial position of the agent, the surrounding environment information, the user behavior characteristics and the state of the agent itself, etc. The surrounding environment information such as personnel density and nearby facility usage, the user behavior characteristics such as the behavior trend and demand preference of nearby users, and the state of the agent itself such as the current working mode and remaining resources of the device. These information are presented in the form of structured data, so that the agent can perceive the current environmental state; Each agent has a policy network, which is built based on a multi-layer perception (MLP), and the policy network takes the state space of the agent as input, and outputs the probability distribution of each possible action in the current state through the calculation of multiple neurons and the action of the activation function. The action space covers all possible interactive actions that the agent can take, such as switching the display screen to display content, playing specific prompts through the voice broadcast system, adjusting the brightness and color of intelligent lighting devices, etc. The agent selects an action to execute according to the action probability distribution output by the policy network, such as selecting the action with the highest probability or randomly selecting an action according to the probability distribution, to achieve interaction with the public space environment and the user. When all agents select and execute actions, these actions collectively constitute joint actions that act on the public space environment. The public space environment changes according to the joint actions, and generates corresponding feedback information, including the user's response to the interactive actions, changes in the space state, etc. The response includes whether the user pays attention to the display screen content, whether the user acts according to the voice prompt, whether the personnel distribution is more reasonable, whether the facility usage efficiency is improved, etc. These feedback information is quantified as a reward value, which is used to measure the execution effect of the joint action. Positive reward indicates that the action has a positive effect on achieving the goal, and negative reward indicates that the action has a negative impact. Each agent obtains its own reward signal and updated state information based on the environment feedback. Based on the obtained reward signal and state information, the MAPPO algorithm optimizes the policy network of the agent using the idea of proximal policy optimization. The core idea is to maximize the cumulative reward while ensuring that the policy update is not too drastic. The algorithm calculates the difference between the current policy and the previous policy, limits the magnitude of policy update to avoid performance degradation caused by drastic fluctuations in the optimization process. In the optimization process, an advantage function is used to evaluate the merits of each action. The advantage function combines the current action reward and future expected reward, which can more accurately reflect the value of the action. According to the advantage function and the action probability output by the policy network, the loss function is calculated, and the parameters of the policy network are updated through the backpropagation algorithm, so that the policy network gradually learns the policy that can obtain higher rewards. After multiple rounds of training and optimization, when the policy network converges to a certain extent, the action selection strategy output by the policy network of each agent is integrated to form an intelligent interaction strategy for the current public space state and user behavior, which is used to guide the actual operation of the interactive devices in the public space.
[0026] Step 105, obtain feedback data, and evaluate and optimize the intelligent interaction strategy based on the feedback data using a multi-objective Bayesian optimization algorithm.
[0027] In this embodiment, feedback data is obtained through feedback buttons on the interactive device, facial recognition of the camera, monitoring of user behavior changes, etc. The feedback data includes direct feedback of the user on the interaction strategy, such as whether to click confirm, whether to choose to ignore, etc., and indirect feedback, such as whether to change the action route after accepting the guidance, whether the facility use efficiency is improved, etc.
[0028] In this embodiment, feedback data is obtained through the interactive device, a probability model of the target function is constructed, new intelligent interaction strategies are continuously sampled and evaluated, and the optimal strategy combination is gradually found to balance multiple optimization objectives, including at least user satisfaction, space use efficiency, and interaction response speed.
[0029] Please refer to Figure 2 The structure of the intelligent public space interaction design system based on space perception provided by the embodiment of the application is shown in the figure, and the system includes: The acquisition module is configured to acquire data in real time in the public space through a plurality of source sensors, and fuse the data using an LIO-SAM algorithm to obtain multi-source fusion data. The construction module is configured to perform semantic segmentation on the multi-source fusion data using a MaskR-CNN model, and construct a three-dimensional dynamic space model. The prediction module is configured to construct a behavior prediction model based on a Social-STGCNN-Transformer network, predict future behavior trends of the personnel based on the multi-source fusion data, and obtain a user behavior analysis result. The execution module is configured to generate an intelligent interaction strategy based on the three-dimensional dynamic space model and the user behavior analysis result, and execute the intelligent interaction strategy through an interactive device in the public space. The optimization module is configured to obtain feedback data, and evaluate and optimize the intelligent interaction strategy based on the feedback data using a multi-objective Bayesian optimization algorithm.
[0030] Figure 3is a structural schematic diagram of an intelligent public space interaction design device based on space perception provided by an embodiment of the present application. The intelligent public space interaction design device based on space perception 300 can have great differences due to different configurations or performances, and can include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, one or more storage media 330 (for example, one or more mass storage devices) storing application programs 333 or data 332. The memory 320 and the storage medium 330 can be temporary storage or persistent storage. The programs stored in the storage medium 330 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the intelligent public space interaction design device based on space perception 300. Further, the processor 310 can be configured to communicate with the storage medium 330 and execute a series of instruction operations in the storage medium 330 on the intelligent public space interaction design device based on space perception 300 to implement the method provided by the above embodiment.
[0031] The intelligent public space interaction design device based on space perception 300 can also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating devices 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and the like. Those skilled in the art can understand that, Figure 3 The structure of the intelligent public space interaction design device based on space perception shown does not constitute a limitation on the computer device provided by the present application, and can include more or fewer components than shown, or combine certain components, or different component arrangements.
[0032] The present application also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium, or a volatile computer readable storage medium. The computer readable storage medium has instructions stored therein, and when the instructions are run on a computer, the computer executes the steps of the intelligent public space interaction design method based on space perception provided by the above embodiments.
[0033] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device or apparatus, unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0034] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0035] The basic principles, main features and advantages of the present application are shown and described above. Those skilled in the art should understand that the present application is not limited by the above embodiments, and the above embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. An intelligent public space interaction design method based on spatial perception, characterized by: The method comprises the following steps: In public spaces, data is collected in real time through multi-source sensors and fused using the LIO-SAM algorithm to obtain multi-source fused data. The Mask R-CNN model is used to perform semantic segmentation on multi-source fusion data and construct a three-dimensional dynamic space model; A behavior prediction model is built based on the Social-STGCNN-Transformer network. Based on multi-source fusion data, the future behavior trends of users are predicted to obtain user behavior analysis results. Generate intelligent interaction strategies based on the 3D dynamic space model and user behavior analysis results, and execute them through interactive devices in public spaces; Obtain feedback data and use a multi-objective Bayesian optimization algorithm based on the feedback data to evaluate and optimize the intelligent interaction strategy.
2. The method for designing intelligent public space interactions based on spatial perception according to claim 1, characterized in that: The multi-source fusion data is obtained by collecting data in real time through multi-source sensors in the public space and fusing them using the LIO-SAM algorithm, including: In public spaces, data is collected in real time through multi-source sensors, and the collected data is pre-processed for time stamp alignment; Match the feature points extracted by the lidar with the map to obtain the constraint information of the lidar pose and form the lidar factor; Construct the IMU factor based on the pose information calculated from the IMU data and the relationship between the poses at adjacent moments; In each iteration, the errors of the lidar factor and the IMU factor are recalculated according to the current pose estimation value, and the pose estimation value of the node in the factor graph is continuously adjusted to minimize the lidar point cloud registration error and the IMU inertial measurement error to obtain multi-source fusion data.
3. The method for designing intelligent public space interactions based on spatial perception according to claim 1, characterized in that: The LIO-SAM algorithm is based on a factor graph optimization framework, where the factor graph consists of nodes and edges. The nodes represent variables including the lidar pose and the IMU pose, and the edges represent the constraints between the variables.
4. The method for designing intelligent public space interactions based on spatial perception according to claim 1, wherein: The Mask R-CNN model is used to perform semantic segmentation on the multi-source fusion data and construct a three-dimensional dynamic space model, including: The multi-source fusion data is input into the Mask R-CNN model, and features are extracted from the multi-source fusion data through a convolutional neural network. The region proposal network is used to generate candidate regions that may contain the target. The candidate regions are aligned through the ROIAlign layer to classify, regress bounding boxes, and predict masks for each target to obtain semantic segmentation results. Based on the semantic segmentation results and combined with the three-dimensional coordinate information of the lidar point cloud data, a three-dimensional dynamic space model is constructed using three-dimensional modeling technology.
5. The method for designing intelligent public space interactions based on spatial perception according to claim 1, characterized in that: The behavior prediction model is constructed based on the Social-STGCNN-Transformer network, and the future behavior trends of people are predicted based on multi-source fusion data to obtain user behavior analysis results, including: Preprocess the multi-source fusion data to extract the movement characteristics of people and the relative position relationship between people, and construct a spatiotemporal graph; The spatiotemporal graph is input into the behavior prediction model, and the spatiotemporal graph convolution layer extracts features from the spatiotemporal graph to capture the local characteristics of human behavior in the spatial and temporal dimensions. The Transformer layer uses a multi-head attention mechanism to model long-distance dependencies between different people and learn the social interaction patterns between them; The behavior prediction model is used to predict the behavior trends of people in the next 10 to 30 seconds and output user behavior analysis results, where the behavior trends include movement direction, probability of staying, and possibility of interacting with others.
6. The method for designing intelligent public space interactions based on spatial perception according to claim 1, characterized in that: Generating an intelligent interaction strategy based on the three-dimensional dynamic space model and user behavior analysis results, and executing the intelligent interaction strategy through interactive devices in the public space, includes: Different interactive objects in the public space are set as multiple intelligent agents. Based on the three-dimensional dynamic space model and user behavior analysis results, a state space is constructed for each intelligent agent. A policy network is constructed based on a multi-layer perceptron. The policy network takes the state space of the agent as input and outputs the probability distribution of each possible action in the current state. The agent selects an action based on the action probability distribution output by the policy network using greedy selection. When all agents select and execute an action, a joint action is formed, and the public space generates a corresponding reward signal. Based on the obtained reward signal and state information, the proximal policy optimization is used to optimize the policy network of the intelligent agent. The action selection strategy output by the policy network of each intelligent agent is integrated to obtain the intelligent interaction strategy.
7. The method for designing intelligent public space interactions based on spatial perception according to claim 1, wherein: The obtaining of feedback data and the use of a multi-objective Bayesian optimization algorithm to evaluate and optimize the intelligent interaction strategy based on the feedback data include: Feedback data is obtained through interactive devices, and a probabilistic model of the objective function is constructed. By continuously sampling and evaluating new intelligent interaction strategies, the optimal strategy combination is gradually found to balance multiple optimization objectives, where the optimization objectives include at least user satisfaction, space utilization efficiency, and interaction response speed.
8. Intelligent public space interactive design system based on spatial perception, characterized by: The system includes: The acquisition module is used to collect data in real time through multi-source sensors in public spaces and fuse them using the LIO-SAM algorithm to obtain multi-source fused data; A construction module is used to perform semantic segmentation on multi-source fusion data using the Mask R-CNN model and construct a three-dimensional dynamic space model; The prediction module is used to build a behavior prediction model based on the Social-STGCNN-Transformer network, predict the future behavior trends of people based on multi-source fusion data, and obtain user behavior analysis results; An execution module is used to generate intelligent interaction strategies based on the three-dimensional dynamic space model and user behavior analysis results, and execute the intelligent interaction strategies through interactive devices in the public space; The optimization module is used to obtain feedback data and evaluate and optimize the intelligent interaction strategy based on the feedback data using a multi-objective Bayesian optimization algorithm.
9. An intelligent public space interactive design device based on spatial perception, characterized in that: The intelligent public space interaction design device based on spatial perception includes a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory to enable the intelligent public space interaction design device based on spatial perception to perform each step of the intelligent public space interaction design method based on spatial perception according to any one of claims 1 to 7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the intelligent public space interaction design method based on spatial perception are implemented.
Citation Information
Patent Citations
Private domain live broadcast peak hot spot prediction and content scheduling method based on deep learning
CN119450099A
Space situation response type generation method and system based on artificial intelligence
CN120105321A
Foreign advertisement putting system for predicting advertisement click rate
CN120181926A
Hospital intelligent space management system and method
CN120221008A
Multi-agent cooperation system and method based on spatial calculation and multi-modal AI fusion
CN120297354A