Virtual hotel guide customization system based on intelligent optimization and multi-modal interaction
By using gridded scene modeling and multimodal interaction, combined with spatiotemporal graph convolutional networks and deep reinforcement learning, personalized congestion-avoidance guide routes are generated, solving the problems of route dynamism and monotonous interaction in existing virtual hotel guide systems, thereby improving user experience and operational efficiency.
Patent Information
- Application Number
- CN202511084148.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-14
AI Technical Summary
Existing virtual hotel navigation systems cannot generate optimal routes based on user preferences and real-time visitor flow dynamics. Their interaction methods are limited, leading to missed or repeated visits to key areas. They also lack immersion and fail to meet the diverse and efficient user experience needs.
By employing gridded scene modeling combined with spatiotemporal graph convolutional congestion prediction and near-end strategy optimization reinforcement learning, personalized congestion avoidance guide routes are generated through multimodal interaction. This integrates user preference data, real-time pedestrian flow data, and multimodal interactive input to achieve path optimization and real-time updates.
It enables dynamic avoidance of congested areas within seconds and adjusts the itinerary according to user interests and time preferences, improving the efficiency and immersive experience of guided tours. It also supports multimodal interaction with voice, gestures, touch, and text, enhancing customer satisfaction and hotel operational efficiency.
Smart Images

Figure CN120953553A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual reality and intelligent tour guide technology, specifically a virtual hotel tour guide customization system based on intelligent optimization and multimodal interaction. Background Technology
[0002] In the hotel industry, with the rapid development of virtual reality, augmented reality, and artificial intelligence technologies, many hotels have introduced virtual tour systems to showcase guest rooms, public areas, and facilities to customers. Currently, mainstream virtual hotel tours typically use pre-recorded panoramic videos or static 3D models, relying on mouse clicks or touchscreen swipes for operation; some systems also incorporate scripted voice narration to create an immersive experience for users. These systems, in project implementation, primarily use fixed routes, preset scenes, and single interactive methods, achieving a basic virtual representation of the hotel space and improving online room selection and marketing efficiency to some extent.
[0003] However, existing technologies still have several shortcomings: the system cannot dynamically generate optimal routes based on user preferences, real-time pedestrian flow, and point-of-interest distribution, often leading to missed or repeated visits to key areas and overall low efficiency; the interaction methods are simplistic, relying solely on mouse clicks, touchscreen swipes, or preset voice prompts, lacking multimodal integration of voice, gestures, and text, failing to provide a natural and flexible interactive experience and lacking immersion. These deficiencies mean that existing virtual hotel guides still cannot meet the diverse and efficient needs in terms of user experience and commercial applications. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a virtual hotel tour customization system based on intelligent optimization and multimodal interaction. It solves the problem of generating personalized congestion-avoidance tour routes in real time under the drive of multimodal interaction by combining gridded scene modeling with spatiotemporal graph convolutional congestion prediction and near-end strategy optimization reinforcement learning, thereby achieving immersive virtual hotel tour customization.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a virtual hotel tour customization system based on intelligent optimization and multimodal interaction, comprising: The user preference collection module is used to acquire user preference data; The scene perception module is used to output the short-term congestion weights of each point of interest based on the hotel's interior floor plan and real-time pedestrian flow data through a spatiotemporal graph convolutional network congestion prediction model. The route optimization module is used to input the user preference data and the short-term congestion weight into a deep reinforcement learning model to generate a tour route that meets the time budget and has the least congestion, and update it in real time. A multimodal interaction module is used to receive voice, gesture, touch and text commands, and send the commands to the path optimization module; The system integration module is used to exchange navigation data with hotel reservation systems and customer relationship management systems through standard interfaces.
[0006] Preferably, the user preference data includes basic user information, historical browsing records, room type click sequence, theme requirements, and time budget.
[0007] Preferably, the scene perception module performs grid-based modeling of the indoor floor plan with a grid precision of 1m×1m, and collects pedestrian density data at time intervals not exceeding 2s.
[0008] Preferably, the spatiotemporal graph convolutional network congestion prediction model outputs the short-term congestion weights through the following steps: 4.1. Convert the hotel's interior floor plan into a graphical structure. ,vertex Representing points of interest, edges Indicates a passable passage; 4.2. Collect real-time pedestrian flow matrix And mapped to the graph structure; 4.3. Execution using a spatiotemporal graph convolutional network: in, For prediction Passenger flow density at each point of interest at any given time For adjacency matrices with self-loops, This is the original adjacency matrix. It is the identity matrix. for The degree matrix, These are the spatial convolution weights and the temporal convolution weights. This is the ReLU activation function.
[0009] Preferably, the deep reinforcement learning model generates the tour route through a proximal policy optimization network, and the formula for the proximal policy optimization network model is: Where t is the time step, The state vector includes the current position, remaining time budget, and congestion weight. In the state Next point of interest selected below Actions for the New and Old Strategies The probability of; probability ratio The dominance function measures the action. The relative merits of the baseline, Let be the distance between the two policy distributions. This is the penalty coefficient.
[0010] Preferably, the system latency for end-to-end command processing by the multimodal interaction module does not exceed 150ms, and the multimodal interaction module automatically switches to the default voice navigation mode if it does not detect a valid command for 10 consecutive seconds.
[0011] Preferably, the system integration module interacts with the hotel booking system through a JSON message body conforming to the REST specification, and uses TLS 1.3 encryption to ensure data security.
[0012] Preferably, the system further includes a historical trajectory storage unit, which is used to record the user's previous tour routes and generate a user profile.
[0013] This invention provides a virtual hotel tour customization system based on intelligent optimization and multimodal interaction. It has the following beneficial effects: This virtual hotel tour guide customization system, based on intelligent optimization and multimodal interaction, deeply couples real-time crowd prediction with personalized path generation through a collaborative architecture of "spatiotemporal graph convolutional network + near-end policy optimization deep reinforcement learning." It can not only dynamically avoid crowded areas within seconds, ensuring a smooth and redundant tour, but also adjust the itinerary in real time according to the user's remaining time and interests, achieving on-demand and precise service for different customer groups. At the same time, gridded scene modeling and 1m×1m precision data acquisition significantly improve the spatial resolution of path planning, enabling the system to maintain high efficiency and high interpretability even in large and complex hotel spaces.
[0014] At the interaction and operation level, multimodal input of voice, gesture, touch and text significantly reduces the operation threshold, and end-to-end response within 150ms ensures an immersive experience; REST interface plus TLS1.3 transmission securely connects tour data to the hotel booking platform, and combined with historical trajectory profiles, it can not only continuously optimize the cold start performance of the model, but also provide hotels with quantifiable passenger flow analysis and precise marketing support, thereby improving overall customer satisfaction and hotel operation efficiency. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the process of realizing the invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1
[0017] like Figure 1 As shown in the figure, this embodiment of the invention provides a virtual hotel tour customization system based on intelligent optimization and multimodal interaction, including a user preference collection module for acquiring user preference data. The user preference data includes basic user information, historical browsing records, room type click sequences, theme requirements, and time budget.
[0018] The scene perception module is used to output the short-term congestion weights of each point of interest based on the hotel's interior floor plan and real-time pedestrian flow data through a spatiotemporal graph convolutional network congestion prediction model.
[0019] The scene perception module performs grid-based modeling of the indoor floor plan with a grid precision of 1m×1m, and collects pedestrian density data at time intervals of no more than 2 seconds.
[0020] The specific implementation method is as follows: A seaside resort hotel with a floor plan of 120 meters long and 80 meters wide was selected as the test scenario. The hotel's floor plan was first gridded at a 1-meter by 1-meter granularity, resulting in 9600 grid cells. The system deployed 54 overhead depth cameras and 112 Bluetooth beacons for pedestrian flow monitoring, with the cameras spaced equidistantly along the corridor centerline to ensure that every grid cell was within the coverage of at least one sensor. The pedestrian flow data refresh cycle was 2 seconds, and the cameras continuously reported the current pedestrian flow count to the central server and automatically performed timestamp calibration.
[0021] User preference collection module: After a user initiates a virtual tour, the system simultaneously collects a questionnaire and tracks user behavior. Questionnaire fields include language / region, membership level, keywords, and desired tour duration. Behavior tracking records the user's browsing time, click order, and scrolling pace on the webpage and in the model. Room numbers 203 and 502, which appear most frequently in the click sequence, are assigned the highest weight. Combining the questionnaire and tracking results, the system generates a 128-bit preference vector and writes it to the real-time session cache for use in path optimization.
[0022] Scene perception module: The central server maps the refreshed pedestrian flow matrix to a graph structure, with vertices corresponding to 352 points of interest (POIs) and edges representing guest room corridors and public passageways. The server triggers a prediction process after collecting 30 frames of pedestrian flow data. A spatiotemporal graph convolutional network extrapolates the congestion level of each POI 5 minutes later, normalizing the output value to 0 to 1. In the actual prediction, restaurant node number 145 has a congestion level of 0.42 at the current moment, with a predicted value of 0.81, while landscape window number 17 maintains a congestion level of 0.12. The server writes the prediction results into the POI attributes, forming short-term congestion weights, and pushes them to the path optimization module within 2 seconds. This accuracy allows the system to update congestion status at a rate of seconds in large, complex spaces, providing a reliable basis for subsequent personalized congestion avoidance.
[0023] The spatiotemporal graph convolutional network congestion prediction model outputs short-term congestion weights through the following steps: 4.1. Convert the hotel's interior floor plan into a graphical structure. ,vertex Representing points of interest, edges Indicates a passable passage.
[0024] 4.2. Collect real-time pedestrian flow matrix And map it to a graph structure.
[0025] 4.3. Execution using a spatiotemporal graph convolutional network: in, For prediction Passenger flow density at each point of interest at any given time For adjacency matrices with self-loops, This is the original adjacency matrix. It is the identity matrix. for The degree matrix, These are the spatial convolution weights and the temporal convolution weights. This is the ReLU activation function.
[0026] The specific implementation method is as follows: Convert interior floor plan to diagram structure: This embodiment selects a seaside resort hotel measuring 120 meters long and 80 meters wide as the test scenario. First, a 2D floor plan accurate to centimeters is exported using Building Information Modeling (BIM), then meshed at a 1m x 1m granularity, resulting in a total of 9600 grid cells. The system marks four types of points of interest (POIs) on the floor plan: guest room entrances, scenic windows, restaurant counters, and fitness center entrances, totaling 352 points. Then, edges are established between POIs based on accessible pathways such as public corridors, evacuation staircases, and connecting bridges, resulting in 1174 valid edges. Walking distance, corridor width, and slope coefficient are added to each edge to form a weighted adjacency matrix; finally, self-loops are added to each vertex to ensure information integrity during subsequent convolution calculations. The final sparse graph structure is stored in a Redis graph database, with an average query latency of 4 milliseconds.
[0027] Real-time pedestrian flow matrix acquisition and mapping: Fifty-four overhead depth cameras are deployed on the hotel ceiling, and 112 Bluetooth cameras are installed in public areas; both work together to count crowd density. The cameras upload the latest number of people within their frames via RTSP streams every 2 seconds, using RSSI thresholds to estimate the number of currently connected devices. The data gateway timestamps and merges these two types of information into an 80x120-column crowd flow matrix, using cubic spline interpolation for missing grid cells. The system projects the density value of each cell in the matrix to the nearest point of interest; if a point of interest falls at the boundary between two grid cells with significantly different density gradients, the eight-neighbor average is used to smooth the noise. After 30 consecutive frames of data are written to a circular buffer, short-term congestion prediction is triggered.
[0028] Spatiotemporal graph convolutional networks predict short-term crowding weights: The prediction model employs a spatiotemporal graph convolutional network consisting of three spatial convolutional layers and two temporal convolutional layers. The number of channels in the spatial convolutional layers is 64, 64, and 32, respectively; the temporal convolutions slide along the time axis for 5 frames after each layer, with a stride of 1 frame. The ReLU activation function is used, and the weights are initialized using a uniform He distribution. The model takes the fused density sequence of the past 30 frames as input and outputs the predicted crowd density values for each point of interest in the next 300 seconds. These values are then linearly compressed and mapped to the 0-1 interval as short-term congestion weights.
[0029] During the training phase, peak operating data from the hotel over the past 30 days, 12 hours per day, was used as the sample. The mean squared error loss function was employed, with a learning rate of 0.0005 and a batch size of 8. The training converged after 32 epochs, and the mean absolute error on the validation set was 0.038 people / square meter. During the inference phase, the single-frame prediction latency was 43 milliseconds.
[0030] In a typical peak dinner test, the current pedestrian density at restaurant node 145 was 0.42 people / square meter, and the model predicted it would rise to 0.81 after 5 minutes; the current density at view window node 17 was 0.10, and the predicted value remained at 0.12; the current density at gym entrance node 321 was 0.35, and the predicted value decreased to 0.22. The system writes these three sets of weights into the point of interest attribute table and makes them available for the path optimization module to use in the next planning cycle, so as to achieve dynamic avoidance of highly crowded areas and adaptive adjustment of the tour order.
[0031] The route optimization module is used to input user preference data and short-term congestion weights into a deep reinforcement learning model to generate a guided route that meets the time budget and minimizes congestion, and updates it in real time.
[0032] The deep reinforcement learning model generates the tour route through a proximal policy optimization network. The formula for the proximal policy optimization network model is as follows: Where t is the time step, The state vector includes the current position, remaining time budget, and congestion weight. In the state Next point of interest selected below Actions for the New and Old Strategies The probability of; probability ratio The dominance function measures the action. The relative merits of the baseline, Let be the distance between the two policy distributions. This is the penalty coefficient.
[0033] The multimodal interaction module receives voice, gesture, touch, and text commands and sends them to the route optimization module. The system latency for end-to-end command processing by the multimodal interaction module does not exceed 150ms. If the multimodal interaction module does not detect a valid command for 10 consecutive seconds, it automatically switches to the default voice navigation mode.
[0034] The specific implementation method is as follows: Overall environment and constraints: The hotel is located in the city center business district, with 28 floors above ground. This guided tour takes place between the 5th-floor atrium and the 8th-floor executive floors. 18:10-19:00 is the hotel's generally recognized peak dinner time, with real-time monitoring showing an average population density of 0.58 people / ㎡ in public areas, peaking at 0.92 in the buffet area, and a fluctuation cycle of approximately 150 seconds. The user is a corporate travel platinum member, wearing the hotel's exclusive MR glasses, using Mandarin for voice interaction, and expects to confirm and complete the reservation for a sea-view room within 12 minutes.
[0035] Detailed process: State vector initialization: Remaining time budget: 720s.
[0036] Crowding weight: The system reads the predicted values of the most recent 30 frames, with the dining aisle at 0.85, the pool bar at 0.73, and the ocean view window at 0.18.
[0037] User preference vector: The theme is "sea view, fast", and the click sequence weights for room types are concentrated at 203 and 502.
[0038] In summary, the above concatenation results in a 482-dimensional state vector input PPO network.
[0039] First action decision (t=0): The PPO network outputs an action sequence, and the system immediately generates a route length of 176m, an estimated walking time of 684s, and a congestion penalty score of 0.22. A voice announcement reads: "A 12-minute fast route to the sea view has been planned for you. Would you like to begin?" The user nods in confirmation.
[0040] Real-time interaction and dynamic recalculation: Upon reaching P017, the user raises their hand and waves it twice quickly. The interaction module recognizes "skip the current point" within 92ms, and the remaining time budget is updated to 487s.
[0041] The path optimization module uses the new state vector to reason forward again, removing the 20-second pause at P017 and shortening the pause at P203 by 15 seconds, reducing the total journey to 592 seconds.
[0042] The system announced, "Your journey has been accelerated; estimated time remaining is 6 minutes." The user expressed satisfaction.
[0043] Automatic downgrade when no further instructions are given: If no new commands are detected for the next 10 seconds, the system triggers a fallback logic and switches to the default voice navigation: "Next stop, business room, follow me." At this point, the interaction thread enters a low-power listening mode, leaving only the voice keywords "pause" and "end" to activate the channel.
[0044] End of guided tour and learning replay: The actual end time was 18:20:03, with a total duration of 603 seconds, which was 117 seconds shorter than the budget.
[0045] The system automatically pops up a single-line rating bar. The user says "Great, five stars," and the satisfaction rating is recorded as 4.8.
[0046] The server writes the dwell sequence, trip duration, and satisfaction level into the replay buffer. The weight update strategy detects a satisfaction level higher than 3, which is considered positive feedback, and the data is directly entered into the PPO network offline learning queue.
[0047] Simultaneously add the "Business Quick Viewing" tag to the CRM to facilitate timely push of price discounts by sales staff.
[0048] Evaluation results: The average walking speed was 1.32 m / s, which is higher than the hotel average of 0.98 m / s.
[0049] The proportion of repeated road sections is only 4%, compared to 18% for traditional static guides.
[0050] With one interaction trigger, the average recognition latency is 101ms.
[0051] Users can ultimately complete room bookings directly in the MR view, reducing conversion time from the traditional 15 minutes to 2 minutes.
[0052] The system integration module is used to exchange navigation data with the hotel reservation system and customer relationship management system through standard interfaces. The system integration module interacts with the hotel reservation system using JSON message bodies conforming to REST specifications and employs TLS 1.3 encryption to ensure data security.
[0053] The system further includes a historical trajectory storage unit, which is used to record the user's previous tour routes and generate user profiles. Example 2
[0054] Unlike Example 1, the application scenario of this example is a parent-child leisure scenario.
[0055] Overall environment and constraints: At 10:30 AM on Sunday, the hotel occupancy rate was 72%, and the overall population density was relatively low, averaging 0.21 people / ㎡. The users were a family of 2 adults and 1 child who used the hotel's tablet to start an AR tour. Their goal was to experience the swimming pool and children's playground and select a family-themed room. Their time budget was 25 minutes, and the language used was a mix of English and Chinese.
[0056] Detailed process: State vector initialization: Remaining time budget: 1500s.
[0057] Crowding weights: pool bar 0.26, children's playground 0.19, buffet area 0.31, overall average 0.21.
[0058] User preference vector: Historical room type preferences are concentrated in 333 and 321.
[0059] Generate a 482-dimensional integrated state vector input PP0.
[0060] First action decision (t=0): The PPO output route is predicted to have a total duration of 1410 seconds, with a congestion penalty of 0.17, which is within budget.
[0061] Multiple interactive adjustments: At the poolside, a child taps a cartoon icon on a tablet to say "Play for another ten minutes." The text instruction takes 118ms to parse, extending the time budget to 2100 seconds. The path optimization module immediately adds P333, bringing the total planned time to 1870 seconds.
[0062] While the child is in the children's playground, the father's voice recognition took 104ms. The path optimization module removed P145 and P017, and added the shortest return segment, with an estimated remaining time of 280s.
[0063] Since no new instructions were detected for 10 consecutive seconds, the system automatically announced: "We will take you directly back to your room. If you need to add a stop, please say 'add' at any time." Safety and comfort control: To prevent children from getting lost, the system activates the "Parent-Child Synchronization" mode, comparing the relative distances of the three members in their field of vision every 5 seconds. If the distance exceeds 4 meters, a voice prompt will say "Please pay attention to your child".
[0064] When the system detects a pool floor water level of 0.5mm, it will avoid slippery areas during path planning and select the dry passage on the right to proceed to the changing room.
[0065] End of guided tour and data writing: The return journey was completed in 11:02:12, with a total walking distance of 312m and a total stay of 1702s.
[0066] Parents clicked the 5-star satisfaction rating and left a message saying "My child is very happy".
[0067] The system count shows that this interaction was triggered 5 times, with an average interaction latency of 109ms and a maximum latency of 132ms.
[0068] Write the "high parent-child engagement" tag into the CRM, trigger the marketing module to generate a parent-child afternoon tea coupon and push it to the user's mobile phone.
[0069] Evaluation results: Highly crowded areas accounted for 12% of visits, compared to 35% for traditional random guided tours during the same period.
[0070] The completion rate was 100%, and the timeout was only 1 minute, so the user did not notice the timeout.
[0071] Parent satisfaction rate: 4.9; children's feedback rating: 5 / 5 (based on smiley faces).
[0072] The system performed ST-GCN inference 42 times and PP0 recalculation 7 times throughout the entire guided tour, with resource utilization of 33% for GPU and 11% for CPU single core, fully meeting real-time requirements.
[0073] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A virtual hotel tour customization system based on intelligent optimization and multimodal interaction, characterized in that, include: The user preference collection module is used to acquire user preference data; The scene perception module is used to output the short-term congestion weights of each point of interest based on the hotel's interior floor plan and real-time pedestrian flow data through a spatiotemporal graph convolutional network congestion prediction model. The route optimization module is used to input the user preference data and the short-term congestion weight into a deep reinforcement learning model to generate a tour route that meets the time budget and has the least congestion, and update it in real time. A multimodal interaction module is used to receive voice, gesture, touch and text commands, and send the commands to the path optimization module; The system integration module is used to exchange navigation data with hotel reservation systems and customer relationship management systems through standard interfaces.
2. The virtual hotel tour customization system based on intelligent optimization and multimodal interaction according to claim 1, characterized in that: The user preference data includes basic user information, browsing history, room type click sequence, theme requirements, and time budget.
3. The virtual hotel tour customization system based on intelligent optimization and multimodal interaction according to claim 1, characterized in that: The scene perception module performs grid-based modeling of the indoor floor plan with a grid precision of 1m×1m, and collects pedestrian density data at time intervals of no more than 2s.
4. The virtual hotel tour customization system based on intelligent optimization and multimodal interaction according to claim 1, characterized in that: The spatiotemporal graph convolutional network congestion prediction model outputs the short-term congestion weights through the following steps: 4.
1. Convert the hotel's interior floor plan into a graphical structure. ,vertex Representing points of interest, edges Indicates a passable passage; 4.
2. Collect real-time pedestrian flow matrix And mapped to the graph structure; 4.
3. Execution using a spatiotemporal graph convolutional network: in, For prediction Passenger flow density at each point of interest at any given time For adjacency matrices with self-loops, This is the original adjacency matrix. It is the identity matrix. for The degree matrix, These are the spatial convolution weights and the temporal convolution weights. This is the ReLU activation function.
5. A virtual hotel tour customization system based on intelligent optimization and multimodal interaction as described in claim 1, characterized in that: The deep reinforcement learning model generates the tour route through a proximal policy optimization network, and the formula for the proximal policy optimization network model is: Where t is the time step, The state vector includes the current position, remaining time budget, and congestion weight. In the state Next point of interest selected below Actions for the New and Old Strategies The probability of; probability ratio For the dominant function, Let be the distance between the two policy distributions. This is the penalty coefficient.
6. The virtual hotel tour customization system based on intelligent optimization and multimodal interaction according to claim 1, characterized in that: The system latency for end-to-end command processing by the multimodal interaction module does not exceed 150ms. If the multimodal interaction module does not detect a valid command for 10 consecutive seconds, it automatically switches to the default voice navigation mode.
7. A virtual hotel tour customization system based on intelligent optimization and multimodal interaction as described in claim 1, characterized in that: The system integration module interacts with the hotel booking system through JSON message bodies conforming to the REST specification and using TLS 1.3 encryption.
8. The virtual hotel tour customization system based on intelligent optimization and multimodal interaction according to claim 1, characterized in that: The system further includes a historical trajectory storage unit, which is used to record the user's previous tour routes and generate a user profile.
Citation Information
Cited By
Real scene three-dimensional modeling method and management system for karst landform scenic area
CN121616765A
Real scene three-dimensional modeling method and management system for karst landform scenic spots
CN121616765B