Method and system for automatic guide and client apparatus

TW202634561AActive Publication Date: 2026-08-16KING ONE INTERACTIVE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW114105631
Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2026-08-16
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Existing methods for browsing information, such as traditional photos and text, fail to provide immersive and efficient navigation in virtual spaces, limiting user engagement and information access.

Method used

An automated navigation system utilizing artificial intelligence (AI) to navigate three-dimensional spatial models through a user terminal device and cloud server, enabling real-time rendering and movement of virtual cameras based on user input, with AI-driven recommendations for panoramic views.

Benefits of technology

Facilitates rapid and intelligent navigation in virtual spaces, providing immersive experiences and efficient access to user-interested content without manual clicks, enhancing user engagement and information delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TA001072358_001
    Figure TWG2TA001072358_001
  • Figure TWG2TA001072358_002
    Figure TWG2TA001072358_002
  • Figure TWG2TA001072358_003
    Figure TWG2TA001072358_003
Patent Text Reader

Abstract

A method and a system for automatic guide and a client apparatus are provided. The method for automatic guide includes: after displaying a three-dimensional space model, obtaining response content corresponding to an input message; when the response content includes a description associated with a virtual object in the three-dimensional space model, displaying the response content and a recommended option corresponding to the virtual object; and when selecting the recommended option, automatically guiding a viewing angle of a user to a suitable angle.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a technique for creating virtual spaces, and more particularly to an automatic navigation method, system, and user terminal device for three-dimensional spatial models. [Previous Technology]

[0002] With the advancement of technology and the rise of the internet, more and more websites are available for people to access the information they need via internet connection. In an era that demands innovation and change, the traditional method of browsing photos and reading text is increasingly unable to meet the needs of the public. As a result, 3D guided tour technology has been developed, which creates a realistic environment to give users a sense of immersion. Automated guided tours are one of the directions currently under research. [Summary of the Invention]

[0003] The present invention provides an automatic navigation method, system and user terminal device, which combine artificial intelligence (AI) technology to enable users to obtain information of interest more quickly.

[0004] The automatic navigation method of the present invention is applicable to an automatic navigation system, the automatic navigation system including a user terminal device and a cloud server. The automatic navigation method includes: communicating with the cloud server through a user interface of the user terminal device and displaying a three-dimensional spatial model provided by the cloud server in the user interface, wherein the three-dimensional spatial model has spatial arrangement data; receiving input messages through an input device of the user terminal device and transmitting the input messages to the cloud server through the user interface; obtaining response content corresponding to the input messages from the spatial arrangement data through the cloud server; if the response content includes a description associated with virtual objects in the three-dimensional spatial model, outputting the response content and recommended options associated with the virtual objects to the user interface for display through the cloud server; if a recommended option is selected through the user interface, automatically sending control commands to the cloud server through the user interface, causing the cloud server to drive a virtual camera corresponding to the three-dimensional spatial model to move to a designated positioning point corresponding to the virtual object, and executing a real-time rendering program to place a designated panoramic view corresponding to the designated positioning point as a material into the three-dimensional spatial model; and displaying the three-dimensional spatial model with the designated panoramic view placed in it through the user interface.

[0005] In one embodiment of the present invention, the above-mentioned three-dimensional space model includes multiple positioning points, and the database of the cloud server includes multiple panoramic images corresponding to the multiple positioning points respectively, with a designated positioning point being one of the positioning points and a designated panoramic image being one of the panoramic images.

[0006] In one embodiment of the present invention, the above-mentioned automatic navigation method further includes: providing a space layout interface through a cloud server; and setting the arrangement of the virtual space corresponding to the three-dimensional space model through the space layout interface, thereby obtaining space layout data.

[0007] In one embodiment of the present invention, the above-mentioned spatial arrangement data includes multiple virtual objects corresponding to multiple exhibits, and each virtual object has a corresponding configuration position and positioning point.

[0008] In one embodiment of the present invention, the above-mentioned automatic navigation method further includes: responding to a communication connection between a user interface of a user terminal device and a cloud server, executing a real-time rendering program through the cloud server to place a preset panoramic image corresponding to a preset location as a material into a three-dimensional space model; and displaying the three-dimensional space model with the preset panoramic image placed in it through the user interface. The three-dimensional space model includes multiple positioning points, and the database of the cloud server includes multiple panoramic images corresponding to the multiple positioning points respectively, wherein the preset location is one of the positioning points, and the preset panoramic image is one of the panoramic images.

[0009] In one embodiment of the present invention, the step of obtaining the response content corresponding to the input message from the spatial arrangement data through the cloud server includes: querying the spatial arrangement data of the three-dimensional spatial model through the large language model provided by the cloud server to obtain the query result corresponding to the input message; and converting the query result into response content that conforms to the prompt word project through the large language model of the cloud server.

[0010] In one embodiment of the present invention, when the response content includes a description associated with a virtual object in a three-dimensional spatial model, the method includes: generating recommended options associated with the virtual object through a front-end program provided by a cloud server, and outputting the response content and recommended options to the user interface.

[0011] The automatic navigation system of the present invention includes: a cloud server and a user terminal device, wherein the user terminal device includes a first storage unit having a user interface, an input device, and a first processor, and the cloud server includes a second processor. In the user terminal device, the first processor is configured to: communicate with the cloud server through the user interface and display a three-dimensional spatial model provided by the cloud server in the user interface, wherein the three-dimensional spatial model has spatial arrangement data; and receive input messages through the input device and transmit the input messages to the cloud server through the user interface. In the cloud server, the second processor is configured to: obtain response content corresponding to the input messages from the spatial arrangement data, and determine whether the response content includes a description associated with virtual objects in the three-dimensional spatial model; and if the response content includes a description associated with virtual objects in the three-dimensional spatial model, output the response content and recommended options associated with the virtual objects to the user interface for display through the cloud server. In the user terminal device, the first processor is configured to: automatically send control commands to the cloud server through the user interface when a recommended option is selected through the user interface. In the cloud server, the second processor is configured to: in response to receiving a control command, move the virtual camera corresponding to the 3D spatial model to the specified positioning point corresponding to the virtual object, and execute a real-time rendering program to place the specified panoramic image corresponding to the specified positioning point as a material into the 3D spatial model; and display the 3D spatial model with the specified panoramic image placed into the user interface.

[0012] The automatic navigation method of the present invention utilizes a processor to perform the following steps: displaying a three-dimensional spatial model through a user interface, wherein the three-dimensional spatial model has spatial arrangement data; receiving input information through an input device; obtaining response content corresponding to the input information through the user interface, wherein the response content is generated based on the spatial arrangement data; if the response content includes a description associated with virtual objects in the three-dimensional spatial model, displaying the response content and recommended options corresponding to the virtual objects in the user interface; if a recommended option is selected through the user interface, driving a virtual camera corresponding to the three-dimensional spatial model to move to a designated positioning point corresponding to the virtual object through the user interface, and driving the execution of a real-time rendering program to place a designated panoramic view corresponding to the designated positioning point as a material into the three-dimensional spatial model; and displaying the three-dimensional spatial model with the designated panoramic view placed in it through the user interface.

[0013] The user terminal device of the present invention includes: a storage device including a user interface; an input device; and a processor coupled to the storage device and the input device and configured to perform the steps of the above-described automatic navigation method.

[0014] Based on the above, the present invention combines artificial intelligence to provide automatic navigation functions, which can proactively provide navigation based on the content that the user is interested in, and can automatically provide navigation without the user having to click, thereby creating a rapid deployment of virtual display and intelligent customer service and navigation functions.

Implementation Method

[0015] Figure 1 is a block diagram of an automated tour guide system according to an embodiment of the present invention. Referring to Figure 1, the automated tour guide system 10 includes a user terminal device 100A and a cloud server 100B. The user terminal device 100A is, for example, an electronic device used by a user, such as a smartphone, tablet computer, laptop computer, or desktop computer. The cloud server 100B is, for example, a computer with powerful computing capabilities or a large amount of disk storage space. The user terminal device 100A and the cloud server 100B can be connected using wired or wireless communication technology.

[0016] The user terminal device 100A includes a first processor 110, an input device 120, a first storage 130, and a first communication connector 140. The first processor 110 is coupled to the input device 120, the first storage 130, and the first communication connector 140. The first storage 130 stores at least one code segment, which is used by the first processor 110 to implement subsequent automatic navigation steps.

[0017] In the user terminal device 100A, the first storage 130 includes a user interface 131. In one embodiment, the user interface 131 is a browser. In another embodiment, the user interface 131 is a graphical interface provided by an application (APP).

[0018] The cloud server 100B includes a second processor 150, a second storage 160, and a second communication connector 170. The second storage 160 stores at least one code segment, which is executed by the second processor 150 to implement the subsequent automatic navigation steps.

[0019] In the cloud server 100B, the second storage 160 includes a three-dimensional spatial model 161, a database 162, a foreground program 163, and a background program 164. In one embodiment, the foreground program 163 and the background program 164 each include one or more code segments, which are executed by the second processor 150 after installation. The foreground program 163 is responsible for the content presented by the user interface 131 of the client device 100A. The background program 164 is responsible for the development, management, and / or editing of the three-dimensional spatial model 161 by developers (administrators).

[0020] Database 162 stores multiple panoramic images corresponding to multiple positioning points included in the 3D spatial model 161, as well as task scripts corresponding to the 3D spatial model 161. The panoramic images can be created, for example, by 3D animation software or by taking photos with a camera. The virtual camera of the 3D animation software can simulate a real-world camera, dividing the 3D scene into multiple angle photos, and then creating a panoramic image from these photos. The task script is a set of events designed for a specific task. For example, a specific task could be one of the following: website page navigation, automatic spatial walking, website visualization, line-of-sight guidance, voice explanation, explanation of the model from various angles, etc.

[0021] The automated navigation system 10 further includes a Large Language Model (LLM) 180. The Large Language Model 180 may be located on a different server than the cloud server 100B. Alternatively, the Large Language Model 180 may also be located on the cloud server 100B; this is not a limitation. The Large Language Model 180 is an artificial intelligence (AI) program built on machine learning, capable of recognizing and interpreting human language or other complex data, thereby generating text and performing other tasks. The Large Language Model 180 consists of an artificial neural network with many parameters, trained on large amounts of text data using self-supervised or semi-supervised learning. For example, the Large Language Model 180 may be ChatGPT, GPT4, LLaMA-65B, PaLM-62B, BERT, T5, or Wenxin Yiyan, etc.

[0022] The first processor 110 and the second processor 150 are, for example, a central processing unit (CPU), a graphics processing unit (GPU), or other programmable microprocessors, digital signal processors (DSPs), programmable controllers, application specific integrated circuits (ASICs), programmable logic devices (PLDs), or other similar devices.

[0023] The first storage 130 and the second storage 160 are, for example, any type of fixed or removable random access memory, read-only memory, flash memory, hard disk or other similar device or combination of these devices.

[0024] The first communication connector 140 and the second communication connector 170 may be chips or circuits employing Local Area Network (LAN) technology, Wireless LAN (WLAN) technology, or mobile communication technology. An example of a LAN is Ethernet. An example of a WLAN is Wi-Fi. Examples of mobile communication technologies include Global System for Mobile Communications (GSM), third-generation (3G), fourth-generation (4G), and fifth-generation (5G).

[0025] Input device 120 is, for example, a mouse, keyboard, touch screen device, voice sensor, etc. The voice sensor is, for example, a microphone array audio in module, used to collect ambient sound and generate corresponding voice signals.

[0026] FIG2 is a flowchart of an automatic navigation method according to an embodiment of the present invention. Referring to FIG1 and FIG2, in step S205, the user interface 131 of the user terminal device 100A communicates with the cloud server 100B, and displays the three-dimensional space model 161 provided by the cloud server 100B in the user interface 131. Here, the three-dimensional space model 161 has spatial arrangement data. For example, if the user interface 131 is a browser, the user can enter a specified URL in the browser of the user terminal device 100A to connect to the cloud server 100B, and the three-dimensional space model 161 provided by the cloud server 100B is displayed in the browser. If the user interface 131 is an application, the user can run the application in the user terminal device 100A, and the application automatically connects to the cloud server 100B based on its internal settings parameters, and the three-dimensional space model 161 provided by the cloud server 100B is displayed in the application interface.

[0027] In response to the communication connection between the user interface 131 of the user terminal device 100A and the cloud server 100B, the cloud server 100B executes a real-time rendering process to place a preset panoramic image corresponding to a preset position into the stereoscopic space model 161 as a material. Furthermore, the stereoscopic space model 161 with the preset panoramic image placed is displayed through the user interface 131. Generally, 3D models are covered with textures; the process of arranging textures on a 3D model is called texture mapping. After texture mapping, the 3D model can be made more detailed and look more realistic. In this embodiment, the panoramic image is placed into the stereoscopic space model 161 as the material for texture mapping. In addition to texture mapping, surface normals can be adjusted to achieve lighting effects, and some surfaces can also use bump mapping and other stereoscopic rendering techniques.

[0028] During the rendering process, the second processor 150 uses the foreground program 163 to locate the relationship between the virtual camera and each virtual object in the stereoscopic space model 161, thereby obtaining multiple rendering parameters. In the real world, the eye sees objects because light is reflected from them. The location of the virtual camera is the location of the "eye". Rendering parameters such as the reflection index between the virtual camera and each pixel in the stereoscopic space model 161 can be calculated using ray tracing. Therefore, after obtaining the specified location point (the destination of the virtual camera's movement), the second processor 150 will further move the virtual camera to the specified location point to obtain the rendering parameters corresponding to this specified location point.

[0029] Figure 3 is a schematic diagram of a three-dimensional spatial model including multiple positioning points according to an embodiment of the present invention. Referring to Figure 3, the three-dimensional spatial model 161 is generated by a 3D model designer using software such as 3D modeling tools according to customer requirements, corresponding to a virtual space. After generating the three-dimensional spatial model 161, the 3D model designer sets multiple positioning points P1 to P29 in the three-dimensional spatial model 161 according to customer requirements, and sets a corresponding panoramic view for each positioning point P1 to P29, and stores it in the database 162. In addition, at least one interactive position can be reserved in the three-dimensional spatial model 161 as needed, so that users can dynamically place interactive elements. Interactive elements are, for example, multimedia files, interactive interfaces, or buttons. In the embodiment shown in Figure 3, 19 interactive positions E1 to E19 are pre-set in the three-dimensional spatial model 161. However, this is only an example and does not limit the number of positioning points or the number of reserved interactive positions.

[0030] The following describes how a user terminal device 100A connects to a cloud server 100B to open a 3D webpage. After receiving the connection notification, the cloud server 100B will start the 3D spatial model 161. In the cloud server 100B, the second processor 150 executes tasks through the foreground program 163 according to the preset task script, interactive operations, etc.

[0031] After executing the stereoscopic space model 161, and before detecting any operation instructions applied to the stereoscopic space model 161, the second processor 150 will first move the lens position of the virtual camera to a preset position (e.g., positioning point P1, but not limited to this) through the foreground program 163, and execute a real-time rendering program to place the panoramic view (preset panoramic view) corresponding to the positioning point P1 into the stereoscopic space model 161 as a material, so that the display screen of the user interface 131 of the user terminal device 100A presents the stereoscopic space model 161 with the preset panoramic view placed in it.

[0032] Next, in step S210, input messages are received through the input device 120 of the user terminal device 100A, and the input messages are transmitted to the cloud server 100B through the user interface 131. Afterwards, in step S215, the response content corresponding to the input messages is obtained from the spatial arrangement data through the cloud server 100B.

[0033] For example, the input device 120 is a voice sensor. The input device 120 receives the user's voice and generates a voice signal (input message). The voice signal is transmitted to the cloud server 100B through the connection between the user terminal device 100A and the cloud server 100B via the user interface 130. In the cloud server 100B, the second processor 150 analyzes the voice signal through the foreground program 163 to obtain the input message, and obtains the response content corresponding to the input message from the spatial arrangement data.

[0034] In one embodiment, the spatial arrangement data can be generated, for example, through a background program 164. The background program 164 provides a spatial arrangement interface to set the arrangement of the virtual space corresponding to the three-dimensional spatial model 161, thereby obtaining the spatial arrangement data. The spatial arrangement data includes multiple virtual objects corresponding to multiple exhibits, each virtual object having a corresponding configuration position and positioning point.

[0035] For example, developers can use a browser (or an application developed for the cloud server 100B) on a user terminal device 100A or an electronic device similar to the user terminal device 100A to connect to the cloud server 100B, and then arrange the virtual space corresponding to the three-dimensional space model 161 through the background program 164. For example, the space arrangement interface provided by the background program 164 guides developers to set up images and videos for the virtual space (similar to the act of posting posters in a physical space), multiple virtual objects corresponding to multiple exhibits, and information for the display window (the content of the window that expands after clicking).

[0036] Figures 4A to 4D are schematic diagrams illustrating the operation of a space arrangement interface according to an embodiment of the present invention. First, in Figure 4A, the developer selects option 410 of the space arrangement interface 400, and is guided in block 412 to prepare data for setting images, videos, and display windows for the virtual space. For example, block 412 provides functional options such as "display materials," "virtual decorations," "theme materials," "display windows," and "external links" for the developer to use, and displays the selected results in block 414.

[0037] In Figure 4B, the developer can select option 420 of the space layout interface 400, and use the function options such as "Quick Create" or "Create a Freely Arranged Scene" in block 422 to create a new project (for space layout), or select an established project 424 in the list of established projects for management or browsing.

[0038] After selecting a project (a new project or an existing project 424), the layout management page 430, as shown in Figure 4C, is accessed. The layout management page 430 provides a list of tools 432 and 434 (including multiple editing tools) for developers to place previously prepared data in the virtual space. After the data is placed in the virtual space, the background program 164 binds the virtual objects (display items, functional showcases) placed in the virtual space and their space information (such as configuration positions) to the virtual space, thereby generating space layout data.

[0039] Furthermore, as the developer completes the space arrangement, the space arrangement interface 400 can further guide the developer to supplement more information about the exhibits and display rooms, as shown in Figure 4D. Accordingly, through the guidance of the space arrangement interface 400, developers can supplement sufficient relevant information while focusing on arranging the virtual space.

[0040] After the virtual space is set up, the background program 164 will automatically generate a knowledge base based on the space layout data of the virtual space. The purpose is to convert the space layout data into vector data and store it in the vector database so that the large language model 180 can search for data in the future.

[0041] The spatial arrangement interface 400 uses previously bound attributes (each arranged virtual object has attributes belonging to the virtual space) to separate the data of the virtual space. In other words, when the large language model 180 searches for data, it will only search for the spatial arrangement data set in the vector database that corresponds to the virtual space, which can avoid the data from different virtual spaces from contaminating each other and causing AI illusion (erroneous or misleading results generated by the AI ​​model).

[0042] After obtaining the response content corresponding to the input message, in step S220, if the response content includes a description associated with the virtual object in the three-dimensional space model 161, the response content and the recommended options associated with the virtual object are output to the user interface 131 for display via the cloud server 100B.

[0043] The large-scale language model 180 analyzes the input information in real time and infers whether the user's preferences are related to the virtual objects arranged in the three-dimensional spatial model 161. If the analysis shows that the virtual objects arranged in the three-dimensional spatial model 161 match the user's preferences, the front-end program 163 generates recommended options and outputs them to the user interface 131 for display. In one embodiment, the recommended options can be presented in the form of "buttons". In addition, if the response content contains descriptions associated with multiple virtual elements, multiple recommended options for the multiple virtual elements will be generated respectively.

[0044] Next, in step S225, when the recommended option is selected through the user interface 131, the user interface 131 automatically sends control commands to the cloud server 100B, so that the cloud server 100B drives the virtual camera corresponding to the three-dimensional space model 161 to move to the designated positioning point corresponding to the virtual object, and executes the real-time rendering program to place the designated panoramic image corresponding to the designated positioning point into the three-dimensional space model 161 as a material.

[0045] Next, in step S230, a 3D space model 161 with the specified panoramic image placed inside is displayed through the user interface 131. That is, in practical applications, when a user wants to view the recommended options, they can manually touch the recommended options in the user interface 131 (the user interface 131 automatically generates control commands after detecting the touch), or use voice control to issue control commands to select the recommended options. Then, the user interface 131 sends the control commands to the cloud server 100B, and the front-end program 163 analyzes the location information of the recommended options in the virtual space, thereby guiding the user's viewing perspective to a suitable perspective. That is, the virtual camera is moved to the specified positioning point corresponding to the virtual object bound to the recommended option, and a real-time rendering program is executed to place the specified panoramic image corresponding to the specified positioning point into the 3D space model 161 as a material.

[0046] FIG5 is a flowchart of an automatic navigation system according to an embodiment of the present invention. Referring to FIG5, when the three-dimensional space model 161 is displayed through the user interface 131 of the user terminal device 100A, in step S501, the user U may, for example, make voice input through the voice sensor (input device 120) and generate a corresponding voice signal accordingly.

[0047] The query results corresponding to the input message are obtained by querying the spatial arrangement data 52 of the three-dimensional spatial model 161 through the large language model 180 provided by the cloud server 100B, and the query results are converted into response content that conforms to the prompt engineering 51 through the large language model 180 provided by the cloud server 100B.

[0048] This stage has two key functions: cue word engineering and Retrieval-Augmented Generation (RAG). The cue word engineering 51 primarily develops general-purpose cue words, giving the large language model 180 characteristics similar to a business or tour guide, and debugging the large language model 180 to be applicable to different types of virtual spaces. Furthermore, in this application scenario, to ensure the large language model 180 provides effective and realistic data, the scope of data that the large language model 180 can access is narrowed in the RAG aspect, thereby developing a suitable response method. For example, user U asks "Have you monitored any related products?" (input message). The large language model 180 will output the corresponding response content based on the space layout data 52.

[0049] After receiving the response content, in step S503, the large language model 180 determines whether the response content includes a description associated with the virtual object in the 3D spatial model 161. If the response content does not include a description associated with the virtual object in the 3D spatial model 161, in step S505, the response content is directly output to the user interface 131 for display. If the response content includes a description associated with the virtual object in the 3D spatial model 161, in step S507, the foreground program 163 generates recommended options associated with the virtual object and obtains the configuration position corresponding to the virtual position. Then, the response content, recommended options, and configuration position are associated and output to the user interface 131.

[0050] During the interaction between a user viewing the 3D spatial model 161 through the user interface 131 and the large language model 180, the large language model 180 analyzes the dialogue in real time to infer whether the user's preferences are related to the exhibits (virtual objects) in the 3D spatial model 161. If the analysis finds that an exhibit in the 3D spatial model 161 matches the user's preferences, the front-end program 163 generates recommended options. For example, another window pops up in the 3D spatial model 161, providing recommended options for the user to choose from. If the user selects one of the recommended options, the front-end program 163 analyzes the coordinates of the virtual object corresponding to the recommended option in the virtual space and guides the user's viewing perspective to a suitable angle.

[0051] In summary, this invention combines artificial intelligence to provide automatic navigation functionality, offering guidance based on content of interest to the user. This creates a rapid deployment of virtual displays and intelligent customer service and navigation functions. A large language model is used to perform recommendation analysis on input information, and the spatial coordinates and movement behavior of recommended options are calculated to achieve automatic navigation. This disclosure achieves the effect of intelligent customer service with virtual spatial navigation capabilities. Compared to existing generative AI models, which are prone to AI illusions, difficult to build, have cumbersome training processes, and are limited to two-dimensional experiences, this disclosure allows generative AI models to query spatial layout data in three-dimensional spatial models, further breaking through the current two-dimensional experience mode. [Simplified Explanation of the Diagram]

[0052] Figure 1 is a block diagram of an automatic navigation system according to an embodiment of the present invention. Figure 2 is a flowchart of an automatic navigation method according to an embodiment of the present invention. Figure 3 is a schematic diagram of a three-dimensional space model including multiple positioning points according to an embodiment of the present invention. Figures 4A to 4D are schematic diagrams of the operation of a spatial arrangement interface according to an embodiment of the present invention. Figure 5 is a flowchart of executing an automatic navigation system according to an embodiment of the present invention.

Claims

1. An automated tour guide method, applicable to an automated tour guide system, the automated tour guide system including a user terminal device and a cloud server, the automated tour guide method comprising: The user interface of the client device communicates with the cloud server and displays a 3D spatial model provided by the cloud server in the user interface. The 3D spatial model has spatial layout data and includes multiple positioning points. A database of the cloud server includes multiple panoramic views corresponding to the positioning points. The user device receives an input message through an input device and transmits the input message to the cloud server through the user interface. The cloud server obtains a response content corresponding to the input message from the spatial layout data. If the response content includes a description associated with a virtual object in the 3D spatial model, the cloud server outputs the response content and a recommended option associated with the virtual object to the user interface for display. The virtual object corresponds to a specified positioning point, which is one of the positioning points. When the recommended option is selected through the user interface, a control command is automatically sent to the cloud server through the user interface, causing the cloud server to move a virtual camera corresponding to the 3D spatial model to the designated positioning point corresponding to the virtual object, and execute a real-time rendering program to place a designated panoramic image corresponding to the designated positioning point as a material into the 3D spatial model, wherein the designated panoramic image is one of the panoramic images; and the 3D spatial model with the designated panoramic image placed is displayed through the user interface.

2. The automatic navigation method as described in claim 1 further includes: This cloud server provides a space layout interface; And through the space layout interface, the arrangement of a virtual space corresponding to the three-dimensional space model is set, thereby obtaining the space layout data.

3. The automatic navigation method as described in claim 1, wherein the spatial arrangement data includes a plurality of virtual objects corresponding to a plurality of exhibits, each of the virtual objects having a corresponding configuration position and one of the positioning points.

4. The automatic navigation method as described in claim 1 further includes: In response to the communication connection between the user interface of the user terminal device and the cloud server, the real-time rendering program is executed through the cloud server to place a preset panoramic image corresponding to a preset position as a material into the three-dimensional space model; and the three-dimensional space model with the preset panoramic image placed in it is displayed through the user interface, wherein the preset position is one of the positioning points and the preset panoramic image is one of the panoramic images.

5. The automatic navigation method as described in claim 1, wherein the step of obtaining the response content corresponding to the input message from the spatial layout data via the cloud server includes: The system queries the spatial layout data of the three-dimensional spatial model through a large language model provided by the cloud server to obtain a query result corresponding to the input message. And through the cloud server, the large language model is used to convert the query results into response content that conforms to a prompt word project.

6. The automatic navigation method as described in claim 1, wherein when the response content includes the description associated with the virtual object in the three-dimensional space model, it includes: The recommended options associated with the virtual object are generated through a front-end program provided by the cloud server, and the response content and the recommended options are output to the user interface.

7. An automated tour guide system, comprising: A cloud server and a client device, wherein the client device includes a first storage unit having a user interface, an input device, and a first processor; the cloud server includes a second processor; the cloud server is configured to provide a three-dimensional spatial model, wherein the three-dimensional spatial model has spatial arrangement data and includes multiple positioning points; a database of the cloud server includes multiple panoramic images corresponding to the positioning points; in the client device, the first processor is configured to: communicate with the cloud server through the user interface and display the three-dimensional spatial model provided by the cloud server, wherein the three-dimensional spatial model has spatial arrangement data, in the user interface; and receive an input message through the input device and transmit the input message to the cloud server through the user interface. In the cloud server, the second processor is configured to: obtain a response content corresponding to the input message from the spatial layout data, and determine whether the response content includes a description associated with a virtual object in the three-dimensional spatial model; and if the response content includes the description associated with the virtual object in the three-dimensional spatial model, output the response content and a recommended option associated with the virtual object to the user interface for display via the cloud server, wherein the virtual object corresponds to a specified location point, and the specified location point is one of the location points; In the client device, the first processor is configured to: automatically send a control command to the cloud server via the user interface when the recommended option is selected via the user interface; In the cloud server, the second processor is configured to: in response to receiving the control command, move a virtual camera corresponding to the 3D spatial model to the designated positioning point corresponding to the virtual object, and execute a real-time rendering program to place a designated panoramic image corresponding to the designated positioning point as a material into the 3D spatial model; and display the 3D spatial model with the designated panoramic image placed into the user interface, wherein the designated panoramic image is one of the panoramic images.

8. An automated navigation method, utilizing a processor to perform the following steps: displaying a three-dimensional spatial model through a user interface, wherein the three-dimensional spatial model has spatial layout data and includes a plurality of location points, each location point corresponding to a plurality of panoramic views; receiving an input message through an input device; obtaining a response content corresponding to the input message through the user interface, wherein the response content is generated based on the spatial layout data; and, if the response content includes a description associated with a virtual object in the three-dimensional spatial model, displaying the response content and a recommended option corresponding to the virtual object in the user interface, wherein the virtual object corresponds to a designated location point, the designated location point being one of the location points; When the recommended option is selected through the user interface, the user interface drives a virtual camera corresponding to the 3D spatial model to move to the designated positioning point corresponding to the virtual object, and drives the execution of a real-time rendering program to place a designated panoramic image corresponding to the designated positioning point as a material into the 3D spatial model, wherein the designated panoramic image is one of the panoramic images; and the user interface displays the 3D spatial model with the designated panoramic image placed in it.

9. A user terminal device, comprising: A storage device, including a user interface; One input device; The system also includes a processor coupled to the storage and the input device, configured to: display a three-dimensional spatial model through the user interface, wherein the three-dimensional spatial model has spatial arrangement data and includes a plurality of positioning points, each corresponding to a plurality of panoramic views; receive an input message through the input device; obtain a response content corresponding to the input message through the user interface, wherein the response content is generated based on the spatial arrangement data; and, if the response content includes a description associated with a virtual object in the three-dimensional spatial model, display the response content and a recommended option corresponding to the virtual object in the user interface, wherein the virtual object corresponds to a specified positioning point, the specified positioning point being one of the positioning points; When the recommended option is selected through the user interface, the user interface drives a virtual camera corresponding to the 3D spatial model to move to the designated positioning point corresponding to the virtual object, and drives the execution of a real-time rendering program to place a designated panoramic image corresponding to the designated positioning point as a material into the 3D spatial model, wherein the designated panoramic image is one of the panoramic images; and the user interface displays the 3D spatial model with the designated panoramic image placed in it.