Immersive scene interpretation system based on AI

By combining AI-driven RFID positioning and voice analysis technology with holographic projection, an immersive and interactive museum tour experience has been achieved, solving the one-way linear problem of traditional tour methods and enhancing the audience's immersive experience and interactivity.

CN121768293APending Publication Date: 2026-03-31SHENZHEN YOUBAO INTERNET OF THINGS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional museum interpretation methods present information in a linear and one-way manner, lacking dynamic interaction and visual presentation with the historical background, scientific principles, or artistic context of the exhibits, thus failing to meet the in-depth cognitive needs of modern audiences.

Method used

The system employs an AI-based immersive scene explanation system. Through RFID positioning and voice semantic analysis, it can identify the location and keywords of the explainer in real time, automatically associate and activate the corresponding scene information and multi-dimensional explanation content from the database, and present the virtual scene through holographic projection, supporting the interaction between the explainer and the virtual scene.

Benefits of technology

It enhances the audience's immersive experience and interactivity, enables dynamic presentation of information and instant interest-based discovery, and improves the interactivity and immersion of museum interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768293A_ABST
    Figure CN121768293A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-based immersive scene explanation system, and relates to the technical field of AI museum explanation, and the system comprises a monitoring center which is in communication connection with a database which is used for storing scene information and explanation information of each point location in a target area; the data interaction module is composed of a plurality of RFID units, is arranged in each point location in the target area, and is in communication connection with the handheld terminal; the AI scene construction module is used for constructing a virtual scene according to the scene information and the explanation information of the point location, and projecting the virtual scene to a holographic scene area in the point location in a holographic manner; the scene interaction module is used for an explainer to interact with a virtual scene in the holographic scene area; according to the method, the explanation information corresponding to the point location is dynamically presented in a holographic projection mode, the immersion experience feeling is enhanced, an explainer can interact with any factor unit through the operation association point on the handheld terminal, and the interactivity between a virtual scene and the explainer is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI museum interpretation technology, specifically an AI-based immersive scene interpretation system. Background Technology

[0002] With the evolution of museum exhibition concepts, traditional interpretation methods are no longer sufficient to meet the in-depth cognitive needs of modern audiences. Currently, museums commonly use audio guides or narration by guides, presenting information linearly and unidirectionally, lacking dynamic interaction and visual representation of the exhibits' historical background, technological principles, or artistic context. This is especially true when showcasing complex structures, evolutionary processes, or abstract concepts, making it difficult for audiences to form a direct and three-dimensional understanding. Although some museums have introduced QR code scanning or fixed-location screen displays as supplementary methods, these interactions are passive, information is fragmented, and they cannot extend content or explore connections based on the audience's immediate interests. Therefore, we now offer an AI-based immersive scene interpretation system. Summary of the Invention

[0003] The purpose of this invention is to provide an AI-based immersive scene narration system.

[0004] The objective of this invention can be achieved through the following technical solution: an AI-based immersive scene narration system, including a monitoring center, wherein the monitoring center is communicatively connected to:

[0005] The database is used to store scene information and explanation information for each point within the target area;

[0006] The data interaction module consists of several RFID units, which are deployed at various points within the target area and communicate with the handheld terminal.

[0007] The AI ​​scene construction module is used to construct virtual scenes based on the scene information and explanation information of the location, and to holographically project the virtual scene onto the holographic scene area within the location.

[0008] The scene interaction module is used for the presenter to interact with the virtual scene within the holographic scene area.

[0009] Furthermore, the target area is equipped with several points;

[0010] The database contains sub-databases associated with each location. Each sub-database stores scene information, explanation information, and factor units for the corresponding location. These factor units are associated with each other, and different associations have different optional connection methods.

[0011] The scene information includes the location of the point, the area covered by the point, and the holographic scene area, wherein the holographic scene area is located within the area covered by the point.

[0012] The explanatory information includes text content, image content, video content, and the correlation between different types of content and scene information and factor units.

[0013] Furthermore, each RFID unit is configured with a corresponding signal coverage area, and each RFID unit is configured with a unique device identifier.

[0014] The handheld terminal is held by the guide. The handheld terminal searches for the signal of the RFID unit in real time. When the handheld terminal is within the signal coverage range of the RFID unit, the signal of the corresponding RFID unit can be searched by the handheld terminal. The guide connects the handheld terminal to the RFID unit for communication based on the signal of the searched RFID unit.

[0015] After the handheld terminal establishes a communication connection with the RFID, the RFID terminal sends an access command to the monitoring center. The monitoring center then activates the sub-database associated with the RFID terminal based on the received access command.

[0016] Furthermore, the handheld terminal is equipped with an audio acquisition unit and an audio processing unit;

[0017] The audio acquisition unit is used to acquire the voice information of the narrator in real time, and then the audio processing unit performs semantic analysis on the acquired voice information to obtain the corresponding voice keywords.

[0018] The obtained voice keywords are uploaded to the monitoring center.

[0019] Furthermore, the AI ​​scene construction module constructs a virtual scene based on the scene information and explanation information of the location, and the process of holographically projecting the virtual scene onto the holographic scene area within the location includes:

[0020] Set trigger keywords for each location;

[0021] Match the obtained voice keywords with the trigger keywords. If the match is successful, then:

[0022] Construct a three-dimensional spatial coordinate system within the holographic scene area;

[0023] The voice keywords obtained from the narrator's voice information are matched with the narration information in the database, and the successfully matched narration information is recorded as the response content.

[0024] Obtain and label scene information and factor units that are related to the response content;

[0025] Based on the marked scene information and factor units, corresponding scene modules and virtual factor modules are generated, and the generated scene modules and virtual factor modules are imported into the handheld terminal. After the narrator selects and confirms the scene modules and virtual factor modules in the handheld terminal, the scene modules and virtual factor modules are projected to the corresponding positions in the three-dimensional spatial coordinate system to obtain a virtual scene synchronized with the narrator's voice information.

[0026] Based on the scene module and virtual factor module projected in the three-dimensional spatial coordinate system, corresponding operation association points are generated in the handheld terminal of the narrator.

[0027] Furthermore, the process of generating scene modules based on scene information includes:

[0028] The holographic scene area is divided into several sub-regions;

[0029] Select at least one sub-region based on the response content, and mark the center point of the sub-region as the reference point corresponding to the response content;

[0030] Based on the specific content of the response, determine the response range of the response content, and generate an initial scene module centered on the reference point based on the response range;

[0031] Sub-regions not fully covered by the response range within the sub-regions occupied by the initial scene module are marked as redundant boundary regions;

[0032] The portion of the redundant boundary area not covered by the response range is recorded as the redundant portion of the redundant boundary area. The portion covered by the response range is taken as the scene module corresponding to the response content, and the scene module is associated with the reference point.

[0033] Furthermore, the process of generating virtual factor modules based on factor units includes:

[0034] The factor units associated with the response content are marked and summarized to obtain a factor unit set;

[0035] Select any factor unit within the factor unit set, obtain other factor units that are related to the selected factor unit, and determine the connection method between the selected factor unit and other factor units based on the existing relationship;

[0036] The selected factor unit is connected to other factor units according to the selected connection method to obtain the factor merging unit corresponding to the selected factor unit;

[0037] Then select another factor unit and obtain the factor merging unit corresponding to the selected factor unit through the above method. In this way, obtain the factor merging units corresponding to all factor units.

[0038] Obtain the frequency of occurrence of each factor unit in all factor merged units, and take the factor unit with the highest frequency as the benchmark unit;

[0039] All factor merging units are merged a second time, overlapping factor units are fitted, and only one is retained to obtain a virtual factor module. The reference unit of the virtual factor module is then associated with a reference point and projected into the scene module.

[0040] Furthermore, the process of the presenter interacting with the virtual scene within the holographic scene area through the scene interaction module includes:

[0041] Based on the virtual factor module and the scene module, corresponding operation association points are generated, and several operation items corresponding to the operation association points are generated.

[0042] Each operation item of the operation association point corresponding to the virtual factor module corresponds to a factor unit, and each operation item is set with several operation contents;

[0043] The presenter can interact with the scene module, virtual factor module, and various factor units in the three-dimensional coordinate system by performing corresponding operations at each operation-related point.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] 1. Through RFID positioning and voice semantic analysis, the system can identify the location of the guide and the keywords of the explanation in real time, and automatically associate and activate the corresponding scene information and multi-dimensional explanation content from the database. It then transforms these into a holographic projection virtual scene synchronized with the explanation language, dynamically presenting the explanation information corresponding to the location in a holographic projection manner, thereby enhancing the immersive experience.

[0046] 2. By associating factor units with scene interaction modules, factor units are grouped into virtual factor modules based on their correlation and projected into the virtual scene. This allows presenters to interact with any factor unit through operation links on their handheld terminals, enhancing the interactivity between the virtual scene and the presenters. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0048] Figure 1 This is a schematic diagram of the present invention. Detailed Implementation

[0049] like Figure 1 As shown, the AI-based immersive scene narration system includes a monitoring center, which has the following communication connections:

[0050] The database is used to store scene information and explanation information for each point within the target area;

[0051] The data interaction module consists of several RFID units, which are deployed at various points within the target area and communicate with the handheld terminal.

[0052] The AI ​​scene construction module is used to construct virtual scenes based on the scene information and explanation information of the location, and to holographically project the virtual scene onto the holographic scene area within the location.

[0053] The scene interaction module is used for the presenter to interact with the virtual scene within the holographic scene area.

[0054] It should be further explained that, in the specific implementation process, several points are set up in the target area;

[0055] The database contains sub-databases associated with each location. Each sub-database stores scene information, explanation information, and factor units for the corresponding location. These factor units are associated with each other, and different associations have different optional connection methods.

[0056] The scene information includes the location of the point, the area covered by the point, and the holographic scene area, wherein the holographic scene area is located within the area covered by the point.

[0057] The explanatory information includes text content, image content, video content, and the correlation between different types of content and scene information and factor units.

[0058] It should be further explained that, in the specific implementation process, the RFID unit is set with a corresponding signal coverage range, and each RFID unit is set with a unique device identifier, and the device identifier is matched with the signal emitted by the RFID unit;

[0059] The handheld terminal is held by the guide. The handheld terminal searches for the signal of the RFID unit in real time. When the handheld terminal is within the signal coverage range of the RFID unit, the signal of the corresponding RFID unit can be searched by the handheld terminal. The guide connects the handheld terminal to the RFID unit for communication based on the signal of the searched RFID unit.

[0060] After the handheld terminal establishes a communication connection with the RFID, the RFID terminal sends an access command to the monitoring center. The monitoring center then activates the sub-database associated with the RFID terminal based on the received access command.

[0061] The handheld terminal is equipped with an audio acquisition unit and an audio processing unit;

[0062] The audio acquisition unit is used to acquire the voice information of the narrator in real time, and then the audio processing unit performs semantic analysis on the acquired voice information to obtain the corresponding voice keywords. It should be further noted that the process of semantic analysis of voice information is a common technical means used by those skilled in the art, and will not be described in detail here.

[0063] The obtained voice keywords are uploaded to the monitoring center.

[0064] It should be further explained that, in the specific implementation process, the AI ​​scene construction module constructs a virtual scene based on the scene information and explanation information of the location, and the process of holographically projecting the virtual scene onto the holographic scene area within the location includes:

[0065] Set trigger keywords for each location;

[0066] Match the obtained voice keywords with the trigger keywords. If the match is successful, then:

[0067] Construct a three-dimensional spatial coordinate system within the holographic scene area;

[0068] The voice keywords obtained from the narrator's voice information are matched with the narration information in the database, and the successfully matched narration information is recorded as the response content.

[0069] Obtain and label scene information and factor units that are related to the response content;

[0070] Based on the marked scene information and factor units, corresponding scene modules and virtual factor modules are generated, and the generated scene modules and virtual factor modules are imported into the handheld terminal. After the narrator selects and confirms the scene modules and virtual factor modules in the handheld terminal, the scene modules and virtual factor modules are projected to the corresponding positions in the three-dimensional spatial coordinate system to obtain a virtual scene synchronized with the narrator's voice information.

[0071] Based on the scene module and virtual factor module projected in the three-dimensional spatial coordinate system, corresponding operation association points are generated in the handheld terminal of the narrator.

[0072] It should be further explained that, in the specific implementation process, the process of generating scene modules based on scene information includes:

[0073] The holographic scene area is divided into several sub-regions; it should be noted that each sub-region is the same size.

[0074] Select at least one sub-region based on the response content, and mark the center point of the sub-region as the reference point corresponding to the response content;

[0075] Based on the specific content of the response, determine the response range of the response content, and generate an initial scene module centered on the reference point based on the response range;

[0076] Sub-regions not fully covered by the response range within the sub-regions occupied by the initial scene module are marked as redundant boundary regions;

[0077] The portion of the redundant boundary area not covered by the response range is recorded as the redundant portion of the redundant boundary area. The portion covered by the response range is taken as the scene module corresponding to the response content, and the scene module is associated with the reference point.

[0078] It should be further explained that, in the specific implementation process, the process of generating virtual factor modules based on factor units includes:

[0079] The factor units associated with the response content are marked and summarized to obtain a factor unit set;

[0080] Select any factor unit within the factor unit set, obtain other factor units that are related to the selected factor unit, and determine the connection method between the selected factor unit and other factor units based on the existing relationship;

[0081] The selected factor unit is connected to other factor units according to the selected connection method to obtain the factor merging unit corresponding to the selected factor unit;

[0082] Then select another factor unit and obtain the factor merging unit corresponding to the selected factor unit through the above method. In this way, obtain the factor merging units corresponding to all factor units.

[0083] Obtain the frequency of occurrence of each factor unit in all factor merged units, and take the factor unit with the highest frequency as the benchmark unit;

[0084] All factor merging units are merged a second time, overlapping factor units are fitted, and only one is retained to obtain a virtual factor module. The reference unit of the virtual factor module is then associated with a reference point and projected into the scene module.

[0085] It should be further explained that, in the specific implementation process, the process by which the presenter interacts with the virtual scene within the holographic scene area through the scene interaction module includes:

[0086] Based on the virtual factor module and the scene module, corresponding operation association points are generated, and several operation items corresponding to the operation association points are generated.

[0087] Each operation item of the operation association point corresponding to the virtual factor module corresponds to a factor unit, and each operation item is set with several operation contents;

[0088] The presenter can perform corresponding operations at each operation-related point to interact with the scene module, virtual factor module, and various factor units in the three-dimensional coordinate system. The operations include movement, rotation, scaling, and interaction between different factor units. The specific functions can be adapted by technical personnel according to the different presentation information, scene information, virtual factor modules, and factor units at different points, which will not be elaborated here.

[0089] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications or equivalent substitutions made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. An AI-based immersive scene narration system comprising a monitoring center, characterized in that, The monitoring center is in communication connection with: a database for storing scene information and explanation information of each point in the target area; a data interaction module composed of a plurality of RFID units, arranged in each point in the target area, and in communication connection with the handheld terminal; an AI scene construction module for constructing a virtual scene according to the scene information and explanation information of the point, and holographically projecting the virtual scene to a holographic scene area in the point; a scene interaction module for the explanation personnel to interact with the virtual scene in the holographic scene area. 2.The AI-based immersive scene illustration system of claim 1, wherein, The target area is provided with a plurality of points; The database is provided with a data sub-library associated with each point, and each data sub-library stores the scene information, explanation information and factor unit of the corresponding point, the factor units are associated, and different association relationships have different selectable connection modes; The scene information includes point position, point coverage area and holographic scene area, wherein the holographic scene area is in the point coverage area; The explanation information includes text content, picture content, video content and the association between different types of content and scene information and factor units. 3.The AI-based immersive scene interpretation system of claim 2, wherein, The RFID unit is provided with a corresponding signal coverage range, and each RFID unit is provided with a unique device identifier; The handheld terminal is held by the explanation personnel, and the handheld terminal searches the signal of the RFID unit in real time. When the handheld terminal is in the signal coverage range of the RFID unit, the signal of the corresponding RFID unit can be searched by the handheld terminal. The explanation personnel searches the signal of the RFID unit, and the handheld terminal is in communication connection with the RFID unit; After the communication connection between the handheld terminal and the RFID is completed, the access instruction is sent to the monitoring center from the RFID terminal, and the monitoring center activates the data sub-library associated with the RFID terminal according to the received access instruction. 4.The AI-based immersive scene interpretation system of claim 3, wherein, The handheld terminal is provided with an audio acquisition unit and an audio processing unit; The audio acquisition unit is used to acquire the voice information of the explanation personnel in real time, and the audio processing unit is used to perform semantic analysis on the acquired voice information to obtain the corresponding voice keyword; The obtained voice keyword is uploaded to the monitoring center. 5.The AI-based immersive scene interpretation system of claim 4, wherein, The process of constructing a virtual scene by the AI scene construction module according to the scene information and explanation information of the point, and holographically projecting the virtual scene to the holographic scene area in the point includes: setting a trigger keyword for each point; matching the obtained voice keyword with the trigger keyword, if the matching is successful, then: constructing a three-dimensional space coordinate system in the holographic scene area; matching the voice keyword obtained from the voice information of the explanation personnel with the explanation information in the data sub-library, and recording the matching successful explanation information part as the response content; obtaining and marking the scene information and factor units associated with the response content; According to the labeled scene information and the factor unit, corresponding scene modules and virtual factor modules are generated, and the generated scene modules and virtual factor modules are imported into the handheld terminal. After the guide personnel selects and confirms the scene modules and virtual factor modules in the handheld terminal, the scene modules and virtual factor modules are projected to corresponding positions in the three-dimensional space coordinate system, and a virtual scene synchronized with the voice information of the guide personnel is obtained. According to the scene modules and virtual factor modules projected in the three-dimensional space coordinate system, corresponding operation association points are generated in the handheld terminal of the guide personnel. 6.The AI-based immersive scene interpretation system of claim 5, wherein, The process of generating a scene module according to scene information includes: dividing the holographic scene area into a plurality of sub-areas; selecting at least one sub-area according to the response content, and marking the center point of the sub-area as a reference point corresponding to the response content; determining the response range of the response content according to the specific content of the response content, and generating an initial scene module centered on the reference point according to the response range; marking the sub-area not completely covered by the response range in the sub-area occupied by the initial scene module as a redundant boundary area; marking the part not covered by the response range in the redundant boundary area as a redundant part of the redundant boundary area, taking the covered part of the response range as a scene module corresponding to the response content, and associating the scene module with the reference point. 7.The AI-based immersive scene interpretation system of claim 6, wherein, The process of generating a virtual factor module according to a factor unit includes: marking and collecting the factor units associated with the response content to obtain a factor unit set; selecting any factor unit in the factor unit set, obtaining other factor units associated with the selected factor unit, and determining the connection mode between the selected factor unit and the other factor units according to the existing association; connecting the selected factor unit and the other factor units according to the selected connection mode to obtain a factor merging unit corresponding to the selected factor unit; selecting another factor unit, obtaining a factor merging unit corresponding to the selected factor unit by the above method, and so on to obtain factor merging units corresponding to all factor units; obtaining the number of occurrences of each factor unit in all factor merging units, and taking the factor unit with the most occurrences as a reference unit; performing secondary merging on all factor merging units, fitting overlapping factor units, and retaining only one to obtain a virtual factor module, and associating the reference unit of the virtual factor module with the reference point and projecting it into the scene module. 8.The AI-based immersive scene interpretation system of claim 7, wherein, The process of the guide personnel interacting with the virtual scene in the holographic scene area through the scene interaction module includes: generating corresponding operation association points according to the virtual factor module and the scene module, and generating a plurality of operation items corresponding to the operation association points; each operation item of the operation association point corresponding to the virtual factor module corresponds to a factor unit, and each operation item is provided with a plurality of operation contents; The guide personnel can interact with the scene module, the virtual factor module and each factor unit in the three-dimensional space coordinate system by performing corresponding operation contents on each operation association point.