A VR / AR Virtual Scene Intelligent Interaction Method and System Based on 3D Models
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]现有方案大多直接将三维图像整体用于模型构建和交互响应,部分复杂场景变化在现实世界中可能感受正常,但是在虚拟场景下用户感知敏感程度不同且虚拟场景下会设置诸多交互指示点,视觉负荷更大
本发明先获取目标场景的三维图像,按照视野角度对所述三维图像进行分区;随后对于划分出的每个视觉区域,根据所有体素点的深度信息筛选得到背景点和场景点,其中,所述场景点包括交互点和非交互点;将所有体素点去除深度信息以得到二维图像,根据二维图像中场景点和背景点的分隔线确定场景轮廓区域,场景轮廓区域中的所有交互点按照坐标排序连接得到交互指示路径;根据场景轮廓区域和二维图像确定场景密集度,根据场景轮廓区域和交互指示路径确定交互指示密集度,根据所述场景密集度和交互指示密集度确定视觉密集特征;然后以相邻区域之间的视觉密集特征的差异满足的视觉负荷标准为目标,对三维图像进行优化。本发明在渲染场景前,对于三维图像进行优化以使视野角度切换时场景差异更小,用户感受更为柔和且视觉压力更小。最后根据优化处理后的三维图像构建目标场景的虚拟三维模型,将所述虚拟三维模型渲染在用户佩戴的VR/AR设备的显示窗口中,用户能够通过操作虚拟或实体手柄的转动元件对所述显示窗口中的实时虚拟场景进行滑移,并且用户能够通过操作虚拟或实体手柄的按钮元件触发所述显示窗口中的交互点指示。本发明的用户能够通过虚拟或实体手柄能够便捷地进行视野切换和智能交互。
Smart Images

Figure CN122569752A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent interaction technology for VR / AR virtual scenes, specifically to a VR / AR virtual scene intelligent interaction method and system based on a three-dimensional model. Background Technology
[0002] With the rapid development of virtual reality (VR) and augmented reality (AR) technologies, immersive virtual scene interaction has been widely applied in many fields requiring visual display. In VR / AR interaction scenarios, users typically wear head-mounted display devices and use handheld controllers (handles) to operate objects and interactive points in the virtual scene. Unlike ordinary two-dimensional visual scenes, in three-dimensional visual scenes, users can slide windows to experience visual images from different angles, and users are more sensitive to scene changes and interactive indicators within the scene.
[0003] Most existing solutions directly use 3D images as a whole for model building and interactive responses. While some complex scene changes may be perceived normally in the real world, user sensitivity differs in virtual scenes, and the numerous interactive markers in virtual scenes increase visual load. When switching perspectives in virtual scenes, the significant differences in dense information can cause visual fatigue or even dizziness, directly impacting the user experience of VR / AR intelligent interaction. Summary of the Invention
[0004] The purpose of this invention is to provide a VR / AR virtual scene intelligent interaction method and system based on a three-dimensional model. Before rendering the virtual three-dimensional model on the display window of the VR / AR device worn by the user, the original three-dimensional image is optimized so that the user can feel a smoother experience when sliding the field of view through the VR / AR device, resulting in a better user experience.
[0005] This invention provides a VR / AR virtual scene intelligent interaction method based on a 3D model, comprising: Acquire a 3D image of the target scene and partition the 3D image according to the field of view; For each defined visual region, background points and scene points are obtained by filtering based on the depth information of all voxel points, wherein the scene points include interactive points and non-interactive points; depth information is removed from all voxel points to obtain a two-dimensional image, and the scene outline region is determined based on the dividing line between the scene points and background points in the two-dimensional image, and all interactive points in the scene outline region are connected according to their coordinates to obtain the interactive indication path. The scene density is determined based on the scene outline region and the two-dimensional image; the interaction indicator density is determined based on the scene outline region and the interaction indicator path; and the visual density features are determined based on the scene density and the interaction indicator density. The three-dimensional image is optimized based on the visual load criterion satisfied by the difference in visual density features between adjacent regions; A virtual 3D model of the target scene is constructed based on the optimized 3D image, and the virtual 3D model is rendered in the display window of the VR / AR device worn by the user. The user can slide the real-time virtual scene in the display window by operating the rotation element of the virtual or physical controller, and the user can trigger the interactive point indication in the display window by operating the button element of the virtual or physical controller.
[0006] Preferably, partitioning the three-dimensional image according to visual angle includes: Centered on a pre-set user viewpoint, the horizontal and vertical field of view are divided equally at preset angle intervals to form multiple fan-shaped or grid-like visual areas.
[0007] As a preferred option, background and scene points are selected based on the depth information of all voxel points, including: A depth threshold is set, and voxel points with a depth value greater than the depth threshold are classified as background points, while voxel points with a depth value less than the depth threshold are classified as scene points.
[0008] Preferably, the interaction instruction path is obtained by connecting all interaction points in the scene outline area according to their coordinates, including: Sort all interactive points in ascending order of their horizontal or vertical coordinates and connect them sequentially to obtain the interactive indication path.
[0009] As a preferred method, determining scene density based on scene outline regions and two-dimensional images includes: Calculate the pixel area ratio of the scene outline region in the two-dimensional image, and use the pixel area ratio as the scene density.
[0010] Preferably, determining the interaction indicator density based on the scene outline region and the interaction indicator path includes: The ratio of the total length of the interaction instruction path to the perimeter of the scene outline region is calculated to determine the interaction instruction density.
[0011] Preferably, determining visual density features based on the scene density and interaction indicator density includes: The scene density and interaction indicator density are normalized and weighted summed to obtain visual density features.
[0012] Preferably, the visual load criteria that the differences in visual density features between adjacent regions need to meet include: The difference in visually dense features between adjacent regions needs to be less than a preset difference threshold.
[0013] Preferably, the optimization of the 3D image is based on the visual load criterion satisfied by the difference in visual density features between adjacent regions, including: For each visual region of the 3D image, the dividing line between the scene point and the background point in the 2D image of that visual region is used as the baseline. Scene points and background points within a preset range on both sides of the baseline are set as free points. The attributes of the free points can be set as scene points or background points. When the set attributes are inconsistent with the initial attributes, after the attributes are set, the color information of the point needs to be adaptively matched according to the nearest point with the same attribute. In the scene outline area, all scene points within a preset range centered on the interaction point are set as indicator candidate points, and each of the indicator candidate points can be set as an interaction point to replace the current interaction point. An iterative algorithm is used to iteratively solve the set states of the free points and the indicator candidate points until the difference in visual density features between adjacent regions meets the visual load standard. The optimized distribution of scene points, background points, and interaction points is mapped back to 3D space to obtain the optimized 3D image.
[0014] This invention also provides a VR / AR virtual scene intelligent interaction system based on a three-dimensional model, which includes: The acquisition module is used to acquire a 3D image of the target scene and partition the 3D image according to the field of view. The determination module is used to, for each divided visual region, filter background points and scene points based on the depth information of all voxel points, wherein the scene points include interactive points and non-interactive points; remove the depth information from all voxel points to obtain a two-dimensional image; determine the scene outline region based on the dividing line between the scene points and background points in the two-dimensional image; connect all interactive points in the scene outline region according to their coordinates to obtain an interactive indication path; determine the scene density based on the scene outline region and the two-dimensional image; determine the interactive indication density based on the scene outline region and the interactive indication path; and determine visual density features based on the scene density and the interactive indication density. The 3D optimization module is used to optimize 3D images based on the visual load criterion satisfied by the difference in visually dense features between adjacent regions. The interaction module is used to construct a virtual 3D model of the target scene based on the optimized 3D image, and render the virtual 3D model in the display window of the VR / AR device worn by the user. The user can slide the real-time virtual scene in the display window by operating the rotation element of the virtual or physical controller, and the user can trigger the interaction point indication in the display window by operating the button element of the virtual or physical controller.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention first acquires a 3D image of the target scene and partitions the 3D image according to the field of view. Then, for each partitioned visual region, background points and scene points are obtained by filtering based on the depth information of all voxel points. The scene points include interactive and non-interactive points. Depth information is removed from all voxel points to obtain a 2D image. The scene outline region is determined based on the dividing line between the center point and background points in the 2D image. All interactive points in the scene outline region are connected according to their coordinates to obtain an interactive indicator path. Scene density is determined based on the scene outline region and the 2D image. Interactive indicator density is determined based on the scene outline region and the interactive indicator path. Visual density features are determined based on the scene density and interactive indicator density. Then, the 3D image is optimized with the visual load standard satisfied by the difference in visual density features between adjacent regions as the target. This invention optimizes the 3D image before rendering the scene to minimize scene differences when switching field of view angles, resulting in a smoother user experience and less visual stress. Finally, a virtual 3D model of the target scene is constructed based on the optimized 3D image, and the virtual 3D model is rendered in the display window of the VR / AR device worn by the user. The user can slide the real-time virtual scene in the display window by operating the rotation element of the virtual or physical controller, and the user can trigger the interaction point indication in the display window by operating the button element of the virtual or physical controller. Users of this invention can conveniently switch fields of view and perform intelligent interactions using virtual or physical controllers. Attached Figure Description
[0016] Figure 1 A flowchart illustrating a VR / AR virtual scene intelligent interaction method based on a 3D model, provided as an embodiment of the present invention; Figure 2 This is a block diagram of a VR / AR virtual scene intelligent interaction system based on a three-dimensional model, provided as an embodiment of the present invention. Detailed Implementation
[0017] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0018] Combination Figure 1 This embodiment provides a VR / AR virtual scene intelligent interaction method based on a 3D model, which includes: Step S1: Obtain a 3D image of the target scene and partition the 3D image according to the field of view. The three-dimensional image is partitioned according to the field of view angle, which includes: taking the user's viewpoint position as the center, dividing the horizontal field of view and the vertical field of view equally according to the preset angle interval to form multiple fan-shaped or grid-shaped visual regions.
[0019] Step S2: For each divided visual region, background points and scene points are obtained by filtering based on the depth information of all voxel points, wherein the scene points include interactive points and non-interactive points; the depth information of all voxel points is removed to obtain a two-dimensional image, and the scene outline region is determined based on the dividing line between the scene points and background points in the two-dimensional image. All interactive points in the scene outline region are connected according to their coordinates to obtain the interactive indication path. The process of filtering background and scene points based on the depth information of all voxel points includes: setting a depth threshold, classifying voxel points with depth values greater than the threshold as background points, and classifying voxel points with depth values less than the threshold as scene points. It should be noted that the depth threshold can be set according to the specific scene.
[0020] The interaction path is obtained by connecting all interactive points in the scene outline area according to their coordinates. This includes sorting all interactive points in ascending order of their horizontal or vertical coordinates and connecting them sequentially.
[0021] Step S3: Determine scene density based on scene outline region and 2D image; determine interaction indicator density based on scene outline region and interaction indicator path; determine visual density features based on scene density and interaction indicator density. Determining scene density based on scene outline regions and 2D images includes: calculating the pixel area ratio of the scene outline regions in the 2D image, and using the pixel area ratio as the scene density. Determining interaction indicator density based on scene outline regions and interaction indicator paths includes: calculating the ratio of the total length of the interaction indicator path to the perimeter of the scene outline regions to determine the interaction indicator density. Determining visual density features based on scene density and interaction indicator density includes: performing a weighted sum of the scene density and interaction indicator density to obtain visual density features. This embodiment comprehensively considers scene factors and interaction indicator factors and constructs visual density features, using these as a benchmark to optimize 3D images, thereby improving the user's visual sensory experience during scrolling.
[0022] Step S4: Optimize the 3D image based on the visual load standard satisfied by the difference in visual dense features between adjacent regions. The visual load standard that the difference in visual dense features between adjacent regions needs to satisfy includes: the difference in visual dense features between adjacent regions needs to be less than a preset difference threshold.
[0023] Optimizing a 3D image with the goal of satisfying a visual load criterion based on the difference in visual density features between adjacent regions includes: For each visual region of the 3D image, using the dividing line between the center point and background point in the 2D image of that visual region as a baseline, scene points and background points within a preset range on both sides of the baseline are set as free points. The attributes of these free points can be set as either scene points or background points. When the set attributes are inconsistent with the initial attributes, after the attributes are set, the color information of the point needs to be adaptively matched based on the nearest point with the same attribute. As an example, the color information of the nearest point with the same attribute can be directly used. In the scene contour region, all scene points within a preset range centered on the interaction point are set as indicator candidate points. These indicator candidate points can all be set as interaction points to replace the current interaction point. An iterative algorithm is used to iteratively solve the setting states of the free points and the indicator candidate points until the difference in visual density features between adjacent regions satisfies the visual load criterion. As an example, the iterative algorithm can be a genetic iterative algorithm or other iterative algorithms with the same function in the prior art. The optimized distribution states of scene points, background points, and interaction points are mapped back to 3D space to obtain the optimized 3D image.
[0024] Step S5: Construct a virtual 3D model of the target scene based on the optimized 3D image, and render the virtual 3D model in the display window of the VR / AR device worn by the user. The user can slide the real-time virtual scene in the display window by operating the rotation element of the virtual or physical controller, and can trigger the interactive point indication in the display window by operating the button element of the virtual or physical controller. The user can switch the field of view and perform intelligent interactive operations through the virtual or physical controller. The interactive operation is convenient and easy, and the visual perception is soft throughout the process, resulting in a good interactive experience.
[0025] This embodiment also provides a VR / AR virtual scene intelligent interaction system based on a 3D model, which includes: The acquisition module is used to acquire a 3D image of the target scene and partition the 3D image according to the field of view. The determination module is used to, for each divided visual region, filter background points and scene points based on the depth information of all voxel points, wherein the scene points include interactive points and non-interactive points; remove the depth information from all voxel points to obtain a two-dimensional image; determine the scene outline region based on the dividing line between the scene points and background points in the two-dimensional image; connect all interactive points in the scene outline region according to their coordinates to obtain an interactive indication path; determine the scene density based on the scene outline region and the two-dimensional image; determine the interactive indication density based on the scene outline region and the interactive indication path; and determine visual density features based on the scene density and the interactive indication density. The 3D optimization module is used to optimize 3D images based on the visual load criterion satisfied by the difference in visually dense features between adjacent regions. The interaction module is used to construct a virtual 3D model of the target scene based on the optimized 3D image, and render the virtual 3D model in the display window of the VR / AR device worn by the user. The user can slide the real-time virtual scene in the display window by operating the rotation element of the virtual or physical controller, and the user can trigger the interaction point indication in the display window by operating the button element of the virtual or physical controller.
[0026] Based on the same inventive concept, this embodiment provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the aforementioned intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model.
[0027] Based on the same inventive concept, this embodiment provides an electronic device, which includes: one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model as described in Embodiment 1.
[0028] In this embodiment of the disclosure, the memory is specifically used to store application code that executes the scheme of the present disclosure, and its execution is controlled by the processor. That is, when the electronic device is running, the processor communicates with the memory via a bus, or the processor communicates with the memory via other means, so that the processor executes the application code stored in the memory, thereby executing the steps of the method described in any of the foregoing embodiments. The memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.
[0029] The processor may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0030] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0031] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0032] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0033] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0034] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0035] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0036] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0037] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A VR / AR virtual scene intelligent interaction method based on a 3D model, characterized in that, include: Acquire a 3D image of the target scene, and partition the 3D image according to the field of view; For each defined visual region, background points and scene points are obtained by filtering based on the depth information of all voxel points, wherein the scene points include interactive points and non-interactive points; depth information is removed from all voxel points to obtain a two-dimensional image, and the scene outline region is determined based on the dividing line between the scene points and background points in the two-dimensional image, and all interactive points in the scene outline region are connected according to their coordinates to obtain the interactive indication path. The scene density is determined based on the scene outline region and the two-dimensional image; the interaction indicator density is determined based on the scene outline region and the interaction indicator path; and the visual density features are determined based on the scene density and the interaction indicator density. The three-dimensional image is optimized based on the visual load criterion satisfied by the difference in visual density features between adjacent regions; A virtual 3D model of the target scene is constructed based on the optimized 3D image, and the virtual 3D model is rendered in the display window of the VR / AR device worn by the user. The user can slide the real-time virtual scene in the display window by operating the rotation element of the virtual or physical controller, and the user can trigger the interactive point indication in the display window by operating the button element of the virtual or physical controller.
2. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, Partitioning the three-dimensional image according to the viewing angle includes: Centered on a pre-set user viewpoint, the horizontal and vertical field of view are divided equally at preset angle intervals to form multiple fan-shaped or grid-like visual areas.
3. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, Based on the depth information of all voxel points, background and scene points are obtained by filtering: A depth threshold is set, and voxel points with a depth value greater than the depth threshold are classified as background points, while voxel points with a depth value less than the depth threshold are classified as scene points.
4. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, All interactive points in the scene outline area, sorted by coordinates and connected together, form the interactive instruction path, which includes: Sort all interactive points in ascending order of their horizontal or vertical coordinates and connect them sequentially to obtain the interactive indication path.
5. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, Determining scene density based on scene outline regions and 2D images includes: Calculate the pixel area ratio of the scene outline region in the two-dimensional image, and use the pixel area ratio as the scene density.
6. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, Determining the density of interaction instructions based on the scene outline area and interaction instruction path includes: The ratio of the total length of the interaction instruction path to the perimeter of the scene outline region is calculated to determine the interaction instruction density.
7. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, Determining visual density features based on the scene density and interaction indicator density includes: The scene density and interaction indicator density are normalized and weighted summed to obtain visual density features.
8. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, The visual load criteria that need to be met for the differences in visual density features between adjacent regions include: The difference in visually dense features between adjacent regions needs to be less than a preset difference threshold.
9. The intelligent interaction method for VR / AR virtual scenes based on a three-dimensional model according to claim 1, characterized in that, Optimization of 3D images, targeting the visual load criterion satisfied by the differences in visual density features between adjacent regions, includes: For each visual region of the 3D image, the dividing line between the scene point and the background point in the 2D image of that visual region is used as the baseline. Scene points and background points within a preset range on both sides of the baseline are set as free points. The attributes of the free points can be set as scene points or background points. When the set attributes are inconsistent with the initial attributes, after the attributes are set, the color information of the point needs to be adaptively matched according to the nearest point with the same attribute. In the scene outline area, all scene points within a preset range centered on the interaction point are set as indicator candidate points, and each of the indicator candidate points can be set as an interaction point to replace the current interaction point. An iterative algorithm is used to iteratively solve the set states of the free points and the indicator candidate points until the difference in visual density features between adjacent regions meets the visual load standard. The optimized distribution of scene points, background points, and interaction points is mapped back to 3D space to obtain the optimized 3D image.
10. A VR / AR virtual scene intelligent interaction system based on a three-dimensional model, characterized in that, include: The acquisition module is used to acquire a 3D image of the target scene and partition the 3D image according to the field of view. The determination module is used to, for each divided visual region, filter background points and scene points based on the depth information of all voxel points, wherein the scene points include interactive points and non-interactive points; remove the depth information from all voxel points to obtain a two-dimensional image; determine the scene outline region based on the dividing line between the scene points and background points in the two-dimensional image; connect all interactive points in the scene outline region according to their coordinates to obtain an interactive indication path; determine the scene density based on the scene outline region and the two-dimensional image; determine the interactive indication density based on the scene outline region and the interactive indication path; and determine visual density features based on the scene density and the interactive indication density. The 3D optimization module is used to optimize 3D images based on the visual load criterion satisfied by the difference in visually dense features between adjacent regions. The interaction module is used to construct a virtual 3D model of the target scene based on the optimized 3D image, and render the virtual 3D model in the display window of the VR / AR device worn by the user. The user can slide the real-time virtual scene in the display window by operating the rotation element of the virtual or physical controller, and the user can trigger the interaction point indication in the display window by operating the button element of the virtual or physical controller.