LOD rendering effect adaptive scene presentation method and device based on large language model
The method leverages a large language model to dynamically allocate resources and control rendering based on user intent, addressing inefficiencies in existing three-dimensional scene rendering technologies by enhancing adaptability and user interaction.
Patent Information
- Application Number
- CN202510470032.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
AI Technical Summary
The existing three-dimensional scene rendering technology is difficult to dynamically adapt to user intentions, resulting in waste of computing resources and poor interactive experience, especially when users pay attention to specific areas, they cannot achieve fine-grained local detail control and smooth transitions.
Through a large language model combined with multimodal data, user instructions are obtained, focus areas are determined, and computing resources are dynamically allocated based on significance. Progressive rendering and transition control are used to achieve semantic-driven rendering effect.
It improves the dynamic understanding and interactive response accuracy of three-dimensional scene rendering, optimizes the utilization of computing resources, realizes the matching of user intentions and rendering effects, and improves rendering efficiency and visual continuity.
Smart Images

Figure CN120318400A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of computer graphics and natural language processing, and in particular, to a method and device for adaptively presenting a scene with LOD rendering effects based on a large language model. Background Art
[0002] In the prior art, the control of the level of detail in 3D scene rendering mainly relies on pre - set multi - level models and static optimization strategies. Traditional methods set model simplification rules at different distances or viewpoints manually. Although the computational load can be reduced, they lack dynamic adaptability. For example, in a user interaction scenario, when a user focuses on a specific area (such as "the head of the lion dance"), the globally unified level of detail (LOD) strategy will cause non - focus areas to still maintain a high polygon count, resulting in waste of computing resources. In addition, existing systems are difficult to understand the semantics of natural language instructions and cannot achieve fine - grained local detail control, leading to a rigid and inefficient interaction experience.
[0003] Although current LOD technologies have made progress in tasks such as object classification and segmentation, their combination with language models still faces significant bottlenecks. The disorder and high - dimensional features of data make it difficult to extract semantic information, and language models usually lack the spatial perception ability of three - dimensional geometric structures. For example, when a user gives an instruction such as "zoom in on the east - side window of the building", traditional methods need to rely on precise coordinate annotations or predefined object labels and cannot directly locate the target area through natural language. At the same time, the cross - modal alignment problem is particularly prominent: it is difficult to effectively map the geometric features and the semantic information of the language description, resulting in the model being difficult to generate rendering parameters that meet the user's intention.
[0004] Existing LOD optimization methods mostly focus on geometric - level simplification and ignore semantic - guided dynamic resource allocation. For example, the projection - based fusion method can combine 2D images and point - cloud data, but it is easy to lose local details when dealing with complex spatial relationships; while unsupervised domain adaptation techniques can reduce the dependence on data annotation, but it is difficult to capture the implicit attention preferences in user interaction. In addition, real - time rendering systems often face response delay problems. Especially when dealing with large - scale scenes, existing transition algorithms (such as linear interpolation) are prone to visual jumps, affecting the user experience.
[0005] The development of interactive multi-modal large models provides new ideas for the above problems, but their application in the field of 3D rendering is still in the exploratory stage. Although existing research has attempted to combine language models with LOD, their output is still limited to environmental assessment reports and fails to directly drive the dynamic parameter adjustment of the rendering engine. The core contradiction that has not been solved by existing technologies lies in: how to construct an end-to-end semantic-geometry mapping system so that the language model can not only parse the user's intention, but also generate LOD control instructions that conform to the characteristics of the 3D space and achieve smooth detail transitions. This technical gap severely restricts the interaction depth and rendering efficiency of application scenarios such as virtual reality and digital twins.
[0006] Regarding the problem of insufficient matching between the 3D scene rendering effect and the user's intention in existing technologies, no effective solution has been proposed yet. Summary of the Invention
[0007] The present invention provides a method and device for adaptively presenting a scene with LOD rendering effects based on a large language model to solve the defect of insufficient matching between the 3D scene rendering effect and the user's intention in existing technologies.
[0008] In the first aspect, the present invention provides a method for adaptively presenting a scene with LOD rendering effects based on a large language model, including: Obtaining multi-modal data of the target scene and annotating the target scene; Obtaining a language instruction, calling the trained large language model, semantically parsing the language instruction, and determining the focus area in the target scene; Dynamically allocating computing resources according to the saliency of the focus area; Based on the focus area, performing progressive rendering and transition control on the target scene to obtain a rendering result.
[0009] According to the method for adaptively presenting a scene with LOD rendering effects based on a large language model provided by the present invention, obtaining multi-modal data of the target scene and annotating the target scene includes: Collecting point cloud data of the target scene through a laser scanner and a panoramic camera; Recording the texture and semantic labels of the target scene; Formulating an annotation tool and mapping the natural language associated with the target scene to a 3D coordinate bounding box through the annotation tool.
[0010] According to the method for adaptively presenting a scene with LOD rendering effects based on a large language model provided by the present invention, training the large language model includes: Obtaining a dialogue data set containing a number of annotated samples; Encode the coordinate information of the target scene into a 768-dimensional vector, and train the large language model in combination with the dialogue dataset.
[0011] According to a method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention, perform semantic parsing on the language instruction to determine the focus area in the target scene, including: Through the large language model, fuse the semantic information of the language instruction with the spatial encoding features of the target scene, and align the semantic information and the coordinates of the target scene through a cross-attention mechanism; Based on the result of coordinate alignment, generate the regional coordinates and level-of-detail level instructions of the focus area.
[0012] According to a method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention, dynamically allocate computing resources according to the saliency of the focus area, including: Based on the attention guidance algorithm, determine the focus area budget according to the saliency of the focus area; Implement an exponential decay strategy for the non-focus area of the target scene to compress the number of patches in the non-focus area.
[0013] According to a method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention, perform progressive rendering and transition control on the target scene to obtain a rendering result, including: Deploy a double-buffer rendering architecture and use an S-shaped curve to control the detail transition of the target scene; Perform multi-threaded resource scheduling optimization on the target scene to obtain the rendering result.
[0014] According to a method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention, perform multi-threaded resource scheduling optimization on the target scene to obtain the rendering result, including: Load a high-precision normal map to the focus area; Perform surface subdivision on the target scene through a geometry shader, and perform frustum culling and instance merging processing on the non-focus area of the target scene to obtain the rendering result.
[0015] According to a method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention, the method further includes: Record the interactive event log, and annotate the data quality in the interactive event log through user satisfaction scoring; the interactive event log contains several historical user interactive rendering information; Perform hot update on the large language model in combination with the interactive event log.
[0016] In a second aspect, the present invention further provides a multi-level-of-detail dynamic rendering device, including: An acquisition module, configured to acquire multimodal data of a target scene and annotate the target scene; An analysis module, configured to obtain a language instruction, call a trained large language model, perform semantic analysis on the language instruction, and determine a focus area in the target scene; An allocation module, configured to dynamically allocate computing resources according to the saliency of the focus area; A rendering module, configured to perform progressive rendering and transition control on the target scene based on the focus area to obtain a rendering result.
[0017] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for adaptively presenting a scene with LOD rendering effect based on a large language model as described in the first aspect above.
[0018] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for adaptively presenting a scene with LOD rendering effect based on a large language model as described in the first aspect above.
[0019] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for adaptively presenting a scene with LOD rendering effect based on a large language model as described in the first aspect above.
[0020] Compared with the prior art, the present invention has the following beneficial effects: The method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention creatively establishes a semantic-driven parameter generation system, determines the focus area of concern according to the user's language instruction, breaks through the limitations of traditional rendering methods in dynamic scene understanding and interaction response accuracy, realizes a comprehensive upgrade from the underlying algorithm to the system architecture, and solves the problem of insufficient matching between the three-dimensional scene rendering effect and the user's intention in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a flowchart of the method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention; Figure 2 It is a schematic diagram of the process of fine-tuning a large language model in an embodiment of the present invention; Figure 3 It is a block diagram of the structure of the device for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention; Figure 4 It is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0024] The present invention provides a method for adaptively presenting a scene with LOD rendering effect based on a large language model, Figure 1 which is a flowchart of the method for adaptively presenting a scene with LOD rendering effect based on a large language model provided by the present invention. As Figure 1 shown, the method includes the following steps: Step S101, obtaining multimodal data of the target scene and annotating the target scene; Step S102, obtaining a language instruction, calling the trained large language model, performing semantic parsing on the language instruction, and determining the focus area in the target scene; Step S103, dynamically allocating computing resources according to the saliency of the focus area; Step S104, based on the focus area, performing progressive rendering and transition control on the target scene to obtain a rendering result.
[0025] In this method, first, multi-modal data of the target scene is obtained. The multi-modal data can clearly represent each region and detail of the target scene. Combining the multi-modal data, different regions and different detail points in the target scene are labeled. Then, the user's language instruction is obtained, and the semantic analysis of the language instruction is carried out through a large language model to analyze the user's intention and determine the focus area that the user needs to pay attention to. Then, the computing resources are reasonably allocated according to the saliency of the focus area to improve the efficiency of dynamic rendering and the utilization rate of computing resources. Finally, based on the focus area, progressive rendering and transition control are performed on the target scene to supplement and highlight the details of the focus area, and the rendering result is obtained. In the above process, a semantic-driven parameter generation system is creatively established. The focus area of concern is determined according to the user's language instruction, breaking through the limitations of traditional rendering methods in dynamic scene understanding and interaction response accuracy, realizing a comprehensive upgrade from the underlying algorithm to the system architecture, and solving the problem of insufficient matching between the three-dimensional scene rendering effect and the user's intention in the existing technology.
[0026] Particularly, in the application scenario of intangible cultural heritage display, the rendering process requirements for three-dimensional scenes are higher. This method is committed to creating the effect of realizing the digital display of intangible cultural heritage through natural language interaction, and completely retaining the core process characteristics such as the "three carvings and two sculptures" of Guangfu architecture. This method deeply integrates geographical information and cultural semantics, constructs a spatial perception model suitable for the complex structure of Lingnan architecture, and realizes smooth visual continuity in ultra-high definition display. It particularly optimizes the rendering of typical elements such as wok ear walls and lime sculptures and colored paintings. By establishing a closed-loop optimization mechanism for feedback from intangible cultural heritage inheritors, the recognition accuracy of Cantonese dialect instructions and regional cultural characteristics of the system is continuously improved, providing a high-fidelity and strong-interaction digital solution for cultural heritage in scenarios such as digital museums and virtual exhibitions. Next, this method will be applied to the Guangfu intangible cultural heritage scenario for elaboration: In one of the embodiments, step S101, obtaining multi-modal data of the target scene and annotating the target scene, includes: collecting point cloud data of the target scene through a laser scanner and a panoramic camera; recording the texture and semantic labels of the target scene; formulating an annotation tool, and mapping the natural language associated with the target scene to a three-dimensional coordinate bounding box through the annotation tool.
[0027] Exemplarily, first, a laser scanner (accuracy 0.1mm) and a panoramic camera are used to collect point cloud data of Guangfu architecture, and the texture and semantic labels (such as "lime sculpture and colored painting", "wok ear wall") are recorded synchronously. Then, a customized annotation tool is developed to map the natural language (such as "zoom in on the carving on the east side of the building") to a three-dimensional coordinate bounding box (x_min, y_min, z_min, x_max, y_max, z_max), and the annotation error is controlled within ±0.3m.
[0028] In some of these embodiments, Figure 2 is a schematic diagram of the process of fine-tuning a large language model in an embodiment of the present invention. As Figure 2 shown, training the large language model includes: obtaining a dialogue dataset containing a number of labeled samples; encoding the coordinate information of the target scene into a 768-dimensional vector, and training the large language model in combination with the dialogue dataset.
[0029] Specifically, the method first establishes a semantic-space conversion mechanism based on the large language model, and uses an improved Transformer architecture to fuse text instructions and three-dimensional coordinate encoding at the input layer. A target region heat map is generated through a Differentiable Region Proposal Network (DRPN), and natural language instructions such as "enhance the carved details of the eaves" are accurately mapped to spatial coordinates, while the topological relationship of building components is analyzed in combination with features. During the training process, a three-dimensional position embedding algorithm is used to convert the three-dimensional coordinate bounding box coordinates into 768-dimensional vectors, which are concatenated with text features and then input into a multi-layer cross-attention module, enabling the large language model to have the spatial perception ability to understand azimuth descriptions such as "the east side of the building", and realizing the deep integration of natural language interaction and LOD control.
[0030] Exemplarily, a dialogue dataset containing 2 million labeled samples is constructed, and the coordinate information is encoded into 768-dimensional vectors using a three-dimensional position embedding algorithm:
[0031] Among them, represents the position encoding that maps the three-dimensional coordinates (x, y, z) to 768-dimensional vectors through sine / cosine functions, k represents the key vector, represents the dimension of the key vector. Then, spatial perception ability is injected into the large language model through a LoRA adapter, freezing 95% of the original parameters and only training the newly added three-dimensional encoding layer.
[0032] On this basis, in step S102, semantic parsing of the language instruction is performed to determine the focus area in the target scene, including: through the large language model, fusing the semantic information of the language instruction with the spatial encoding features of the target scene, and aligning the semantic information and the coordinates of the target scene through a cross-attention mechanism; based on the result of the coordinate alignment, generating the regional coordinates and the level-of-detail hierarchy instruction of the focus area.
[0033] Exemplarily, the fine-tuned large language model is called, and text embedding and spatial encoding features are fused at the input layer, and semantic-coordinate alignment is achieved through a cross-attention mechanism:
[0034] Among them, represents the output calculated through the cross-attention mechanism, which is used to align text semantics with three-dimensional spatial features. For example, in a dialogue system, this output can fuse the text information of the user's instruction and the coordinate information of the scene. represents the text query matrix (Query Matrix), which is generated by linearly transforming the text embedding: . Among them is the word embedding of the text, is the learnable weight matrix. represents the three-dimensional spatial key matrix (Key Matrix), which is generated by the position encoding of the three-dimensional coordinates , is the three-dimensional position encoding (such as the 768-dimensional vector defined above), is the learnable weight matrix, represents the three-dimensional spatial value matrix (Value Matrix), which is also generated by the position encoding: , through weighted summation , the model can extract the spatial features related to the text, is the learnable weight matrix, represents the dimension of the key vector (usually the same as the query / value dimension), which is used to scale the dot product result: , for example, if 8-head attention is used, , the superscript T represents the matrix transpose. The output layer of the large language model configures the structured parameter head to generate a JSON instruction containing the coordinates of the focus area (such as X[215.4,216.8]) and the LOD level (LOD4 in the core area).
[0035] In some of these embodiments, in step S103, the computing resources are dynamically allocated according to the saliency of the focus area, including: based on the attention guidance algorithm, determining the focus area budget according to the saliency of the focus area; implementing an exponential decay strategy for the non-focus areas of the target scene to compress the number of patches in the non-focus areas.
[0036] In this embodiment, a spatial perception-based resource allocation algorithm (Aggregate Resource Allocation, AGRA) is designed to generate JSON instructions containing geometric and texture parameters through a structured output header. This algorithm is based on visual saliency calculation, adopts an adaptive normalization strategy to dynamically allocate patch budgets, preferentially allocates 80% of the computing resources to the focus area, and implements hierarchical attenuation control according to the grid bounding box overlap. The focus area is enhanced to 20,000 patches using the tessellation algorithm, the transition area is smoothly degraded according to an S-shaped curve, and the non-focus area compresses the number of patches to 30% of the baseline value through instancing merging technology. The resource monitoring module continuously detects the video memory occupancy and automatically triggers the texture degradation strategy when it exceeds the threshold to ensure that the key area maintains 8K ultra-high definition rendering.
[0037] Exemplarily, the resource allocation algorithm is executed to dynamically allocate computing resources according to the focus area saliency. The specific formula is as follows:
[0038] An exponential decay strategy is implemented for the non-focus area, compressing the number of patches to 30%-50% of the original scale and reducing the video memory occupancy by 42%.
[0039] To achieve visual continuity optimization, in some of these embodiments, in step S104, progressive rendering and transition control are performed on the target scene to obtain a rendering result, including: deploying a double-buffer rendering architecture and using an S-shaped curve to control the detail transition of the target scene; performing multi-threaded resource scheduling optimization on the target scene to obtain a rendering result.
[0040] Specifically, performing multi-threaded resource scheduling optimization on the target scene to obtain a rendering result includes: loading a high-precision normal map into the focus area; performing tessellation on the target scene through a geometry shader, and performing frustum culling and instancing merging processing on the non-focus area of the target scene to obtain a rendering result.
[0041] In this embodiment, a hybrid LOD transition engine is developed, which adopts a double-buffer architecture and a spatio-temporal consistency algorithm. While the main thread maintains the basic LOD control, the background thread preloads high-precision resources and dynamically generates supplementary geometric details through quadtree subdivision. The transition process is controlled by a parameterized S-shaped curve, and the transition factor k related to the viewing distance dynamically adjusts the interpolation rate. A 200ms fade window is set in the VR scene to eliminate visual jumps. For the intangible cultural heritage digital exhibition scene, the transition logic of complex components such as wok ear walls is specially optimized, and the temporal anti-aliasing technology is used to handle the detail mutations of carved patterns.
[0042] Exemplarily, first, a double-buffered rendering architecture is deployed. The main thread maintains the basic LOD control (number of patches: 5000), and the background thread preloads high-precision resources (number of patches: 20000). And a parametric S-shaped curve is used to control the detail transition. The specific formula is as follows:
[0043] where represents the number of patches at the detail level at time t, represents the number of patches of the basic LOD (fixed at 5000 patches), represents the number of patches of the target high-precision LOD (fixed at 20000 patches), k represents the transition factor, k = 8.0 (VR scene), represents the time variable (unit: second), usually counted from the start of the transition, represents the central time point of the transition process, that is, when t = t0, LOD(t) reaches the intermediate value between the basic value and the target value, and a visual smooth switch is completed within 0.8 seconds of the transition time.
[0044] Then, multi-threaded resource scheduling optimization is carried out. Specifically: First, load a high-precision normal map (2048×2048 resolution) to the focus area. Then, start the geometry shader for tessellation, and the tessellation level is 4. Finally, perform frustum culling and instance merging on the non-focus area to obtain the rendering result.
[0045] In addition, in order to improve the rendering accuracy of this method, the method also constructs a closed-loop optimization mechanism including user feedback, specifically including: recording the interaction event log, and annotating the data quality in the interaction event log through user satisfaction scores; the interaction event log contains several historical user interaction rendering information; combining the interaction event log to perform hot updates on the large language model.
[0046] In this embodiment, the interaction events and rendering parameters are recorded by the log analysis module to form an optimization data set containing more than 20 indicators such as viewpoint trajectory, LOD level, and GPU load. In each round of iteration, a contrastive learning strategy is used to update the network parameters, and the dialect feature extraction layer is optimized for Cantonese dialect instructions to improve the recognition accuracy of professional terms such as "lime sculpture and colored painting". During the deployment stage, the optimization effect is verified through A / B testing. After three rounds of iteration, the dialect instruction parsing accuracy is successfully increased from 78% to 86% in the digital museum scenario, and at the same time, the rendering latency is reduced by 42%.
[0047] Exemplarily, first, an interaction event log (including 20 parameters such as timestamp, viewpoint coordinates, LOD level, etc.) is recorded, and the data quality is labeled by the user satisfaction score (1 - 5 points). An incremental training set is generated every 24 hours, and the curriculum learning strategy is adopted to optimize low-scoring samples first. Then, the LoRA adapter is used for the hot update of the large language model, freezing 90% of the original parameters and only training the newly added semantic mapping layer. In the intangible cultural heritage display scenario, after 3 rounds of iteration, the recognition accuracy of Cantonese dialect instructions has increased from 88% to 93.7%.
[0048] In summary, this method focuses on solving the cross-modal fusion problem of three-dimensional coordinate encoding and semantic understanding, developing a dynamic resource allocation algorithm with independent intellectual property rights, and establishing a self-optimizing closed-loop feedback mechanism to provide a new generation of intelligent rendering solutions for the digital content interaction field. In a typical scenario, the "display of the details of the lion dance face" instruction can complete the entire process from semantic parsing to 8K texture loading within 300 ms, and the geometric subdivision error < 0.5 mm.
[0049] The present invention also provides a device for adaptively presenting the LOD rendering effect based on a large language model. The multi-level detail dynamic rendering device provided by the present invention will be described below. The multi-level detail dynamic rendering device described below can be correspondingly referred to the method for adaptively presenting the LOD rendering effect based on a large language model described above. Figure 3 is the structural block diagram of the device for adaptively presenting the LOD rendering effect based on a large language model provided by the present invention, as Figure 3 shown. The device includes: An acquisition module 301, configured to acquire multi-modal data of a target scene and label the target scene; An analysis module 302, configured to acquire a language instruction, call the trained large language model, perform semantic analysis on the language instruction, and determine the focus area in the target scene; An allocation module 303, configured to dynamically allocate computing resources according to the saliency of the focus area; A rendering module 304, configured to perform progressive rendering and transition control on the target scene based on the focus area to obtain a rendering result.
[0050] When this device is in use, first, the acquisition module 301 acquires the multimodal data of the target scene, and the multimodal data can clearly represent each area and detail of the target scene. Combining the multimodal data, different areas and different detail points in the target scene are labeled. Then, the parsing module 302 acquires the user's language instruction, and performs semantic parsing on the language instruction through a large language model, analyzes the user's intention, and determines the focus area that the user needs to pay attention to. The allocation module 303 then reasonably allocates computing resources according to the saliency of the focus area, improving the efficiency of dynamic rendering and the utilization rate of computing resources. Finally, the rendering module 304 performs progressive rendering and transition control on the target scene based on the focus area, supplements and highlights the details of the focus area, and obtains the rendering result. In the above process, a semantic-driven parameter generation system is creatively established, the focus area to be concerned is determined according to the user's language instruction, breaking through the limitations of traditional rendering methods in dynamic scene understanding and interaction response accuracy, realizing a comprehensive upgrade from the underlying algorithm to the system architecture, and solving the problem of insufficient matching between the three-dimensional scene rendering effect and the user's intention in the prior art.
[0051] Figure 4 Schematic diagram of the physical structure of an electronic device is illustrated, as Figure 4 shown. The electronic device may include: a processor 401, a communications interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communications interface 402, and the memory 403 communicate with each other through the communication bus 404. The processor 401 can call the logical instructions in the memory 403 to execute the LOD rendering effect adaptive scene rendering method based on a large language model. The method includes: Acquire the multimodal data of the target scene and label the target scene; Acquire the language instruction, call the trained large language model, perform semantic parsing on the language instruction, and determine the focus area in the target scene; Dynamically allocate computing resources according to the saliency of the focus area; Based on the focus area, perform progressive rendering and transition control on the target scene to obtain the rendering result.
[0052] In addition, when the logical instructions in the above-mentioned memory 403 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0053] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for adaptively presenting a scene with LOD rendering effects based on a large language model provided by the above-mentioned various methods. The method includes: Obtain multi-modal data of the target scene and annotate the target scene; Obtain a language instruction, call the trained large language model, perform semantic parsing on the language instruction, and determine the focus area in the target scene; Dynamically allocate computing resources according to the saliency of the focus area; Based on the focus area, perform progressive rendering and transition control on the target scene to obtain a rendering result.
[0054] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for adaptively presenting a scene with LOD rendering effects based on a large language model provided by the above-mentioned various methods. The method includes: Obtain multi-modal data of the target scene and annotate the target scene; Obtain a language instruction, call the trained large language model, perform semantic parsing on the language instruction, and determine the focus area in the target scene; Dynamically allocate computing resources according to the saliency of the focus area; Based on the focus area, perform progressive rendering and transition control on the target scene to obtain a rendering result.
[0055] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0056] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive scene rendering method for LOD rendering effects based on large language models, characterized in that, Including: Obtain multimodal data of the target scene and annotate the target scene; Obtain a language instruction, call the trained large language model, perform semantic parsing on the language instruction, and determine the focus area in the target scene; Dynamically allocate computing resources according to the saliency of the focus area; Based on the focus area, perform progressive rendering and transition control on the target scene to obtain a rendering result.
2. The method for adaptively presenting a scene with LOD rendering effect based on a large language model according to claim 1, wherein Obtain multimodal data of the target scene and annotate the target scene, including: Collect point cloud data of the target scene through a laser scanner and a panoramic camera; Record the texture and semantic labels of the target scene; Develop an annotation tool, and map the natural language associated with the target scene to a three-dimensional coordinate bounding box through the annotation tool.
3. The method for adaptively presenting a scene with LOD rendering effect based on a large language model according to claim 1, wherein Train the large language model, including: Obtain a dialogue dataset containing a number of annotated samples; Encode the coordinate information of the target scene into a 768-dimensional vector, and train the large language model in combination with the dialogue dataset.
4. The method for adaptively presenting a scene with LOD rendering effect based on a large language model according to claim 1, wherein Perform semantic parsing on the language instruction to determine the focus area in the target scene, including: Through the large language model, fuse the semantic information of the language instruction with the spatial encoding features of the target scene, and align the semantic information and the coordinates of the target scene through a cross-attention mechanism; Based on the result of coordinate alignment, generate the regional coordinates and level-of-detail hierarchy instructions of the focus area.
5. The method for adaptively presenting a scene with LOD rendering effect based on a large language model according to claim 1, wherein Dynamically allocate computing resources according to the saliency of the focus area, including: Based on the attention guidance algorithm, determine the focus area budget according to the saliency of the focus area; Implement an exponential decay strategy for the non-focus areas of the target scene to compress the number of patches in the non-focus areas.
6. The method for adaptively presenting a scene with LOD rendering effect based on a large language model according to claim 1, wherein Perform progressive rendering and transition control on the target scene to obtain a rendering result, including: Deploy a double-buffer rendering architecture and use an S-shaped curve to control the detail transition of the target scene; Perform multi-threaded resource scheduling optimization on the target scene to obtain the rendering result.
7. The method for adaptively presenting a scene with LOD rendering effect based on a large language model according to claim 6, wherein Perform multi-threaded resource scheduling optimization on the target scene to obtain the rendering result, including: Load a high-precision normal map to the focus area; Perform tessellation on the target scene through a geometry shader, and perform frustum culling and instance merging processing on the non-focus areas of the target scene to obtain the rendering result.
8. The method for adaptively presenting a scene with LOD rendering effect based on a large language model according to claim 1, wherein The method further includes: Record the interaction event log, and annotate the data quality in the interaction event log through user satisfaction scoring; the interaction event log contains a number of historical user interaction rendering information; Perform hot update on the large language model in combination with the interaction event log.
9. An LOD rendering effect adaptive scene presentation device based on a large language model, characterized in that, Including: An acquisition module for obtaining multimodal data of the target scene and annotating the target scene; An analysis module for obtaining a language instruction, calling the trained large language model, performing semantic parsing on the language instruction, and determining the focus area in the target scene; An allocation module for dynamically allocating computing resources according to the saliency of the focus area; A rendering module, configured to perform progressive rendering and transition control on the target scene based on the focus area, and obtain a rendering result.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for adaptively presenting a scene with LOD rendering effects based on a large language model according to any one of claims 1 to 8.
Citation Information
Cited By
AI-based three-dimensional model display method and system
CN120823327A
Multi-modal convergence media generation system and method based on dynamic emotion map
CN121030019A
Three-dimensional model compression transmission method and system based on dynamic feature perception
CN121284272A
Three-dimensional model compression transmission method and system based on dynamic feature perception
CN121284272B