A display system for designing an effects presentation

CN120723076BActive Publication Date: 2026-09-08JIAXING MAMMOTH IND DESIGN CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510949307.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2026-09-08
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

[0005]然而,以上述专利为代表的现有技术方案虽然在视觉呈现和成本控制上取得了一定的进步,但仍存在明显局限:首先,用户的体验多为被动观察,交互方式单一,缺乏对场景的主动控制和探索能力;其次,其沉浸感主要依赖于视觉和现场灯光的结合,忽略了声音、触觉等其他感官维度的融合,沉浸体验不够全面;最后,场景内容完全依赖于预设的设计模型库 ,无法根据用户的实时反馈或需求动态生成或调整,系统的智能化和自适应能力不足

Benefits of technology

[0034]By integrating VR vision, spatial audio, intelligent lighting synchronization, and finely detailed tactile feedback based on airflow, this system provides a highly realistic multi-sensory experience. In particular, the airflow tactile feedback unit can simulate the unique tactile feel of different materials, greatly enhancing the realism of the virtual environment and making the user feel as if they are in a real design space. Intelligent lighting synchronization ensures that the lighting effects in the virtual and real worlds are consistent, eliminating the discomfort caused by sensory discrepancies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723076B_ABST
    Figure CN120723076B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of scene display, and relates to a display system for design effect presentation, which comprises a live layer in communication with a cloud system layer, the live layer comprising a VR module and a plurality of sensory feedback devices, the system layer comprising a multimodal interaction analysis module, a customer state evaluation module, an AI experience generation engine and a multisensory rendering synchronization module, the multimodal interaction analysis module synchronously collecting and processing voice, gesture and gaze data of customers, the customer state evaluation module analyzing fixation point staying time and saccade path of customers, quantifying interest and understanding degree indexes of customers on current scenes in real time, the AI experience generation engine comprising a semantic parameter translation unit, a conditional content generation unit and a situational UI generation and layout unit, and the multisensory rendering synchronization module rendering scene parameters output by the AI experience generation engine into visual images, the present application can enhance the reality of a virtual environment, has strong interactivity, and makes users feel as if they are in a real design space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of scene display, and more specifically, to a display system for presenting design effects. Background Technology

[0002] In the field of modern design, whether it's industrial design, architectural design, or interior design, efficiently and accurately presenting the final effect of a design solution to clients or project stakeholders remains a crucial step. Traditional design presentation methods, such as two-dimensional renderings, physical sand table models, or animated walkthroughs with pre-set paths, while capable of showcasing design concepts to some extent, often suffer from limitations such as high communication costs, unintuitive information delivery, and a lack of immersion. Users often struggle to truly understand the scale, materials, and realistic feel of the space.

[0003] With the rapid development of computer graphics and virtual reality (VR) technologies, interactive and immersive 3D scene display systems have emerged, aiming to address the pain points of traditional methods. These systems construct virtual environments that are 1:1 scale with the real world, allowing users to "place" themselves in future design products, greatly improving the efficiency and depth of design communication.

[0004] Publication No. CN 114115523 B discloses an "Immersive Scene Display System Combining Static and Dynamic Elements." This solution proposes a two-tier architecture consisting of a "system layer" located in the cloud and a "site layer" located at the display site. Its core concept is to rapidly generate immersive effects on the cloud server through the visual algorithm unit and design model library unit of the system layer, and then send instructions to the site layer. The site layer includes not only a site VR module for displaying virtual reality projection scenes, but also a lighting control unit capable of receiving instructions and adjusting the physical lighting fixtures on site in real time, as well as a scene hardware management module. A key feature of this system is the use of a visual enhancement module to track the viewing angle of personnel within the site VR module and enhance the rendering of the visual focus area. This method aims to reduce reliance on the performance of on-site hardware, enabling enhanced rendering even on lower-performance devices, thereby effectively reducing the cost of on-site deployment.

[0005] However, while existing technologies, represented by the aforementioned patents, have made some progress in visual presentation and cost control, they still have significant limitations: First, the user experience is mostly passive observation, with a single interaction method and a lack of active control and exploration capabilities over the scene; second, the immersive experience relies primarily on the combination of vision and ambient lighting, neglecting the integration of other sensory dimensions such as sound and touch, resulting in an incomplete immersive experience; finally, the scene content is entirely dependent on a pre-set design model library, unable to be dynamically generated or adjusted based on real-time user feedback or needs, indicating insufficient intelligence and adaptability of the system. Therefore, developing a more interactive, multi-sensory integrated, and intelligent immersive display system is a pressing issue that needs to be addressed in the current technological field. Summary of the Invention

[0006] Therefore, the purpose of this invention is to provide a display system for presenting design effects, which can enhance the realism of the virtual environment, enhance interactivity, and make users feel as if they are in a real design space.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A display system for presenting design effects includes a live layer that communicates with a cloud-based system layer. The live layer includes a VR module and multiple sensory feedback devices. The system layer includes:

[0009] The multimodal interaction parsing module is configured to synchronously collect and process the customer's voice, gesture, and gaze data. It includes a dominant channel recognition unit to determine that voice is the primary command channel and uses gesture and gaze data as contextual references for deciphering pronouns in the commands.

[0010] The customer status assessment module is configured to quantify, in real time, a customer's interest in and understanding of the current scene by analyzing the duration of their gaze and their scanning path.

[0011] The AI ​​experience generation engine includes a semantic parameter translation unit, a conditional content generation unit, and a contextual UI generation and layout unit. The semantic parameter translation unit is used to translate the customer's vague verbal requirements into a set of executable scene parameter adjustment instructions.

[0012] The conditional content generation unit is used to call resources from the design knowledge base or create new design options through a generative model based on the customer's instructions and interest indicators.

[0013] The contextual UI generation and layout unit is used to create and dynamically lay out interactive information cards and data panels based on interest and comprehension metrics.

[0014] The multi-sensory rendering synchronization module is used to render the scene parameters output by the AI ​​experience generation engine into visual images and drive the sensory feedback devices of the scene layer to provide sound, light, and tactile feedback that are synchronized with visual time and space.

[0015] The present invention is further configured such that: the semantic parameter translation unit has a built-in dual encoder model, which is used to map subjective descriptive words and objective parameter sets in the same high-dimensional vector space;

[0016] When a descriptive word is received, it is encoded into a target semantic vector. By calculating the difference between the current scene parameter vector and the target semantic vector, a set of parameter adjustments is calculated and an instruction set is generated.

[0017] The present invention is further configured such that: the conditional content generation unit includes a stylized transfer model based on generative adversarial networks. This model can retain the core structural layout of the current space after receiving a client instruction, and only perform overall and consistent style replacement on the materials, colors, lighting atmosphere and decorative elements in the scene, and provide an interactive interface to support users to adjust style parameters for fine-tuning.

[0018] The present invention is further configured such that the customer status assessment module and the AI ​​experience generation engine are coupled and associated in the following manner:

[0019] When the customer status assessment module determines that the customer's understanding index is continuously below the threshold, the module sends an activation signal to the AI ​​experience generation engine to instantiate an AI intelligent guide.

[0020] When the customer status assessment module determines that the customer's interest in a certain object has significantly increased based on the duration and frequency of gaze and the direction of the gesture, the module sends another activation signal to the contextual UI generation and layout unit to instantiate an interactive information card about that object.

[0021] The present invention is further configured such that the processing flow of the AI ​​intelligent guide includes:

[0022] Context injection: The real-time context injector inputs the current viewpoint, the ID of the object being observed, and the recent interaction history from the VR scene into the prompts of the localized large language model in a structured manner;

[0023] Answer generation: The localized large language model, combined with the context, generates answers to the customer's natural language questions;

[0024] Non-conflict layout: The layout optimization algorithm is executed by the contextual UI generation and layout unit. Based on the real-time gaze heatmap from the customer state assessment module, the algorithm calculates the conflict value of multiple candidate display positions and selects the position with the smallest conflict value to present the answer in the form of augmented reality labels.

[0025] The present invention is further configured such that: the tactile feedback provided by the multi-sensory rendering synchronization module is implemented by an airflow tactile feedback unit, which is connected to a material tactile knowledge base. When the user's gesture touches the virtual object, the unit calls the corresponding airflow parameters from the knowledge base according to the material ID of the object and drives a miniature airflow jet device to simulate the tactile sensation.

[0026] The present invention is further configured such that: the system layer includes a design knowledge base stored in a graph database structure, wherein the database nodes include 3D models, materials, commercial attributes and tactile parameters, and the edge relationships of the design knowledge base are defined to describe the adaptability, compatibility and spatial logic between the nodes.

[0027] The present invention is further configured such that: the multi-sensory rendering synchronization module further includes a lighting intelligent synchronization unit, which includes a light dynamic calculation module and a signal controller.

[0028] The dynamic lighting calculation module combines pre-calculated light maps and dynamic lighting models, and optimizes rendering based on the gaze point heatmap from the customer status assessment module, to calculate the real-time lighting effect of the VR scene; the signal controller then sends the calculated parameters to the adjustable LED array lights on site.

[0029] The present invention is further configured such that: the system layer supports an immersive side-by-side comparison mode, which is executed by the contextual UI generation and layout unit after receiving a specific instruction, the execution action including:

[0030] Generate a split-screen rendering view in the VR field of view to render two different design schemes simultaneously;

[0031] A differentiated data panel is generated on one side of the view, and the differences between the two solutions are compared and displayed in real time in the form of charts by querying the design knowledge base.

[0032] The present invention is further configured such that: the system layer includes a multi-user state synchronization engine, used to manage and synchronize the position, posture and interaction events of the virtual avatars of remote users in a cross-regional collaborative experience mode, and has a built-in spatial audio communication module, which adjusts the voice volume and channels in real time according to the distance and orientation between avatars.

[0033] Compared with the shortcomings of the prior art, the beneficial effects of the present invention are as follows:

[0034] By integrating VR vision, spatial audio, intelligent lighting synchronization, and finely detailed tactile feedback based on airflow, this system provides a highly realistic multi-sensory experience. In particular, the airflow tactile feedback unit can simulate the unique tactile feel of different materials, greatly enhancing the realism of the virtual environment and making the user feel as if they are in a real design space. Intelligent lighting synchronization ensures that the lighting effects in the virtual and real worlds are consistent, eliminating the discomfort caused by sensory discrepancies.

[0035] The multimodal interaction parsing module can process the user's voice, gestures, and gaze simultaneously. In particular, it uses gestures and gaze as contextual references for voice commands, effectively eliminating ambiguity in natural language. This allows users to interact with the virtual environment in a more natural and intuitive way, without having to memorize complex commands or menu operations, thus significantly improving interaction efficiency. Attached Figure Description

[0036] Figure 1 This is a system framework diagram of the present invention. Detailed Implementation

[0037] Reference Figure 1 The present invention provides a further description of an embodiment of a display system for presenting design effects, the system comprising an on-site layer and a system layer on the cloud.

[0038] The on-site layer is the part that users directly interact with, and its core function is to provide an immersive experience and collect user data. It includes a VR module that provides the core visual immersion. Commercially available VR headsets can be used; these devices typically integrate high-definition displays, high refresh rates, and precise head and hand tracking systems, capable of capturing the user's head posture, position, and hand movements in real time.

[0039] Multi-sensory feedback devices can enhance immersion by providing feedback in addition to visual and auditory feedback.

[0040] Haptic feedback devices are primarily used to simulate the tactile sensation of virtual objects.

[0041] The auditory feedback device is provided by high-quality spatial audio headphones built into the VR headset, which can simulate the source and direction of sound propagation to enhance the sense of realism.

[0042] The cloud system layer is responsible for handling complex data computation, AI inference, content generation, and data storage.

[0043] The system layer is deployed on a high-performance cloud server cluster, using GPU-accelerated computing instances to meet the needs of real-time processing of large amounts of data and running complex AI models. Communication with the field layer is achieved through high-speed, low-latency network connections (such as Wi-Fi 6 or wired Ethernet).

[0044] The system layer includes a multimodal interaction parsing module, a customer status assessment module, an AI experience generation engine, a multisensory rendering synchronization module, a design knowledge base, a multi-user status synchronization engine, and supports an immersive parallel comparison mode.

[0045] The multimodal interaction parsing module is responsible for synchronously collecting, integrating, and parsing various modal data from the user. It collects the user's voice data through the VR headset's built-in microphone; it collects the user's gesture data (e.g., grasping, pointing, swiping) through the VR headset and hand trackers; and it collects the user's gaze data (gaze coordinates, saccade path) through the VR headset's eye trackers. These data streams are timestamped and then fed into this module for preprocessing, such as speech noise reduction, gesture skeleton extraction, and eye-tracking data smoothing.

[0046] The dominant channel identification unit analyzes the characteristics of different modal data to determine which modality currently carries the user's core command intent.

[0047] In most cases, user commands are expressed through speech. This unit prioritizes recognizing and parsing speech content. For example, when it detects a user uttering a clear command phrase (such as "Change this wall to blue" or "Show a modern style"), the system determines that speech is the current dominant channel. This can be achieved using a deep learning-based speech recognition model (such as a Transformer or RNN-T model), which is trained to recognize the features of command language.

[0048] Gesture and gaze data are used as contextual references for deciphering pronouns in voice commands. When a voice command contains pronouns (such as "this", "it", "there"), the dominant channel recognition unit uses synchronously acquired gesture and gaze data to determine the specific object to which the pronoun refers.

[0049] If a user points to an object in the virtual scene while speaking, the system will prioritize using the object the gesture is pointing to as the reference for the pronoun. This is achieved by analyzing hand tracking data and calculating the intersection of the hand pointing vector and the bounding box of the object in the scene.

[0050] If the user does not explicitly point with a gesture, the system analyzes the synchronized eye-tracking data. The object the user's gaze lingers on for the longest time or the object traversed by the saccade is likely the object referred to by the pronoun. This can be achieved by calculating the cumulative time the gaze lingers on different objects or by analyzing the degree of overlap between the saccade path and the object's bounding box.

[0051] When gestures and eye contact are present simultaneously, the system can assign different weights or priorities. For example, gestures can have higher priority than eye contact, or when both gestures and eye contact point to the same object, the system can further enhance confirmation of that object's selection. If gestures and eye contact point to different objects, the system can further analyze the user's tone of voice, level of hesitation, and other characteristics to try to determine the more likely intent, or issue a confirmation prompt to the user when uncertain.

[0052] The customer status assessment module is responsible for real-time monitoring and understanding of users' psychological state during immersive experiences, particularly their interest in and comprehension of the current scene. It utilizes eye-tracking data from the VR headset to analyze customer gaze duration and saccade paths.

[0053] Record the duration of a user's gaze on different objects or areas within a virtual scene. Longer gaze durations typically indicate that the user is interested in the object or is observing it closely. A threshold can be set to determine "longer gaze duration." For example, a continuous gaze lasting more than 2 seconds might be considered a signal of interest.

[0054] Analyzing the trajectory of a user's eye movement reveals that frequent scanning may indicate that the user is searching for information or is confused about the current content, while smooth scanning paths may indicate that the user is browsing purposefully.

[0055] Interest metrics are quantified based on the cumulative time a gaze point spends on a specific object or area, the frequency of gaze, and whether it is accompanied by other positive signals (such as smiling, nodding, etc., which can be obtained through facial expression recognition technology, if the VR headset supports it).

[0056] Comprehension metrics are based on the duration of gaze on key information areas (such as labels or data charts), saccade patterns (e.g., whether the main part of the information is covered), and whether the user displays confused facial expressions or speech (which can be obtained through speech emotion analysis or facial expression recognition). Lower comprehension may manifest as frequent saccades, short dwell time on key information areas, or expressions of confusion.

[0057] These metrics are updated in real time, typically calculated multiple times per second, so that the system can respond quickly to changes in the user's status.

[0058] The AI ​​experience generation engine is responsible for generating personalized design content and interactive interfaces based on user instructions and status.

[0059] The semantic parameter translation unit transforms users' vague, colloquial needs into precise, executable system instructions. It translates these vague, verbal requests into a set of actionable scenario parameters. Users often use highly subjective and imprecise language to describe their design preferences, such as "I want a cozy living room" or "Make this room look more modern." This unit is responsible for understanding these vague descriptions and mapping them to specific, adjustable scenario parameters, such as wall color (warm tones), furniture material (fabric), and lighting brightness and color temperature (low brightness, warm color temperature).

[0060] The built-in dual-encoder model consists of two independent encoders: one for encoding subjective descriptive words (e.g., "warm" or "modern"), and the other for encoding objective parameter sets (e.g., RGB color values, material type IDs, and light brightness values). The two encoders are jointly trained so that in the model's high-dimensional vector space, the descriptive word vectors and parameter set vectors related to the same design intent are close to each other.

[0061] The dual-encoder model uses the Transformer as its foundation, building a descriptor encoder and a parameter set encoder on top of it. The descriptor encoder encodes the input text sequence into a fixed-dimensional vector.

[0062] The conditional content generation unit creates or calls new design content based on user instructions and assessed interest levels.

[0063] This unit accesses resources from a design knowledge base, which stores a vast collection of 3D models, materials, textures, lighting presets, and other design elements. It translates the unit's output instructions (e.g., "replace the sofa with a modern style sofa") based on semantic parameters and searches the knowledge base for matching resources (e.g., searching for 3D models tagged "sofa" and "modern"). Search results are sorted by relevance and presented to the user for selection, or directly loaded into the scene.

[0064] For design elements that cannot be directly obtained from the knowledge base, or for designs that require high personalization, the unit can call upon the generative model to create new design options.

[0065] Stylization transfer models based on generative adversarial networks can make overall and consistent changes to the visual style of a scene while preserving the original scene's spatial layout and object structure.

[0066] Style transfer models such as CycleGAN, StyleGAN2, or diffusion-based models can be used. These models learn image features of different styles (e.g., modern, classical, industrial) and apply them to the input image. In this embodiment, the input is a rendered image of the VR scene or a scene description (containing geometry, material, and lighting information), and the output is a new scene rendered image or updated scene parameters with the target style applied. To preserve the core structural layout, a structural similarity loss function can be introduced during model training, or scene depth information and semantic segmentation information can be used as conditional inputs. The model receives the current scene image / description and the target style description as input. The style encoder extracts features of the target style. The content encoder extracts content features of the current scene (geometry, object positions, etc.). The style transfer module integrates the style features with the content features to generate a new feature representation. The decoder restores the new feature representation to the rendered image or updated scene parameters. The discriminator is used to distinguish the generated image from the real target style image, driving the generator to learn and generate more realistic stylized results.

[0067] Stylized transfer models typically have adjustable parameters, such as style intensity and preference for specific materials. Contextual UI generation and layout units can generate an adjustment panel in the user's view based on the generated results, allowing the user to fine-tune these parameters using sliders or buttons, observe the effect in real time, and continue until satisfied. For example, the user can adjust the intensity of the "metallic feel" or the proportion of "exposed brick walls" in an industrial style.

[0068] The contextual UI generation and layout unit intelligently creates and places interactive interface elements based on the user's state and the context, providing information or allowing user interaction. It also creates and dynamically lays out interactive information cards and data panels based on interest and comprehension metrics.

[0069] When the customer status assessment module determines that a user's interest in a specific object (such as a lamp or a piece of furniture) has significantly increased, this unit generates an information card near that object. This card can display detailed information about the object, such as brand, model, material, price, and size. Users can interact with the card through gaze or gestures, such as clicking to view more images, adding it to their favorites, or obtaining a purchase link. The card's placement takes into account the user's gaze point to avoid obscuring the content the user is currently focusing on. When the customer status assessment module determines that the user's comprehension index is consistently below a threshold, or when the user actively requests more information, this unit generates data panels. These panels can display key data related to the current scene in the form of charts, lists, or text, such as space dimensions, lighting analysis, and material cost estimates. The panel layout considers the user's gaze point heatmap, placing it in a position that is easily noticed by the user but does not interfere with the core visual experience. This unit has a built-in layout optimization algorithm, which takes into account the current user's gaze point heatmap (from the customer status assessment module) and the UI elements to be placed (information cards, data panels). The algorithm aims to find the optimal layout positions where UI elements are easily noticed by the user without obscuring the area the user is currently most focused on, while also avoiding overlap between different UI elements. This is achieved by calculating the "conflict value" between each candidate position and the gaze point heatmap; a higher conflict value indicates that the position is more likely to obscure the area the user is focused on. The algorithm selects the position with the minimum conflict value, and this algorithm is based on reinforcement learning to find the optimal layout.

[0070] The multi-sensory rendering synchronization module is responsible for transforming the abstract scene parameters output by the AI ​​experience generation engine into a multi-sensory experience—visual images—that the user can perceive. It ensures synchronization between different senses. Based on the updated scene parameters output by the AI ​​engine (such as object position, rotation, scaling, materials, and lighting settings), it uses a high-performance rendering engine for real-time rendering, generating stereoscopic visual images corresponding to the left and right eyes, and displaying them through a VR headset. A high frame rate (e.g., 72Hz or 90Hz) is required to reduce motion blur and dizziness. It also drives the sensory feedback devices in the scene layer, providing sound, light, and tactile feedback synchronized with visual spatiotemporal perception.

[0071] Based on the position, material, and spatial structure of objects in the scene, the propagation and reflection effects of sound are simulated to generate spatial audio. For example, the sound is louder closer to the sound source, and the echo is more obvious in an enclosed space, which is achieved by the spatial audio system built into the VR headset.

[0072] The intelligent lighting synchronization unit synchronizes the lighting changes in the virtual scene with the adjustable LED lights in the real environment; the airflow tactile feedback unit provides corresponding airflow tactile simulation when the user's hand touches the virtual object, ensuring that the changes in sound, light, tactile feedback and visual images are highly synchronized in time.

[0073] The coupling between the customer status assessment module and the AI ​​experience generation engine: When the customer status assessment module determines that the customer's understanding index is continuously below the threshold, the module sends an activation signal to the AI ​​experience generation engine to instantiate an AI intelligent guide.

[0074] The customer status assessment module continuously calculates the user's comprehension metric. A threshold is set (e.g., 0.4). If the user's comprehension metric remains below this threshold within a set time window (e.g., 30 seconds), the system determines that the user may be experiencing difficulties or is confused about the current content.

[0075] The customer status assessment module generates an activation signal that includes the current scene ID, the user's viewpoint, and the IDs of possible objects / areas that could lead to reduced comprehension, and sends it to the AI ​​experience generation engine.

[0076] After receiving the activation signal, the AI ​​experience generation engine initiates the AI ​​intelligent guide's processing flow. The intelligent guide will attempt to understand the user's points of confusion and proactively provide assistance. When the customer status assessment module determines that the customer's interest in a certain object has significantly increased based on the duration and frequency of gaze and gestures, the module sends another activation signal to the contextual UI generation and layout unit to instantiate an interactive information card about that object. The customer status assessment module continuously calculates the user's interest index. When the user's gaze duration on a specific object exceeds a set threshold (e.g., 3 seconds), the gaze frequency increases, or there is a gesture pointing at the object, the system determines that the user has a strong interest in that object. The customer status assessment module generates an activation signal containing the ID of the object being observed and sends it to the contextual UI generation and layout unit.

[0077] After receiving the activation signal, the contextual UI generation and layout unit retrieves relevant information from the design knowledge base based on the object ID, and generates and displays the object's information card according to the aforementioned layout algorithm.

[0078] AI-powered intelligent guidance is a proactive assistance provided by the system when users encounter difficulties or require additional information. Context injection is implemented by a real-time context injector, a data processing unit responsible for collecting and structuring real-time information related to the user's current state and the scene. It inputs the current viewpoint from the VR scene, the ID of the observed object, and the recent interaction history into the prompts of the localized large language model.

[0079] Current viewpoint, including the user's head position and orientation.

[0080] The ID of the object being watched, the ID of the object the user is currently watching.

[0081] Recent interaction history is the user's most recent actions, such as adjusting wall colors, moving furniture, and issuing voice commands.

[0082] Structured input involves organizing this information into a format that the large language model can understand. This structured information is embedded in the prompts sent to the large language model as contextual information, helping the model better understand the user's intent and current situation.

[0083] Answer generation: The localized large language model, combined with context, generates answers to customers' natural language questions. If the user's low comprehension is due to a lack of understanding of a certain design concept or system operation, the model will generate explanatory or guiding answers based on the injected context information.

[0084] If a user's low comprehension is due to unfamiliarity with the system's functions, the model will provide operational guidance.

[0085] The model may also proactively provide relevant design suggestions based on users' historical interactions and preferences. Non-conflict layouts are generated by contextual UI and executed by layout units.

[0086] The layout optimization algorithm, based on a real-time gaze heatmap provided by the customer status assessment module, calculates the conflict value of multiple candidate display positions. The algorithm generates multiple potential display positions for the AI-powered intelligent guide's answer text or prompt window. For each position, it calculates the degree of overlap or distance with the user's current high-attention area (high-density area on the gaze heatmap). Higher overlap and closer distance result in a higher conflict value. The position that minimizes interference with the user's current visual focus is selected to display the intelligent guide's information. The intelligent guide's answer or prompt does not directly cover the entire screen but appears as a semi-transparent, floating text box or label within the scene, always facing the user.

[0087] The multi-sensory rendering synchronization module provides haptic feedback, which is implemented by an airflow haptic feedback unit. This unit mainly consists of a miniature air pump or fan connected to a flexible duct, and one or more miniature nozzles located at the user's fingers or palm. When haptic sensation needs to be simulated, the air pump activates, spraying airflow onto the user's skin through the duct. This connection is linked to a material haptic knowledge base, which stores airflow parameters corresponding to different materials.

[0088] When a user touches a virtual object, the VR system's hand tracking module detects that the user's virtual hand collides with an object in the virtual scene. The unit then retrieves the corresponding airflow parameters from the knowledge base based on the object's material ID and drives a miniature airflow jet device to simulate the tactile sensation.

[0089] The system layer includes a design knowledge base stored in a graph database structure. This base stores all design elements used to construct the virtual scene and their interrelationships. The graph database structure efficiently expresses complex object relationships and attributes, storing data in the form of nodes and edges. Nodes represent entities, and edges represent relationships between entities.

[0090] The database nodes include 3D models, materials, commercial attributes, and tactile parameters.

[0091] The 3D model stores the geometric shape and appearance data of various furniture, building components, decorations, plants, etc. The attributes of the nodes can include model ID, name, category (e.g., sofa, chair, wall), style tag, size, number of polygons, etc.

[0092] Materials store various physical properties of materials, such as texture maps, colors, reflectivity, transparency, and normal maps. Node attributes can include material ID, name, type (e.g., wood, metal, cloth, glass), style tags, and physical property parameters.

[0093] Business attributes are related to business information associated with design elements. Node attributes can include brand, model, price, supplier, place of manufacture, inventory status, etc.

[0094] The tactile parameters store the airflow tactile feedback parameters corresponding to the material. The node's attributes can include parameters such as airflow intensity, frequency, and temperature associated with the material ID.

[0095] Edges are used to describe the logical relationships, adaptability, compatibility, etc. between nodes.

[0096] The multi-sensory rendering synchronization module also includes a lighting intelligent synchronization unit, which synchronizes lighting changes in the virtual scene to adjustable LED lights in the real environment, further enhancing the sense of immersion.

[0097] The intelligent lighting synchronization unit comprises a dynamic lighting calculation module and a signal controller. The dynamic lighting calculation module combines pre-calculated lightmaps and dynamic lighting models. VR scene rendering typically combines pre-calculation (e.g., baked lightmaps to simulate static lighting effects such as global illumination and ambient occlusion) and dynamic calculation (to handle real-time changing lighting, such as moving light sources and material specular highlights). The dynamic lighting calculation module acquires the real-time lighting results of the virtual scene calculated by the rendering engine and optimizes the rendering based on the gaze heatmap from the client state assessment module. The user's gaze heatmap reflects the area currently being focused on by the user. Based on the heatmap, the dynamic lighting calculation module can prioritize or improve the lighting accuracy of the user's focused areas, thereby optimizing computing resources and improving rendering efficiency. It calculates the real-time lighting effects of the VR scene and outputs real-time lighting parameters for various areas in the virtual scene, such as light intensity, color, direction, and shadow information.

[0098] The signal controller sends the calculated parameters to the adjustable LED array lights on site: The signal controller converts the real-time illumination parameters output by the dynamic lighting calculation module into control signals, which are then sent to the adjustable LED array lights on site via DMX512, Art-Net, or proprietary wireless protocols.

[0099] The LED lights installed in the on-site environment have independent control capabilities and adjustable color and brightness functions. The position and orientation of these lights need to correspond to the light sources in the virtual scene or the areas that need to be synchronously illuminated.

[0100] As the sun sets in the virtual scene, the light color changes to a warmer tone and the brightness decreases. The LED lights on-site will also change synchronously to produce the same lighting effect, further enhancing the user's sense of immersion in the virtual environment. If a specific light fixture in the virtual scene is turned on or off, the corresponding LED light on-site will also change accordingly.

[0101] The signal controller needs to have enough channels to control the LED lights on site and support low-latency communication protocols.

[0102] The lighting calculation module needs to provide sufficiently high illumination accuracy while meeting real-time requirements. Different calculation algorithms can be selected based on hardware performance.

[0103] The system supports an immersive side-by-side comparison mode, allowing users to compare two different design options simultaneously for better decision-making. This mode is executed by the contextual UI generation and layout unit upon receiving specific instructions: users can activate this mode via voice commands (e.g., "Compare these two options") or gestures / UI operations, and the execution actions include:

[0104] A split-screen rendering view is generated within the VR field of view to simultaneously render two different design schemes: logically or visually dividing the VR headset's display area into two parts. Scheme A is displayed on the left, and Scheme B on the right. The two schemes can be different modifications to the same base scene, or they can be completely different scenes. The rendering engine needs to render both scenes simultaneously and composite the results into a single image. This requires high rendering performance.

[0105] A differentiated data panel is generated on one side of the view, and by querying the design knowledge base, the differences between the two solutions are displayed in real time in the form of charts. The contextual UI generation and layout unit generates a data panel next to the split-screen view (for example, between two solutions or in a floating window on one side). The system queries the design knowledge base to obtain the key attributes and parameters of solution A and solution B, and compares the differences between the two solutions in key indicators.

[0106] The system layer includes a multi-user state synchronization engine, which is responsible for managing and synchronizing the position, posture, and interaction events of the virtual avatars of each user participating in the collaborative experience.

[0107] Each user's VR headset and hand tracker data (position, rotation, gestures) is sent to the cloud. The multi-user state synchronization engine receives this data and copies and updates it on other users' local systems, thus displaying each other's virtual avatars and their actions in the VR scene of all participants. Interaction events (e.g., user A moves an object, user B changes the wall color) are also synchronized to ensure that all users see the same version of the scene state.

[0108] A built-in spatial audio communication module allows for natural voice communication between collaborating users. After the user's voice data is collected, it is sent to the cloud. The spatial audio communication module calculates the attenuation and spatialization effect of the voice in real time based on the relative position and orientation of the virtual avatars of the sender and receiver in the scene.

[0109] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any ordinary changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included within the protection scope of the present invention.

Claims

1. A display system for presenting design effects, comprising a live layer communicating with a cloud system layer, the live layer including a VR module and multiple sensory feedback devices, characterized in that, The system layer includes: The multimodal interaction parsing module is configured to synchronously collect and process the customer's voice, gesture, and gaze data. It includes a dominant channel recognition unit to determine that voice is the primary command channel and uses gesture and gaze data as contextual references for deciphering pronouns in the commands. The customer status assessment module is configured to quantify, in real time, a customer's interest in and understanding of the current scene by analyzing the duration of their gaze and their scanning path. The AI ​​experience generation engine includes a semantic parameter translation unit, a conditional content generation unit, and a contextual UI generation and layout unit. The semantic parameter translation unit is used to translate the customer's vague verbal requirements into a set of executable scene parameter adjustment instructions. The conditional content generation unit is used to call resources in the design knowledge base or create new design options through a generative model based on the customer's instructions and interest indicators. The contextual UI generation and layout unit is used to create and dynamically lay out interactive information cards and data panels based on interest and comprehension metrics. The multi-sensory rendering synchronization module is used to render the scene parameters output by the AI ​​experience generation engine into visual images and drive the sensory feedback devices of the scene layer to provide sound, light, and tactile feedback synchronized with visual time and space. The tactile feedback provided by the multi-sensory rendering synchronization module is implemented by an airflow tactile feedback unit. This unit is connected to a material tactile knowledge base. When the user's gesture touches a virtual object, the unit retrieves the corresponding airflow parameters from the knowledge base based on the object's material ID and drives a miniature airflow jet device to simulate tactile sensation. The multi-sensory rendering synchronization module also includes a lighting intelligent synchronization unit, which comprises a light dynamic calculation module and a signal controller. The dynamic lighting calculation module combines pre-calculated light maps and dynamic lighting models, and optimizes rendering based on the gaze point heatmap from the customer status assessment module, to calculate the real-time lighting effect of the VR scene; the signal controller then sends the calculated parameters to the adjustable LED array lights on site.

2. The display system for presenting design effects according to claim 1, characterized in that, The semantic parameter translation unit has a built-in dual encoder model, which is used to map subjective descriptive words and objective parameter sets in the same high-dimensional vector space; When a descriptive word is received, it is encoded into a target semantic vector. By calculating the difference between the current scene parameter vector and the target semantic vector, a set of parameter adjustments is calculated and an instruction set is generated.

3. A display system for presenting design effects according to claim 2, characterized in that, The conditional content generation unit includes a stylized transfer model based on generative adversarial networks. After receiving a client instruction, the model can retain the core structural layout of the current space and only perform a holistic and consistent style replacement on the materials, colors, lighting atmosphere, and decorative elements in the scene. It also provides an interactive interface to support users in fine-tuning the style parameters.

4. A display system for presenting design effects according to claim 3, characterized in that, The customer status assessment module and the AI ​​experience generation engine are coupled and linked in the following ways: When the customer status assessment module determines that the customer's understanding index is continuously below the threshold, the module sends an activation signal to the AI ​​experience generation engine to instantiate an AI intelligent guide. When the customer status assessment module determines that the customer's interest in a certain object has significantly increased based on the duration and frequency of gaze and the direction of the gesture, the module sends another activation signal to the contextual UI generation and layout unit to instantiate an interactive information card about that object.

5. A display system for presenting design effects according to claim 4, characterized in that, The processing flow of the AI ​​intelligent guide includes: Context injection: The real-time context injector inputs the current viewpoint, the ID of the object being observed, and the recent interaction history from the VR scene into the prompts of the localized large language model in a structured manner; Answer generation: The localized large language model, combined with context, generates answers to customers' natural language questions; Non-conflict layout: The layout optimization algorithm is executed by the contextual UI generation and layout unit. Based on the real-time gaze heatmap from the customer state assessment module, the algorithm calculates the conflict value of multiple candidate display positions and selects the position with the smallest conflict value to present the answer in the form of augmented reality labels.

6. A display system for presenting design effects according to claim 1, characterized in that, The system layer includes a design knowledge base stored in a graph database structure. Its database nodes include 3D models, materials, commercial attributes, and tactile parameters. The edge relationships of the design knowledge base are defined to describe the adaptability, compatibility, and spatial logic between the nodes.

7. A display system for presenting design effects according to claim 6, characterized in that, The system layer supports an immersive side-by-side contrast mode, which is executed by the contextual UI generation and layout unit after receiving specific instructions. These actions include: Generate a split-screen rendering view in the VR field of view to render two different design schemes simultaneously; A differentiated data panel is generated on one side of the view, and the differences between the two solutions are compared and displayed in real time in the form of charts by querying the design knowledge base.

8. A display system for presenting design effects according to claim 7, characterized in that, The system layer includes a multi-user state synchronization engine, which manages and synchronizes the location, posture, and interaction events of the virtual avatars of remote users in a cross-regional collaborative experience mode. It also includes a built-in spatial audio communication module, which adjusts the voice volume and channels in real time based on the distance and orientation between avatars.

Citation Information

Patent Citations

  • An immersive scene display system combining movement and stillness

    CN114115523B

  • Exhibition display system applying virtual reality technology

    CN117784929A

  • Intelligent home decoration system and method based on mixed reality

    CN119848974A

  • Method and system for the recognition of reading skimming and scanning from eye-gaze patterns

    US6873314B1