An intelligent user guidance method for digital twin projects
Patent Information
- Application Number
- CN202511883530.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-12-15
AI Technical Summary
其主要技术缺陷在于:耦合度高且难以复用:在引导逻辑嵌入业务代码中时,每开发新项目需重新编写引导代码,无法实现跨项目的模块化部署
[0043]1.极大简化开发流程,将引导系统开发化繁为简。本发明的核心在于提供了一套标准化的软件框架,它将引导系统的开发工作从复杂的“编程任务”转变为简单的“配置任务”。开发者不再需要为每个项目编写和维护繁琐的引导逻辑代码,而是通过对UI元素进行简单的“标记”,并填写一份标准化的JSON“清单”即可。这种配置驱动的模式,让原本需要数周编码和调试的工作,缩短为数小时的配置和导入,从而将开发者从重复性的工作中解放出来,极大降低了研发成本与人力投入。同时,该框架作为成熟资产,可在不同项目中即插即用,实现高度复用,也方便进行各类修改。
Smart Images

Figure CN121742927B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and more specifically, to an intelligent user guidance method for digital twin projects. Background Technology
[0002] Digital twin and industrial simulation software typically features high-fidelity modeling, real-time data-driven operation, and complex business logic simulation. The user interface of this type of software often contains numerous function panels, parameter controls, and multi-level interactive logic, resulting in high operational complexity. Furthermore, this type of software is highly dependent on industry knowledge; the terminology, data indicators, and operational procedures in the interface are often closely related to specific processes or equipment principles. For non-domain experts, simply mastering the interface operation steps without the background knowledge makes it difficult to correctly understand the software's functions, often requiring frequent interruptions to consult external resources, leading to low human-computer interaction efficiency.
[0003] Existing user guidance technologies for this type of software mainly include the following: First, offline documents or video tutorials. These solutions separate the information carrier from the actual software operating environment and lack interactivity. Users need to frequently switch between the user interface and the documentation, unable to obtain immediate feedback on specific difficulties encountered during the current operation. Furthermore, when software version iterations lead to changes in the user interface layout, static materials need to be recreated, resulting in high maintenance costs.
[0004] Second, there's hard-coded in-program bootstrapping. This approach typically binds the bootstrapping logic to the user interface's hierarchical path during the development phase. Its main technical drawbacks are: high coupling and difficulty in reuse: when bootstrapping logic is embedded in business code, the bootstrapping code must be rewritten for each new project, making cross-project modular deployment impossible. Rigid logic: it can only provide pre-defined one-way operation guidance, unable to respond to user's independent exploration behavior, and unable to handle dynamically generated interface elements at runtime. Lack of semantic understanding: it cannot understand users' natural language questions and cannot establish semantic connections between operation instructions and industry knowledge.
[0005] Third, it relies on manual training. This method is limited by labor costs and time and space constraints, and cannot meet users' immediate needs for assistance.
[0006] In summary, existing technologies for solving user onboarding problems in complex digital twin applications suffer from several drawbacks, including a disconnect between onboarding information and the operating environment, high coupling between onboarding logic and business code, and a lack of intelligent question-answering capabilities based on industry context. This invention aims to address these issues by providing a low-code, reusable, and context-aware intelligent user onboarding method. Summary of the Invention
[0007] To address at least one of the aforementioned problems, the present invention first provides an intelligent user guidance method for digital twin projects, the intelligent user guidance method comprising the following:
[0008] A front-end system is established, which is based on a 3D graphics engine and consists of a plug-in system that goes from static configuration to dynamic mapping. A real-time index is created by combining externally defined industry knowledge with dynamically generated interface elements at runtime using unique identifiers. The functional attributes, guidance logic, and associated industry knowledge of the user interface are abstracted and persistently stored as static configuration files. The abstract data in these static configuration files is mapped to specific object instances in the 3D engine's rendering pipeline. An explicit identification mechanism is set up at the component system level of the 3D engine. Simultaneously, to balance the human-like interactive experience with lightweight system operation, a two-dimensional virtual proxy image is constructed at the user interface level as a unified medium for user-system interaction. This two-dimensional virtual proxy in screen space serves as the interaction entry point, executing structured linear guidance.
[0009] A backend system is established, and the deployment of the backend system is based on the intelligent service generated by retrieval enhancement. When a user initiates a natural language question, the frontend system captures the context label of the current interaction focus. Based on this, the backend retrieves professional documents within a limited scope and generates a business-aware answer by combining it with a large language model, thus realizing an interactive closed loop from operation navigation to professional Q&A.
[0010] Optionally, the metadata definition method of the static configuration file involves defining a data graph to describe the bootstrap node. The data exchange format file internally maintains a linear array structure, where each element is an independent data object representing a specific, bootstrap user interface functional unit in the application. Each data object contains the following field definitions:
[0011] A unique identifier field, which is used to uniquely identify a user interface element globally;
[0012] The bootstrap sequence index field is used to define the time execution order of the functional units in the deterministic linear bootstrap process. The logic control unit in the backend system will determine the triggering order of the bootstrap events based on the arithmetic ascending order of the time execution order values.
[0013] The voice broadcast content field stores the text information that the anthropomorphic agent needs to output to the user when executing structured guidance. The text content will be directly passed to the text-to-speech engine for audio waveform synthesis.
[0014] The knowledge base index tag set field, which serves as the core metadata of the backend system, stores a set of domain-specific terms that are semantically strongly related to the functions of the user interface. These terms will be used as filtering conditions for vector retrieval to limit the knowledge search scope of the large language model in the backend system.
[0015] The context system prompt field stores a natural language description used to define the business boundaries, preconditions, or potential impacts of the user interface function at the semantic level. The content of the context system prompt field will be injected into the dialogue context as a system-level instruction when building the large language model prompt project, so as to establish the business background of the dialogue.
[0016] Optionally, the component system identifier and data association mechanism of the 3D engine at runtime is as follows:
[0017] The data of the identification component is encapsulated, and a custom script component class is developed. The script component class inherits from the base component class of the graphics engine and can be mounted on scene objects. The script component class declares and serializes a public string variable. The name of the variable is consistent with the unique identifier field in the static configuration file. During the editing and building phase of the application, the string variable is manually filled to establish the logical correspondence between the scene objects and specific data entries in the static configuration file.
[0018] The global registry is constructed in memory. A global singleton registry manager is built in the runtime memory of the application. The registry manager maintains an efficient key-value pair container based on a hash algorithm. The key type of the efficient key-value pair container is constrained to be a string, corresponding to the unique identifier field. The value type of the efficient key-value pair container is constrained to the general object reference type of the graphics engine.
[0019] The automated registration process manages the lifecycle using the lifecycle callback interface provided by the graphics engine. The automatic registration logic is triggered the instant the identifier component is instantiated and enters memory. This logic reads the identifier string stored internally within the component instance and writes it as a key-value pair with the memory address reference of the object to which the component belongs. If the same key already exists in the hash container, an overwrite update operation is performed. This ensures that regardless of whether the user interface object is a static resource loaded with the scene or a dynamic resource generated by code logic, as long as the identifier component is mounted and initialized by the graphics engine, it will be automatically included in the index scope of the global registry without additional external intervention.
[0020] Optionally, the visual overlay and layout method of the virtual agent image is as follows:
[0021] An independent visual display area is established at the top of the application's display interface. Within this area, a two-dimensional digital character image with a transparent background is rendered. This image is fixedly positioned in the non-operational core area of the screen display to ensure that continuous, supportive guidance is provided without obscuring the underlying business operation interface or data display content.
[0022] Optionally, a multimodal interaction and performance mechanism for the two-dimensional virtual agent image is established:
[0023] The multimodal interaction includes voice interaction capability. The two-dimensional virtual agent image is equipped with a complete voice input and output interface, which can receive the user's voice commands and perform natural language broadcast through speech synthesis technology.
[0024] The performance mechanism includes facial expressions and actions. The system maintains a state mapping table for the two-dimensional virtual agent image. The state mapping table can dynamically update the visual presentation effect of the digital human according to the current interaction state. The visual presentation effect includes switching between different emotional portrait images, playing preset basic two-dimensional animation sequences, and mouth opening and closing actions that match the rhythm of voice broadcasting, thereby achieving a vivid and natural anthropomorphic interaction effect.
[0025] Optionally, the structured linear guidance of the front-end system relies on local logic control, guiding the user through the key functions of the software according to a predefined logical order, as follows:
[0026] When the bootstrapping process is activated, the logic controller reads the global configuration data. The front-end system arranges all relevant boot nodes in ascending order according to the preset boot order index in the data, generates a linear task execution queue, and prepares to start execution from the first step.
[0027] The execution steps involve the front-end system sequentially reading the current guidance node in the task execution queue. First, it performs target localization and highlighting, which involves reading the unique identifier field in the current guidance node, finding the corresponding user interface object position in the scene, and generating or moving a highlighted indicator box at the position to visually mark the operation area that needs attention. Then, it performs voice explanation, which involves extracting the voice broadcast content in the current guidance node, driving the two-dimensional virtual agent image to read the text through voice synthesis technology, and coordinating with corresponding body or lip movements to complete the explanation.
[0028] In the process flow control, after the current step's highlighted mark and voice explanation are completed, the front-end system automatically pauses the process and displays a "Next" interactive button. When the user clicks the button, the front-end system removes the current highlighted mark, moves the execution pointer to the next node in the queue, and repeats the above execution steps until all content in the queue has been guided.
[0029] Optionally, a context-aware intelligent question-answering system is established between the front-end system and the back-end system. The context-aware intelligent question-answering system realizes intelligent interaction between the user's natural language input and the large language model in the back-end system. The context-aware intelligent question-answering system constructs a communication carrier containing precise business context and performs knowledge retrieval based on the context tags in the back-end system. The specific method is as follows: first, the interaction focus of the front-end system is captured and the context is extracted; then, the back-end system generates retrieval enhancement logic based on tag filtering; and finally, closed-loop feedback and display are performed.
[0030] Optionally, when a user triggers the voice question function, the front-end system starts a recording device to capture an audio stream and calls an automatic speech recognition interface to convert the captured audio stream into a text string. Then, the context-aware intelligent question-answering system executes key context capture logic:
[0031] The ray-based focus detection algorithm uses the input module of the graphics engine to obtain the screen coordinates of the current mouse pointer, emits a three-dimensional ray from the camera position to the screen coordinates, and uses the ray projection interface of the physics engine or user interface event system to detect the collision between the ray and the user interface objects in the scene. Then, it returns the first valid interactive object that the ray intersects as the current focus object.
[0032] In the reverse lookup of the identifier component, the context-aware intelligent question answering system attempts to obtain the reference of the custom identifier component on the focused object. If the acquisition fails, the current context is marked as empty, and only the user's question text is retained. If the acquisition is successful, the unique identifier stored in the custom identifier component is read. Then, the unique identifier is used to perform a reverse lookup in the static configuration file to locate the corresponding configuration data object.
[0033] The communication payload is structured and encapsulated. The context-aware intelligent question-answering system extracts the content of the knowledge base index tag set field and the context system prompt word field from the located configuration data object, and constructs a POST request packet conforming to the Hypertext Transfer Protocol standard. The data body of the request packet is serialized into a JSON format string, containing the following key-value pairs: query: the user's original question text; context_id: the unique identifier of the focused object; context_tags: the extracted tag array; system_prompt: the extracted business description text; the request packet is sent to the service interface of the backend system through a network socket.
[0034] Optionally, the method for generating enhanced retrieval logic based on tag filtering in the backend system is as follows:
[0035] The backend system maintains a vector database for conditional filtering retrieval. This database stores pre-processed and quantized industry technical document slices. Each document slice is tagged with specific metadata tags upon entry into the database. The retrieval module first reads the context_tags array from the request packet, constructs a filtering expression, and instructs the vector database to search only in the document set whose metadata tags are contained within the context_tags array. Subsequently, the retrieval module converts the user's query text into a high-dimensional vector, calculates the cosine similarity in the aforementioned limited document set, and recalls the top-N document slices with the highest similarity scores.
[0036] The prompt word engineering involves dynamic assembly. The backend proxy module constructs the final prompt words sent to the large language model based on the search results. The structure of the prompt words strictly follows the template below:
[0037] At the system instruction layer, the system_prompt content in the request packet is injected to force the model to enter a specific business role setting.
[0038] Referring to the knowledge layer, the plain text content of all recalled document slices is concatenated as factual evidence;
[0039] In the user command layer, inject the user's query text;
[0040] Generative reasoning and response: The assembled prompt words are submitted to the large language model reasoning interface. The large language model generates response text based on the limited context environment. The backend system encapsulates the text into an HTTP response packet and returns it to the frontend system.
[0041] Optionally, the closed-loop feedback and display method is as follows: after the front-end network module receives the response text, it distributes it to the virtual proxy display module. The front-end system calls the text-to-speech interface to play the answer audio, and instantiates a text bubble control on one side of the virtual proxy to display the answer text in a visual typewriter effect, thereby completing a complete, context-aware intelligent question-and-answer interaction closed loop.
[0042] Compared to existing technologies, the intelligent user guidance method for digital twin projects in this invention has the following advantages:
[0043] 1. Significantly simplifies the development process, reducing the complexity of bootstrapping system development to a simple task. The core of this invention lies in providing a standardized software framework that transforms bootstrapping system development from a complex "programming task" into a simple "configuration task." Developers no longer need to write and maintain cumbersome bootstrapping logic code for each project; instead, they simply "mark" UI elements and fill in a standardized JSON "manifestation." This configuration-driven model reduces what used to require weeks of coding and debugging to hours of configuration and import, freeing developers from repetitive tasks and significantly reducing R&D costs and manpower investment. Furthermore, as a mature asset, this framework can be plugged and played across different projects, achieving high reusability and facilitating various modifications.
[0044] 2. Breaking down industry knowledge barriers and deeply integrating operation guidance with knowledge transfer. One of the core innovations of this invention lies in its hierarchical knowledge data structure, which strongly binds UI elements with deterministic "operation instructions" and unstructured "domain knowledge." This allows the digital human guide to not only tell users "where to click," but also, through linkage with the backend RAG system, answer users' deeper professional questions in real time and accurately, such as "what is this?" and "why is this value?" It acts like a domain expert embedded within the software, effectively solving the operational interruptions and comprehension difficulties caused by users' lack of professional background knowledge, ensuring the continuity and smoothness of the workflow.
[0045] 3. Significantly improves user experience, providing an adaptive and interactive learning path. This invention integrates structured guidance and intelligent question-and-answer modes, giving users a high degree of freedom. Users can systematically learn software functions by following preset processes, or interrupt the process at any point to ask questions in natural language about points of confusion on the current interface. This flexibility and adaptability can meet the personalized needs of users with different knowledge levels and usage habits, transforming passive "cramming" teaching into active "exploratory" learning, thereby greatly enhancing users' learning interest and acceptance.
[0046] To ensure the accuracy and reliability of guidance information and meet the stringent requirements of industrial applications, this invention addresses the potential "fact illusion" problem inherent in general-purpose large language models. By employing a retrieval-enhanced generation architecture and strictly limiting its knowledge source to a professional knowledge base provided by the developer, the accuracy and reliability of all question-and-answer content are guaranteed. This is crucial for serious application scenarios such as industrial digital twins, where data and operational error tolerance is extremely low, ensuring that the guidance information received by users is professional and trustworthy. Attached Figure Description
[0047] Figure 1 A flowchart for an intelligent user onboarding method for digital twin projects;
[0048] Figure 2 A general framework diagram for intelligent user guidance methods for digital twin projects;
[0049] Figure 3 A framework diagram defining the metadata for static configuration files;
[0050] Figure 4 A framework diagram of the runtime component identification and data association mechanism for static configuration files;
[0051] Figure 5 A framework diagram of the multimodal interaction and presentation mechanism for static configuration files;
[0052] Figure 6 A framework diagram for capturing front-end interactive focus and extracting context for static configuration files;
[0053] Figure 7 This is a framework diagram of the tag-based filtering retrieval enhancement logic generated for the backend of a static configuration file. Detailed Implementation
[0054] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0055] This invention provides an intelligent user guidance method for digital twin projects, the intelligent user guidance method comprising the following:
[0056] like Figure 1 and Figure 2As shown, a front-end system is established. This front-end system is based on a 3D graphics engine and consists of a plug-in system that combines static configuration with dynamic mapping. A real-time index is established by combining externally defined industry knowledge with dynamically generated interface elements at runtime through a unique identifier. The functional attributes, guidance logic, and associated industry knowledge of the user interface are abstracted and persistently stored as a static configuration file. The abstract data in the static configuration file is mapped to specific object instances in the 3D engine rendering pipeline. An explicit identification mechanism is set in the component system layer of the 3D engine. At the same time, in order to balance the anthropomorphic interactive experience with the lightweight operation of the system, a two-dimensional virtual proxy image is built at the user interface layer as a unified medium for user interaction with the system. The two-dimensional virtual proxy in screen space is used as the interaction entry point to execute structured linear guidance.
[0057] A backend system is established, and the deployment of the backend system is based on the intelligent service generated by retrieval enhancement. When a user initiates a natural language question, the frontend system captures the context label of the current interaction focus. Based on this, the backend retrieves professional documents within a limited scope and generates a business-aware answer by combining it with a large language model, thus realizing an interactive closed loop from operation navigation to professional Q&A.
[0058] like Figure 3 As shown, optionally, the metadata definition method of the static configuration file involves defining a data graph (physically represented as a lightweight data exchange JSON file) to describe the bootstrap node. The data exchange JSON file internally maintains a linear array structure, where each element is an independent data object representing a specific, bootstrap user interface functional unit in the application. Each data object contains the following field definitions:
[0059] A unique identifier field (data type string) is used to uniquely identify a user interface element globally;
[0060] The bootstrap sequence index field (data type is long integer or integer) is used to define the time execution order of the functional unit in the deterministic linear bootstrap process. The logic control unit in the backend system will determine the triggering order of the bootstrap events based on the arithmetic ascending order of the time execution order values.
[0061] The voice broadcast content field (data type is long string) stores the text information that the humanoid agent needs to output to the user when executing structured guidance. The text content will be directly passed to the text-to-speech engine for audio waveform synthesis.
[0062] The knowledge base index tag set field (data type is string array or delimited string) serves as the core metadata of the backend system. It stores a set of domain-specific terms that are semantically strongly related to the user interface functions. These terms will be used as filtering conditions for vector retrieval to limit the knowledge search scope of the large language model in the backend system.
[0063] The context system prompt word field (data type is long string) stores a natural language description used to define the business boundaries, preconditions or potential impacts of the user interface function at the semantic level. The content of the context system prompt word field will be injected into the dialogue context as a system-level instruction when building the large language model prompt project to establish the business background of the dialogue.
[0064] like Figure 4 As shown, optionally, the component system identifier and data association mechanism of the 3D engine at runtime is as follows:
[0065] The data of the identification component is encapsulated, and a custom script component class is developed. The script component class inherits from the base component class of the graphics engine and can be mounted on scene objects. The script component class declares and serializes a public string variable. The name of the variable is consistent with the unique identifier field in the static configuration file. During the editing and building phase of the application, the string variable is manually filled to establish the logical correspondence between the scene objects and specific data entries in the static configuration file.
[0066] The global registry is constructed in memory. A global singleton registry manager is built in the runtime memory of the application. The registry manager maintains an efficient key-value pair container based on a hash algorithm. The key type of the efficient key-value pair container is constrained to be a string, corresponding to the unique identifier field. The value type of the efficient key-value pair container is constrained to the general object reference type of the graphics engine.
[0067] The automated registration process manages the lifecycle using the lifecycle callback interfaces (including initialization and activation callbacks) provided by the graphics engine. The automatic registration logic is triggered the instant the identifier component is instantiated and enters memory. This logic reads the identifier string stored internally within the component instance and writes it as a key-value pair with the memory address reference of the object to which the component belongs. If the same key already exists in the hash container, an overwrite update operation is performed. This ensures that regardless of whether the user interface object is a static resource loaded with the scene or a dynamic resource generated by code logic, as long as the identifier component is mounted and initialized by the graphics engine, it will be automatically included in the global registry index without additional external intervention.
[0068] Optionally, the visual overlay and layout method of the virtual agent image is as follows:
[0069] A separate visual display area is established at the top of the application's display interface. Within this area, a two-dimensional digital character image with a transparent background is rendered. This image is fixedly positioned in a non-core operational area of the screen (such as a corner) to ensure that continuous, supportive guidance is provided without obscuring the underlying business operation interface or data display content.
[0070] like Figure 5 As shown, optionally, a multimodal interaction and performance mechanism for the two-dimensional virtual agent image can be established:
[0071] The multimodal interaction includes voice interaction capability. The two-dimensional virtual agent image is equipped with a complete voice input and output interface, which can receive the user's voice commands and perform natural language broadcast through speech synthesis technology.
[0072] The performance mechanism includes facial expressions and actions. The system maintains a state mapping table for the two-dimensional virtual agent image. The state mapping table can dynamically update the visual presentation effect of the digital human according to the current interaction state (such as "idle standby", "listening", "explaining"). The visual presentation effect includes switching between different emotional portrait images (such as normal, confused, happy), or playing preset basic two-dimensional animation sequences (such as blinking, head movement, gesture indication), as well as lip opening and closing movements that match the rhythm of voice broadcasting, thereby achieving a vivid and natural anthropomorphic interaction effect.
[0073] Optionally, the structured linear guidance of the front-end system relies on local logic control, guiding the user through the key functions of the software according to a predefined logical order, as follows:
[0074] When the bootstrapping process is activated, the logic controller reads the global configuration data. The front-end system arranges all relevant boot nodes in ascending order according to the preset boot order index in the data, generates a linear task execution queue, and prepares to start execution from the first step.
[0075] The execution steps involve the front-end system sequentially reading the current guidance node in the task execution queue. First, it performs target localization and highlighting, which involves reading the unique identifier field in the current guidance node, finding the corresponding user interface object position in the scene, and generating or moving a highlight indicator box (such as a flashing circle or border) at the position to visually mark the operation area that needs attention. Then, it performs voice explanation, which involves extracting the voice broadcast content in the current guidance node, driving the two-dimensional virtual agent image to read the text through voice synthesis technology, and coordinating with corresponding body or lip movements to complete the explanation.
[0076] In the process flow control, after the current step's highlighted mark and voice explanation are completed, the front-end system automatically pauses the process and displays a "Next" interactive button. When the user clicks the button, the front-end system removes the current highlighted mark, moves the execution pointer to the next node in the queue, and repeats the above execution steps until all content in the queue has been guided.
[0077] Optionally, a context-aware intelligent question-answering system is established between the front-end system and the back-end system. The context-aware intelligent question-answering system realizes intelligent interaction between the user's natural language input and the large language model in the back-end system. The context-aware intelligent question-answering system constructs a communication carrier containing accurate business context and performs knowledge retrieval based on the context tags in the back-end system. The specific method is as follows: first, the interaction focus of the front-end system is captured and the context is extracted; then, the back-end system generates retrieval enhancement logic based on tag filtering; and finally, closed-loop feedback and display are performed.
[0078] like Figure 6 As shown, optionally, when a user triggers the voice question function, the front-end system starts a recording device to capture an audio stream and calls an automatic speech recognition interface to convert the captured audio stream into a text string. Then, the context-aware intelligent question-answering system executes key context capture logic:
[0079] The ray-based focus detection algorithm uses the input module of the graphics engine to obtain the screen coordinates of the current mouse pointer, emits a three-dimensional ray from the camera position to the screen coordinates, and uses the raycast interface (RaycastAll) of the physics engine or the user interface event system to detect the collision between the ray and the user interface objects in the scene. Then, it returns the first valid interactive object that the ray intersects as the current focus object.
[0080] In the reverse lookup of the identifier component, the context-aware intelligent question answering system attempts to obtain the reference of the custom identifier component on the focused object. If the acquisition fails, the current context is marked as empty, and only the user's question text is retained. If the acquisition is successful, the unique identifier stored in the custom identifier component is read. Then, the unique identifier is used to perform a reverse lookup in the static configuration file to locate the corresponding configuration data object.
[0081] The communication payload is structured and encapsulated. The context-aware intelligent question-answering system extracts the content of the knowledge base index tag set field and the context system prompt word field from the located configuration data object, and constructs a POST request packet conforming to the Hypertext Transfer Protocol (HTTP) standard. The data body of the request packet is serialized into a JSON format string, containing the following key-value pairs: query: the user's original question text; context_id: the unique identifier of the focused object; context_tags: the extracted tag array; system_prompt: the extracted business description text; the request packet is sent to the service interface of the backend system through a network socket.
[0082] like Figure 7 As shown, optionally, the method of the backend system based on tag filtering and retrieval enhancement generation (RAG) logic is as follows:
[0083] The backend system maintains a vector database for conditional filtering retrieval. This database stores pre-processed and quantized industry technical document slices. Each document slice is tagged with specific metadata tags upon entry into the database. The retrieval module first reads the context_tags array from the request packet, constructs a filtering expression, and instructs the vector database to search only in the document set whose metadata tags are contained within the context_tags array. Subsequently, the retrieval module converts the user's query text into a high-dimensional vector, calculates the cosine similarity in the aforementioned limited document set, and recalls the top-N document slices with the highest similarity scores.
[0084] The prompt word engineering involves dynamic assembly. The backend agent module constructs the final prompt words (Prompt) sent to the large language model based on the search results. The structure of the prompt words strictly follows the following template:
[0085] At the system instruction layer, the system_prompt content in the request packet is injected to force the model to enter a specific business role setting.
[0086] Referring to the knowledge layer, the plain text content of all recalled document slices is concatenated as factual evidence;
[0087] In the user command layer, inject the user's query text;
[0088] Generative reasoning and response: The assembled prompt words are submitted to the large language model reasoning interface. The large language model generates response text based on the limited context environment. The backend system encapsulates the text into an HTTP response packet and returns it to the frontend system.
[0089] Optionally, the closed-loop feedback and display method is as follows: after the front-end network module receives the response text, it distributes it to the virtual proxy display module. The front-end system calls the text-to-speech interface to play the answer audio, and instantiates a text bubble control on one side of the virtual proxy to display the answer text in a visual typewriter effect, thereby completing a complete, context-aware intelligent question-and-answer interaction closed loop.
[0090] In this embodiment:
[0091] Taking Unity3D as an example, if the technical solution of this invention is applied to a digital twin project of a "central monitoring system for intelligent factory reactors," the aim is to assist operators in mastering the use of core temperature control and safety devices, and to provide real-time decision support based on safety standards. The specific implementation details and construction process are described below:
[0092] 1. Construction of the backend industry knowledge base and preparation for targeted retrieval
[0093] Before the front-end application was released, the implementers first built an industry-specific knowledge base on the back-end server side, which is the foundation for realizing Search Enhancement Generation (RAG).
[0094] Unstructured data ingestion: The backend system first imports two core technical documents: "Maintenance Manual for High-Pressure Equipment in Chemical Industrial Park" and "ESD Emergency Shutdown System Operating Procedure (SOP)".
[0095] Semantic Slicing and Metadata Injection: The preprocessing script of the backend system semantically segments the above document into paragraphs. In particular, the section on "Emergency Stop Button Operation Specifications" is segmented into a separate text block, and the content clearly states: "The ESD system can only be activated when the core temperature of the reactor exceeds the critical value (850°C) and the backup cooling pump is confirmed to have failed."
[0096] Tag Mapping (Key Step): In the database management system, technicians explicitly tag the text block with a set of business metadata tags, such as: ["Chemical Safety", "ESD System", "Emergency Shutdown"]. This set of tags is designed to strictly correspond to the definitions in the front-end system configuration file.
[0097] Vectorized storage: The text blocks and their tags are converted into high-dimensional vectors and stored in a vector database, building a dedicated index library that supports "tag-based filtering".
[0098] 2. Configuration of front-end static bootstrap data
[0099] Developers create a configuration file named guide_data.json in the StreamingAssets directory of their Unity project. For key interaction points in the scene, define the following data structures:
[0100] {
[0101] "guide_nodes": [
[0102] {
[0103] "ui_id": "Btn_Emergency_Stop_01",
[0104] "sequence_index": 5,
[0105] "speech_content": "This is the emergency stop button, also known as the ESD system switch. Please note that this operation is irreversible; clicking it will forcibly disconnect the power to all feed pumps and open the pressure relief valve."
[0106] "rag_tags": ["Chemical Safety", "ESD System", "Emergency Shutdown"],
[0107] "system_prompt": "The current user is concerned with the plant's highest priority safety device—the emergency shutdown system. The operation of this device involves significant safety responsibilities. Responses must be extremely rigorous, prioritizing the red-line clauses of the safe operating procedures, and ambiguous suggestions are strictly prohibited."
[0108] },
[0109] {
[0110] "ui_id": "Panel_Temp_Monitor_Core",
[0111] "sequence_index": 1,
[0112] "speech_content": "This is the core temperature monitoring area, which displays the average temperature data collected by thermocouples inside the reactor in real time."
[0113] "rag_tags": ["thermocouple", "PID control", "temperature monitoring"],
[0114] "system_prompt": "The current user is viewing core temperature data, a key indicator for assessing the reaction progress."
[0115] } ]
[0117] }
[0118] 3. Implementation of the scene object identification and registration mechanism
[0119] In the Unity editor environment, developers need to configure the UI elements of the control room scene as components.
[0120] Component development: Write the UIGuideIdentity script, which inherits from MonoBehaviour.
[0121] Static binding: Select the red "Emergency Stop Button" GameObject in the scene and attach the UIGuideIdentity script. In the Inspector panel, manually fill in the UniqueID field exposed by the script as "Btn_Emergency_Stop_01" to ensure that it is completely consistent with the JSON configuration.
[0122] Dynamic registration logic: This script includes registration logic in the Awake() function. It automatically calls the singleton manager when the program starts.
[0123] GuidanceManager.Instance.Register("Btn_Emergency_Stop_01",this.gameObject).
[0124] For the "real-time alarm scrolling list" dynamically generated by code on the right side of the scene, the developer attaches the script to the prefab of the list item and writes code to dynamically assign the ID (such as "Alarm_Item_" + deviceID) during instantiation, thereby ensuring that the dynamically generated UI elements can also be indexed by the system.
[0125] 4. Presentation layer integration of two-dimensional virtual proxies
[0126] To achieve lightweight and intuitive human-like interaction, this embodiment constructs an independent two-dimensional visual digital human interaction layer in the Unity scene. The specific integration process is as follows:
[0127] Visual hierarchy construction: Create an independent UI canvas object in the scene, configure its rendering mode to "Screen Space - Overlay", and set it to the highest sorting level to ensure that the virtual agent always appears above all business interfaces. Instantiate an image component in a non-operational area of the canvas (such as the lower right corner of the screen), load the static portrait resource of the virtual engineer, and use it as the main visual element for user interaction.
[0128] Audio-visual synchronization mechanism: To give static character portraits a dynamic "speaking" effect, the system integrates an audio-driven visual logic. 1) State definition: Two image resources are prepared in advance: "closed mouth (default state)" and "open mouth (speaking state)". 2) Driving logic: The system monitors the audio source component (AudioSource) mounted on the virtual proxy in real time. In each frame's logic update, the system obtains the real-time volume intensity of the currently playing audio. 3) Dynamic switching: A volume threshold is set. When the audio is playing and the real-time volume intensity exceeds this threshold, the system automatically replaces the image resource of the image component with the "open mouth" state; conversely, when the volume is below the threshold or playback stops, it reverts to the "closed mouth" state. This mechanism achieves a humanoid lip-syncing effect that matches the speech rhythm without complex skeletal animation.
[0129] End-to-end interactive integration: This virtual agent module is registered as the system's output terminal. When the backend intelligent service returns the response text, the system automatically triggers the following processes: 1) Speech synthesis: The TTS interface is called to convert the text into an audio stream and assign it to the audio source component. 2) Visual feedback: While playing the audio, the above audio-visual synchronization driving logic is activated, causing the virtual agent to start "speaking". 3) Text display: A dialog bubble UI control is instantiated simultaneously next to the virtual agent, presenting the response content in a word-by-word typewriter effect.
[0130] 5. Runtime demonstration of the structured bootstrapping process
[0131] When a new operator initiates the "Onboarding Guide" mode, the front-end system executes the following logic:
[0132] Queue parsing: The system reads the JSON file and sorts the nodes according to the sequence_index.
[0133] Step 1: The system identifies the ID corresponding to sequence 1 as Panel_Temp_Monitor_Core. The system then searches for the UI object corresponding to this ID in the global registry and obtains its screen coordinates (RectTransform).
[0134] Highlighting and Explanation: The system instantiates a highlighted prefab with a yellow warning border, precisely selecting the temperature monitoring panel. Simultaneously, the virtual agent announces via TTS: "This is the core temperature monitoring area..." After the user clicks "Next," the system automatically proceeds to the next step.
[0135] 6. The complete closed loop of context-aware intelligent question answering (core scenario)
[0136] In actual monitoring, operators may notice abnormal fluctuations in temperature readings but be unsure if the emergency shutdown criteria have been met. In this situation, the operator can press and hold the "voice question" button and directly ask, "The temperature is fluctuating significantly right now, may I press this button?"
[0137] Front-end interaction focus capture: The system uses EventSystem.current.RaycastAll to emit a ray and detects that the mouse pointer is currently hovering over the "Emergency Stop Button" with the attached identifier component. The system reads the component's ID: "Btn_Emergency_Stop_01". The system uses this ID to look up the configuration data in memory and extract the context labels: ["Chemical Safety", "ESD System", "Emergency Shutdown"], as well as the corresponding system prompts.
[0138] Communication payload encapsulation: The front-end constructs an HTTP POST request, with the body containing: query: "The temperature is fluctuating greatly now, can I press this button?" context_tags: ["Chemical Safety", "ESD System", "Emergency Shutdown"]system_prompt: "The current user is concerned with the highest priority safety device in the entire plant..."
[0139] Backend RAG Inference: 1) Tag Filtering Retrieval: The backend vector database receives the request and uses context_tags to perform filtering, only searching within the subset of documents tagged "ESD System". This directly excludes irrelevant content in the knowledge base such as "canteen management" and "attendance system". 2) Knowledge Recall: Within the limited scope, the search engine calculates similarity and recalls the paragraph about "850℃ critical value" in the SOP document. 3) Response Generation: The large language model combines the recalled SOP paragraph, system prompts, and user questions to generate the response: "Please be extremely cautious! According to the safety operating procedures, this button should only be operated when the temperature exceeds 850 degrees and the backup cooling system is confirmed to be ineffective. Currently, the temperature is only fluctuating; pressing the button is strictly prohibited. It is recommended to check the coolant flow rate first."
[0140] Multimodal feedback: The front end receives the response text. The virtual agent immediately broadcasts the warning in a serious tone via TTS, while a semi-transparent text bubble pops up in the lower right corner of the screen, displaying the content word by word with a typewriter animation effect, completing the interactive loop from the user's vague question to a professional and accurate answer.
[0141] Compared with existing technologies, the "Adaptive Intelligent User Guidance Method for Digital Twin Projects" proposed in this invention, through its unique system architecture and data organization method, successfully overcomes many shortcomings of existing technologies and has the following significant effects and advantages:
[0142] 1. Significantly simplifies the development process, reducing the complexity of bootstrapping system development to a simple task. The core of this invention lies in providing a standardized software framework that transforms bootstrapping system development from a complex "programming task" into a simple "configuration task." Developers no longer need to write and maintain cumbersome bootstrapping logic code for each project; instead, they simply "mark" UI elements and fill in a standardized JSON "manifestation." This configuration-driven model reduces what used to require weeks of coding and debugging to hours of configuration and import, freeing developers from repetitive tasks and significantly reducing R&D costs and manpower investment. Furthermore, as a mature asset, this framework can be plugged and played across different projects, achieving high reusability and facilitating various modifications.
[0143] 2. Breaking down industry knowledge barriers and deeply integrating operation guidance with knowledge transfer. One of the core innovations of this invention lies in its hierarchical knowledge data structure, which strongly binds UI elements with deterministic "operation instructions" and unstructured "domain knowledge." This allows the digital human guide to not only tell users "where to click," but also, through linkage with the backend RAG system, answer users' deeper professional questions in real time and accurately, such as "what is this?" and "why is this value?" It acts like a domain expert embedded within the software, effectively solving the operational interruptions and comprehension difficulties caused by users' lack of professional background knowledge, ensuring the continuity and smoothness of the workflow.
[0144] 3. Significantly improves user experience, providing an adaptive and interactive learning path. This invention integrates structured guidance and intelligent question-and-answer modes, giving users a high degree of freedom. Users can systematically learn software functions by following preset processes, or interrupt the process at any point to ask questions in natural language about points of confusion on the current interface. This flexibility and adaptability can meet the personalized needs of users with different knowledge levels and usage habits, transforming passive "cramming" teaching into active "exploratory" learning, thereby greatly enhancing users' learning interest and acceptance.
[0145] To ensure the accuracy and reliability of guidance information and meet the stringent requirements of industrial applications, this invention addresses the potential "fact illusion" problem inherent in general-purpose large language models. By employing a retrieval-enhanced generation architecture and strictly limiting its knowledge source to a professional knowledge base provided by the developer, the accuracy and reliability of all question-and-answer content are guaranteed. This is crucial for serious application scenarios such as industrial digital twins, where data and operational error tolerance is extremely low, ensuring that the guidance information received by users is professional and trustworthy.
[0146] The above embodiments are merely illustrative of several implementation methods of this disclosure, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept of this disclosure, and these modifications and improvements all fall within the protection scope of this disclosure.
Claims
1. A method for intelligent user guidance in digital twin projects, characterized in that, The intelligent user guidance method includes the following: A front-end system is established, which is based on a 3D graphics engine and consists of a plug-in system that goes from static configuration to dynamic mapping. A real-time index is created by combining externally defined industry knowledge with dynamically generated interface elements at runtime using unique identifiers. The functional attributes, guidance logic, and associated industry knowledge of the user interface are abstracted and persistently stored as static configuration files. The abstract data in these static configuration files is mapped to specific object instances in the 3D engine's rendering pipeline. An explicit identification mechanism is set up at the component system level of the 3D engine. Simultaneously, to balance the human-like interactive experience with lightweight system operation, a two-dimensional virtual proxy image is constructed at the user interface level as a unified medium for user-system interaction. This two-dimensional virtual proxy in screen space serves as the interaction entry point, executing structured linear guidance. A backend system is established, and the deployment of the backend system is based on the intelligent service generated by retrieval enhancement. When a user initiates a natural language question, the frontend system captures the context label of the current interaction focus. Based on this, the backend retrieves professional documents within a limited scope and generates a business-aware answer by combining a large language model, thus realizing an interactive closed loop from operation navigation to professional Q&A. The component system identifier and data association mechanism of the 3D engine at runtime is as follows: The data of the identification component is encapsulated, and a custom script component class is developed. The script component class inherits from the base component class of the graphics engine and can be mounted on scene objects. The script component class declares and serializes a public string variable. The name of the variable is consistent with the unique identifier field in the static configuration file. During the editing and building phase of the application, the string variable is manually filled to establish the logical correspondence between the scene objects and specific data entries in the static configuration file. The global registry is constructed in memory. A global singleton registry manager is built in the runtime memory of the application. The registry manager maintains an efficient key-value pair container based on a hash algorithm. The key type of the efficient key-value pair container is constrained to be a string, corresponding to the unique identifier field. The value type of the efficient key-value pair container is constrained to the general object reference type of the graphics engine; The automated registration process manages the lifecycle using the lifecycle callback interface provided by the graphics engine. The automatic registration logic is triggered the instant the identifier component is instantiated and enters memory. This logic reads the identifier string stored internally within the component instance and writes it as a key-value pair with the memory address reference of the object to which the component belongs. If the same key already exists in the hash container, an overwrite update operation is performed. This ensures that regardless of whether the user interface object is a static resource loaded with the scene or a dynamic resource generated by code logic, as long as the identifier component is mounted and initialized by the graphics engine, it will be automatically included in the index scope of the global registry without additional external intervention.
2. The intelligent user guidance method for digital twin projects according to claim 1, characterized in that, The metadata definition method of the static configuration file is to define a data graph used to describe the bootstrap node. The data exchange format file internally maintains a linear array structure, where each element is an independent data object representing a specific, bootstrap user interface functional unit in the application. Each data object contains the following field definitions: A unique identifier field, which is used to uniquely identify a user interface element globally; The bootstrap sequence index field is used to define the time execution order of the functional units in the deterministic linear bootstrap process. The logic control unit in the backend system will determine the triggering order of the bootstrap events based on the arithmetic ascending order of the time execution order values. The voice broadcast content field stores the text information that the anthropomorphic agent needs to output to the user when executing structured guidance. The text content will be directly passed to the text-to-speech engine for audio waveform synthesis. The knowledge base index tag set field, which serves as the core metadata of the backend system, stores a set of domain-specific terms that are semantically strongly related to the functions of the user interface. These terms will be used as filtering conditions for vector retrieval to limit the knowledge search scope of the large language model in the backend system. The context system prompt field stores a natural language description used to define the business boundaries, preconditions, or potential impacts of the user interface function at the semantic level. The content of the context system prompt field will be injected into the dialogue context as a system-level instruction when building the large language model prompt project, so as to establish the business background of the dialogue.
3. The intelligent user guidance method for digital twin projects according to claim 1, characterized in that, The visual coverage and layout method of the virtual agent image is as follows: An independent visual display area is established at the top of the application's display interface. Within this area, a two-dimensional digital character image with a transparent background is rendered. This image is fixedly positioned in the non-operational core area of the screen display to ensure that continuous, supportive guidance is provided without obscuring the underlying business operation interface or data display content.
4. The intelligent user guidance method for digital twin projects according to claim 1, characterized in that, Establish a multimodal interaction and performance mechanism for the aforementioned two-dimensional virtual agent image: The multimodal interaction includes voice interaction capability. The two-dimensional virtual agent image is equipped with a complete voice input and output interface, which can receive the user's voice commands and perform natural language broadcast through speech synthesis technology. The performance mechanism includes facial expressions and actions. The system maintains a state mapping table for the two-dimensional virtual agent image. The state mapping table can dynamically update the visual presentation effect of the digital human according to the current interaction state. The visual presentation effect includes switching between different emotional portrait images, playing preset basic two-dimensional animation sequences, and mouth opening and closing actions that match the rhythm of voice broadcasting, thereby achieving a vivid and natural anthropomorphic interaction effect.
5. The intelligent user guidance method for digital twin projects according to claim 1, characterized in that, The structured linear guidance of the front-end system relies on local logic control, guiding the user through the key functions of the software according to a predefined logical order, as follows: When the bootstrapping process is activated, the logic controller reads the global configuration data. The front-end system arranges all relevant boot nodes in ascending order according to the preset boot order index in the data, generates a linear task execution queue, and prepares to start execution from the first step. The execution steps involve the front-end system sequentially reading the current guidance node in the task execution queue. First, it performs target localization and highlighting, which involves reading the unique identifier field in the current guidance node, finding the corresponding user interface object position in the scene, and generating or moving a highlighted indicator box at the position to visually mark the operation area that needs attention. Then, it performs voice explanation, which involves extracting the voice broadcast content in the current guidance node, driving the two-dimensional virtual agent image to read the text through voice synthesis technology, and coordinating with corresponding body or lip movements to complete the explanation. In the process flow control, after the current step's highlighted mark and voice explanation are completed, the front-end system automatically pauses the process and displays a "Next" interactive button. When the user clicks the button, the front-end system removes the current highlighted mark, moves the execution pointer to the next node in the queue, and repeats the above execution steps until all content in the queue has been guided.
6. The intelligent user guidance method for digital twin projects according to any one of claims 1-5, characterized in that, A context-aware intelligent question-answering system is established between the front-end system and the back-end system. The context-aware intelligent question-answering system realizes intelligent interaction between the user's natural language input and the large language model in the back-end system. The context-aware intelligent question-answering system constructs a communication carrier containing precise business context and performs knowledge retrieval based on the context tags in the back-end system. The specific method is as follows: first, the interaction focus of the front-end system is captured and the context is extracted; then, the back-end system generates retrieval enhancement logic based on tag filtering; and finally, closed-loop feedback and display are performed.
7. The intelligent user guidance method for digital twin projects according to claim 6, characterized in that, When a user triggers the voice question function, the front-end system starts the recording device to capture the audio stream and calls the automatic speech recognition interface to convert the captured audio stream into a text string. Then, the context-aware intelligent question-answering system executes the key context capture logic: The ray-based focus detection algorithm uses the input module of the graphics engine to obtain the screen coordinates of the current mouse pointer, emits a three-dimensional ray from the camera position to the screen coordinates, and uses the ray projection interface of the physics engine or user interface event system to detect the collision between the ray and the user interface objects in the scene. Then, it returns the first valid interactive object that the ray intersects as the current focus object. In the reverse lookup of the identifier component, the context-aware intelligent question answering system attempts to obtain the reference of the custom identifier component on the focused object. If the acquisition fails, the current context is marked as empty, and only the user's question text is retained. If the acquisition is successful, the unique identifier stored in the custom identifier component is read. Then, the unique identifier is used to perform a reverse lookup in the static configuration file to locate the corresponding configuration data object. The communication payload is structured and encapsulated. The context-aware intelligent question-answering system extracts the content of the knowledge base index tag set field and the context system prompt word field from the located configuration data object, and constructs a POST request packet conforming to the Hypertext Transfer Protocol standard. The data body of the request packet is serialized into a JSON format string, containing the following key-value pairs: query: the user's original question text; context_id: the unique identifier of the focused object; context_tags: the extracted tag array; system_prompt: the extracted business description text; the request packet is sent to the service interface of the backend system through a network socket.
8. The intelligent user guidance method for digital twin projects according to claim 7, characterized in that, The method for generating enhanced search logic based on tag filtering in the backend system is as follows: The backend system maintains a vector database for conditional filtering retrieval. This database stores pre-processed and quantized industry technical document slices. Each document slice is tagged with specific metadata tags upon entry into the database. The retrieval module first reads the context_tags array from the request packet, constructs a filtering expression, and instructs the vector database to search only in the document set whose metadata tags are contained within the context_tags array. Subsequently, the retrieval module converts the user's query text into a high-dimensional vector, calculates the cosine similarity in the aforementioned limited document set, and recalls the top-N document slices with the highest similarity scores. The prompt word engineering involves dynamic assembly. The backend proxy module constructs the final prompt words sent to the large language model based on the search results. The structure of the prompt words strictly follows the template below: At the system instruction layer, the system_prompt content in the request packet is injected to force the model to enter a specific business role setting. Referring to the knowledge layer, the plain text content of all recalled document slices is concatenated as factual evidence; In the user command layer, inject the user's query text; Generative reasoning and response: The assembled prompt words are submitted to the large language model reasoning interface. The large language model generates response text based on the limited context environment. The backend system encapsulates the text into an HTTP response packet and returns it to the frontend system.
9. The intelligent user guidance method for digital twin projects according to claim 8, characterized in that, The closed-loop feedback and display method is as follows: after the front-end network module receives the response text, it distributes it to the virtual proxy display module. The front-end system calls the text-to-speech interface to play the answer audio, and instantiates a text bubble control on the side of the virtual proxy to display the answer text in a visual typewriter effect, thereby completing a complete, context-aware intelligent question-and-answer interaction closed loop.
Citation Information
Patent Citations
AI extensions and intelligent model validation for an industrial digital twin
CN111562769A
Multi-mode collaborative output virtual human intelligent question and answer method and system
CN121051203A