Behavior insight and perception enhancement method and system based on digital twinning

By using recursive semantic fusion and localized application of large language models, the semantic gap and privacy issues in behavior analysis are resolved, enabling dynamic evolution of user profiles and context-aware interaction, thereby improving the efficiency of user behavior understanding and management.

CN122064955APending Publication Date: 2026-05-19GUANGZHOU SUPER ENTROPY INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU SUPER ENTROPY INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-02-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing behavioral analysis technologies suffer from semantic gaps, lack of holistic digital profiles, lack of dynamic evolution and contextual memory, lack of proactive intelligent intervention, and privacy and computational challenges in the application of large language models. They are unable to transform low-semantic-density computer operation logs into high-semantic-density user personality profiles and provide context-aware interaction.

Method used

It adopts a recursive semantic fusion mechanism, builds a local behavior log repository through non-intrusive data collection, uses a large language model for semantic reasoning and state mapping to achieve fully localized data display, and provides context-aware intervention through virtual expert interaction.

Benefits of technology

It has achieved the transformation from low semantic density data to high semantic density profiles, has the ability to evolve throughout the entire lifecycle, provides enhanced user self-awareness with natural interaction and privacy protection, and improves the efficiency of user behavior understanding and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064955A_ABST
    Figure CN122064955A_ABST
Patent Text Reader

Abstract

The invention discloses a behavior insight and perception enhancement method and system based on digital twinning, and relates to the technical field of computer data processing and artificial intelligence. The method comprises the following steps: firstly collecting a computer terminal operation behavior log of a user as basic data mapping; performing semantic understanding and feature extraction on the discrete behavior logs by using a large language model as a semantic calculation engine and combining a preset cue word engineering technology, and constructing a digital behavior twin model capable of mapping behavior habits, intentions and psychological states of the user; based on the twinning model, generating a behavior insight report including multiple dimensions such as personal portraits, staged review, target tracking prediction and the like; and finally, by integrating an interactive interface of the visual board and the natural language dialogue ability, the user can visually view and interact with the digitized behavior of the user. According to the method, low-value original behavior data can be converted into high-value semantic intelligence, and the perception depth and the self-optimization capability of a user on a digital behavior mode of the user are remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer data processing, artificial intelligence (AI) applications, human-computer interaction (HCI), and digital personal management. Specifically, this invention relates to a method, system, electronic device, and computer-readable storage medium that utilizes Large Language Model (LLM) technology to process heterogeneous user desktop behavior data, constructs a high-dimensional user digital twin model, and achieves user self-perception enhancement and behavioral performance optimization through recursive semantic analysis, dynamic visualization rendering, and context-aware dialogue interaction. This invention is particularly suitable for applications requiring deep self-awareness, personal performance management, career development planning, psychological state monitoring, educational guidance, and corporate human resource optimization. By performing non-intrusive data collection at the operating system level and utilizing generative artificial intelligence technology for semantic-level deep insights, this invention aims to solve the challenge of mapping from low-semantic-density raw behavioral logs to high-semantic-density psychological and behavioral profiles. Background Technology

[0002] In today's highly digitalized information age, personal computers (PCs) are no longer merely computing tools, but rather the core carriers of modern human work, study, entertainment, social interaction, and even emotional expression. Every interaction a user makes within a computer operating system—including but not limited to keyboard typing rhythm, mouse movement, window focus switching, application launches and closures, and file system read / write operations—constitutes a holographic projection of their behavior in the digital world, known as a "digital behavioral footprint." According to theories of behavioral psychology and data science, this massive, continuous, and fine-grained behavioral data theoretically contains deeper information about users' behavioral patterns, thinking habits, cognitive styles, efficiency bottlenecks, emotional fluctuations, and even potential mental health states. However, despite the increasing maturity of data collection technologies, existing behavioral recording and analysis technologies still face significant technical bottlenecks and limitations in deeply mining the intrinsic value of this data. These limitations are mainly reflected in the following key dimensions: 2.1 The Semantic Gap in Data This is the primary technical obstacle preventing users from understanding their own behavioral data. Traditional logging software, time tracking tools, or system monitoring software typically only generate raw log files (such as CSV, JSON, or database records) with technical identifiers at their core. These logs usually contain precise timestamps, process names, window titles, and window class names. For computer operating systems, this data is structured and clear; however, for ordinary users, psychologists, or business managers—people without a technical background—this data presents a significant "semantic gap." For example, class names and identifiers recorded in logs, such as Qt51514QWindowIcon, Chrome_WidgetWin_1, afxFrameOrView140s, or OpusApp, have almost no intuitive semantic meaning for human observers. Even window titles with slightly better readability (such as "Document1 - Word") fail to reflect the user's true intent due to a lack of contextual information. Specific Case Analysis: When user behavior logs show frequent switching between the "Visual Studio Code" (code editor) and "Stack Overflow - Google Chrome" (technical Q&A website) windows within a short period, traditional tools can only record a series of "window activation" events. They cannot understand the high-level semantics behind this behavioral sequence—that the user is in the specific task context of "encountering specific technical difficulties in coding and actively seeking external solutions." Traditional technologies also cannot distinguish the essential difference between "engaging in creative programming" (spending a long time in the editor) and "engaging in mechanical debugging" (frequently switching between consulting documentation). This "low semantic density" data leads to a significant gap between low-level data collection and high-level value insights, making it difficult for users to directly use the massive amounts of data to guide their behavioral improvements, even though they possess such data. 2.2 Lack of Holistic Digital Personas Existing behavioral analysis tools mostly remain at the superficial stage of descriptive statistics. Their output is usually in the form of static dashboards, displaying fragmented and atomic indicators such as "Application A was used for 30 minutes today," "Website B was visited 10 times this week," and "Efficiency score 85 points." 1 . This statistical approach to analysis has serious limitations: it cannot aggregate discrete behavioral data points into a richly detailed, semantically meaningful "personality profile" or "digital twin." Users cannot use these tools to see whether their "thinking patterns" are divergent or focused, cannot identify in which professional field they exhibit a "flow state" to infer their "potential talents," and cannot extract their "core identity tags" (whether they are a cultivator, an explorer, or a browser). This lack of holistic modeling and abstraction results in fragmented, cold numbers that fail to translate into effective feedback that promotes users' metacognition update. 2.3 Lack of Dynamic Evolution and Contextual Memory Human behavior is continuous and evolving, deeply influenced by the cumulative effects of historical states and environmental changes. A person's "today" is a continuation and development of "yesterday." However, existing behavioral analysis systems typically employ a "sliced" processing logic, analyzing data independently for static time periods (such as "today" or "this week"), lacking a mechanism to organically semantically integrate "previous state" with "current behavior." Specific technical pain points: For example, a user's behavior a month ago indicated they were in the "early stages of postgraduate entrance exam preparation," frequently browsing university websites; while today's behavior shows they are browsing "adjustment information" or "job recruitment websites." If the system lacks historical memory and evolutionary capabilities, it will only isolate the user's behavior as "browsing the web" or "looking for a job"; a system with evolutionary capabilities should be able to identify this as "strategic adjustment after encountering setbacks in the exam preparation stage" or "risk signals." The fragmented nature of existing technologies means that the analysis results cannot reflect the user's trajectory over time or the trend of state decline, making it difficult to build a digital twin with a full lifecycle perspective. This static analysis cannot adapt to the dynamic complexity of human behavior. 2.4 Lack of Proactive Intelligent Intervention Traditional behavior analytics systems are typically positioned as "passive recorders." They faithfully record data, generate reports, and then completely delegate the responsibility of interpreting the data, identifying problems, and developing improvement plans to the user. For most users who lack data analysis skills or the willpower for self-management, such passive systems are unlikely to produce actual behavior change. The system lacks an intelligent interactive mechanism that can understand the context of a user's current behavior and provide immediate feedback, emotional support, strategic advice, and risk warnings, much like a real expert (such as a graduate school tutor, CEO coach, psychologist, or career planner). For example, when a user spends a long time on entertainment websites during work hours, a traditional system can only generate a "low-efficiency" weekly report afterward; an ideal system should be able to detect this state in real time and intervene through an anthropomorphic role (such as a "strict mentor" or a "gentle partner"), asking the user if they have encountered difficulties or are feeling down. The gap between "post-event review" and "in-event intervention" is one that current technology has failed to bridge. 2.5 Challenges in Applying Existing Large Model Techniques With breakthroughs in Large Language Model (LLM) technology, exemplified by the Transformer architecture, computers have for the first time gained the ability to understand complex contexts, perform logical reasoning, and generate natural language. This provides a new technological path for solving the aforementioned problems. However, directly applying LLM to user behavior analysis still faces many challenges: 1. Privacy issues: User desktop behavior data is extremely sensitive, and directly uploading it to the cloud LLM poses a huge risk of privacy leakage. 2. Context window limitation: The context window of an LLM is limited and cannot handle the massive amount of logs that a user may have accumulated over several months or even years at once. 3. Computational costs and latency: Real-time analysis of full data will result in high token costs and computational latency. 4. Practicality of technical solutions: How to efficiently transform streaming, low-density behavior logs into high-density user profiles, and how to build an evolvable, locally stored "digital twin system" to ensure privacy, are the technical challenges that urgently need to be addressed. Summary of the Invention

[0003] This invention aims to solve the aforementioned technical problems by providing a method and system for behavior insight and perception enhancement based on digital twins. This invention creatively introduces a "Recursive Semantic Fusion" mechanism, utilizing a large language model to transform local behavior logs into an evolvable HTML-formatted digital twin model, and achieving fully localized data closure and interactive display through the Electron desktop application framework. 3.1 Core Technology Issues The present invention mainly solves the following technical problems: 1. How to transform low-semantic-density computer operation logs (time, window title, class name) into high-semantic-density user profiles and behavioral insights. 2. How to achieve continuous evolution and dynamic updating of user profiles over time when the LLM context window is limited. 3. How to achieve localized data collection, analysis, and visualization while protecting user privacy. 4. How to provide context-aware, human-like expert interaction and intervention based on real-time behavioral data. 3.2 Technical Solution To solve the above-mentioned technical problems, the present invention adopts the following technical solution: 3.2.1 A method for behavioral insight and perception enhancement based on digital twins The method includes the following steps: Step S1: Non-intrusive Heterogeneous Behavioral Data Stream Acquisition and Structured Processing. A file system watcher is deployed at the operating system level of the user terminal. Utilizing kernel-level file change notification mechanisms (such as inotify, FSEvents, and ReadDirectoryChangesW), the watcher captures the user's interactive event stream in the graphical user interface (GUI) in real time. This interactive event stream includes at least a timestamp sequence, the foreground active window title (WindowTitle), and the window class name. Through a stream writing mechanism, these heterogeneous interactive events are appended in real time to a local behavior log repository (CSV Repository), forming a time-series-based raw behavioral data stream. Step S2: Construction and Evolution of the Recursive Digital Twin State Vector. In response to change events in the behavior log repository, the latest N behavior records within a preset time window are read as the current behavior slice. A system prompt template containing system-level role settings, task instructions, and output constraints is constructed. The previous-time digital twin state file (Persona State File) is read from the local state storage. The system prompt, the content of the previous-time digital twin state file, and the current behavior slice are contextually fused to construct a composite input sequence. Step S3: Semantic Reasoning and State Mapping Based on a Large Language Model. The composite input sequence is sent to the pre-trained Large Language Model (LLM) inference interface. Utilizing the semantic understanding and logical reasoning capabilities of the LLM, the user intent, thought patterns, and behavioral characteristics implicit in the behavior slices are analyzed. Combined with the state information from the previous moment, updated digital twin model data is generated. The digital twin model data is represented in Structured Markup Language (HTML / JSON) format, covering a multi-dimensional feature space encompassing basic information, core identity tags, thought patterns, and potential needs. Step S4: Dynamic rendering and feedback of the enhanced perception interface. The structured markup language code returned by the large language model undergoes security cleaning and format validation. The cleaned code is injected into the application's front-end rendering container (WebView) to dynamically update the user profile interface, target list, or analysis report, thereby intuitively presenting the current state of the user's digital twin on the user's terminal and achieving digital enhancement of the user's self-perception. 3.2.2 Refined Scheme for Recursive State Construction The recursive digital twin state vector construction in step S2 specifically includes the following sub-steps: S2.1 State Regression: Check if a historical profile file exists locally. If it exists, read its complete HTML text content as the "previous state context"; if it does not exist, initialize it to an empty state. S2.2 Incremental Awareness: Set a data reading threshold (e.g., 180 lines) to read only incremental data at the end of the behavior log file, avoiding redundant consumption of computing resources caused by full data processing. S2.3 Prompt Engineering: Dynamically assemble Prompts, which explicitly define the instruction model: "Based on the latest user computer operation logs and previous user profiles, analyze their behavior and develop a more comprehensive user profile that is compatible with previous ones." S2.4 State Iteration: The new HTML content output by the model is overwritten and saved as the current state file, and a historical version file with a timestamp is generated at the same time, forming the full life cycle evolution trajectory of the digital twin. 3.2.3 Construction of Multidimensional Feature Space The digital twin model is constructed covering at least 18 orthogonal analytical dimensions, including but not limited to: basic information, core identity tags, behavioral characteristics, behavioral patterns, thinking patterns, mentality, typical scenarios, professional fields and corresponding levels, personality (based on the Enneagram model) and habits, potential and talents, life cycle stage, personal development prediction, tool preferences, potential needs and pain points, risk points, environment, recommended items, and comparison with the behavior of world-renowned figures. For each dimension, the system maintains independent prompt word templates and status files to achieve fine-grained twin characterization. 3.2.4 Virtual Expert Interaction Mechanism The method also includes real-time interactive intervention steps based on specific roles: S4.1 Role Instantiation: In response to the user's mode selection command in the dialogue interface, instantiate a preset virtual expert agent, which includes a companion motivator, a postgraduate entrance examination tutor, a corporate CEO mentor, a crowdfunding strategist, or an influence building mentor. S4.2 Context Dynamic Injection: When a user sends a natural language message, the system automatically extracts the top M (e.g., 120) behavior log records at the current moment. S4.3 Behavior-Aware Response: User messages and captured behavior logs are used together as context input to the large language model, enabling the virtual expert agent to provide context-aware feedback or suggestions based on the user's real-time behavioral state (such as whether they are entertaining or working). 3.2.5 A Behavioral Insight and Perception Enhancement System Based on Digital Twin The system includes: ●Data Acquisition Module: Used to monitor and record user window operation behavior in the operating system in real time, and generate a behavior log stream in CSV format. ● Semantic Analysis Engine: Integrates a large language model API client, responsible for the dynamic construction of prompts, sending API requests and receiving response data, and supports switching and configuration of multiple model backends (OpenAI, Gemini, DeepSeek, Ollama). ● Digital Twin State Library: A storage structure built on the local file system for persistently storing multi-dimensional HTML profile files and their historical versions. ●Rendering and Display Module: An embedded browser container based on Webview technology, used to parse and render HTML code generated by the semantic analysis engine, providing visualized personality profiles, goal lists, and analysis reports. ● Interactive Control Module: Provides a graphical user interface, including a primary function navigation, a secondary dimension list, and a main content interaction area. It supports user customization of data slice size, API parameters, and system prompts. 3.3 Beneficial Effects Compared with the prior art, the present invention has the following significant advantages: 1. Significantly improved semantic interpretability: Transforming obscure computer system logs (such as class names and process IDs) into natural language descriptions and intuitive visualizations greatly lowers the cognitive barrier for users to understand their own behavioral data. Users are no longer faced with cold, hard data, but with a digital twin that "understands" them. 2. Fully Localized Privacy Protection Architecture: All log monitoring, file storage, and HTML rendering in this system are completed on the user's local device. The LLM serves only as a stateless inference computing unit. Users can choose to deploy models privately (such as Ollama), ensuring that the original behavior logs never leave the device, thus maximizing the protection of user behavior privacy. 3. Efficient utilization of computing resources: Employing file streaming watcher and incremental reading technologies, only the latest N data entries are processed for recursive updates. This mechanism avoids memory overflow and computational lag caused by loading GB-level historical logs, ensuring smooth system operation even on ordinary PC configurations, and possesses extremely high engineering practicality. 4. Natural and Immersive Interaction: The HTML / Tailwind CSS code generated by LLM is rendered directly through WebView, eliminating the need for intermediate data parsing and third-party chart library conversions. This makes the generated content extremely flexible, allowing for dynamic adjustments to layout, color, and interactive elements based on analysis results, providing a smooth, self-reflective experience similar to browsing a modern webpage. 5. Possesses full lifecycle evolution capability: Through "recursive semantic fusion", the system can remember and accumulate the user's historical state, so that the digital profile is continuously enriched, corrected and grown over time, truly reflecting the user's development trajectory. Attached Figure Description To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort. Figure 1This is a system overall logical architecture diagram provided in this embodiment of the invention. The diagram illustrates the data flow and interaction between the local file system, the Electron main process, the renderer process, and the LLM API, demonstrating the dual-process communication mechanism and external API calling logic under the Electron framework. Figure 2 This is a schematic diagram of the recursive digital twin profile update process in an embodiment of the present invention. The diagram focuses on how the "current profile (old context)" and "new behavioral data (trigger source)" are fused by the "profile update engine" to generate the "new profile," which is then used as the recursive closed-loop logic for the next update context (focusing on how the old profile participates in the generation of the new profile as context). Figure 3 This is a diagram illustrating the structure of the Prompt generated in this embodiment of the invention, representing the generation of an 18-dimensional personal profile. The diagram details the components of the Prompt, including multi-source raw data input, feature extraction and vectorization, dimension mapping and construction (showing 18 specific dimensions), the structured Prompt output (instruction header, data body, generation constraints), and LLM processing and result generation. Figure 4 This is a sequence diagram of data injection and context awareness in the virtual expert role dialogue system of this invention. The diagram describes the entire process from user input query (T1) to data injection module acquiring external knowledge / real-time data (T3-T5), and then to constructing rich context (T6) and AI generating answer (T8). Figure 5 This is a schematic diagram of the system user interface (UI) layout provided in an embodiment of the present invention. The diagram shows a classic three-column layout: a primary navigation bar on the left (personal profile, perception enhancement, goal management, data, settings), a secondary list bar in the middle (a list of 18 profile dimensions), and a main content display area on the right (a profile report rendered in HTML), reflecting an intuitive human-computer interaction design (showing the three-column structure and main functional areas). Figure 6 This is a schematic diagram illustrating the generation logic and HTML output example of the "Efficiency Growth Analysis Report" in this embodiment of the invention. The diagram shows the complete pipeline from raw logs to data preprocessing, core indicator calculation, AI insight analysis, and finally the generation of an HTML report containing trend charts and key conclusions. Detailed Implementation The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the terms "digital twin," "large language model," "Webview," and "Electron" used in this invention should be interpreted based on the general understanding of those skilled in the art and in conjunction with the specific definitions in this specification. 5.1 System Hardware Environment and Software Architecture Design 5.1.1 Hardware Operating Environment The method described in this invention can be used on various types of electronic devices, including but not limited to desktop computers, laptops, and workstations. Electronic devices typically include: ● Processor: A single-core or multi-core CPU (such as Intel Core i5 / i7, AMD Ryzen) used to execute the Electron main process and rendering process. ●Memory: Random access memory (RAM), 8GB or more is recommended, for loading applications and caching behavior log slices; read-only memory (ROM) or hard disk (HDD / SSD) for persistent storage of behavior log files (CSV) and generated digital twin state files (HTML). ●Input devices: Keyboard, mouse, touchpad, etc., used to generate user interaction events. ● Display device: Liquid crystal display (LCD) or organic light-emitting diode display (OLED) for displaying sensory enhancement interfaces. ● Network interface: Wired or wireless network adapter for communicating with the cloud-based LLM API (offline if using a local model). 5.1.2 Software Architecture Design This system is developed using the Electron framework, leveraging its cross-platform compatibility (Windows / macOS / Linux) and the powerful capabilities of the Chromium rendering engine. The system architecture is logically divided into the following core modules, such as... Figure 1 As shown: 1. Main Process Controller: Runs in the Electron main process. It is the "brain" of the system and is responsible for: ● Manage the application lifecycle (startup, exit). ● Perform native file I / O operations in the Node.js environment (read and write CSV, save HTML). ●Manage system configuration (persist user settings via electron-store). ● Handles inter-process communication (IPC), receiving instructions from the rendering process and returning data. 2. Data Collection Module: Responsible for monitoring underlying behavioral data. It utilizes the chokidar or fs.watch library, directly mounted to the operating system's file system, to monitor changes to the target CSV file. 3. Semantic Analysis Engine: Running in the main process, it integrates an HTTP / HTTPS client (such as axios or fetch). It is responsible for sending structured prompts to the configured LLM API endpoints and processing the returned Stream or JSON data. 4. Rendering Module: Runs within the Electron renderer process. Based on React, Vue, or a pure HTML / JS technology stack, it is responsible for UI rendering. The core component is a full-screen WebView or div container used for seamless rendering of the HTML code generated by the LLM. 5. Digital Twin State Repository: A storage structure based on local file directories. The system maintains subdirectories such as persona / , targets / , and practice / under the userdata directory, storing HTML files and their historical versions for different dimensions. 5.2 Step S1: Acquisition and Structured Processing of Non-Intrusive Heterogeneous Behavioral Data Streams This step aims to obtain raw behavioral data of the user during computer operation, which is the foundation for building a digital twin. 5.2.1 Deployment of the File System Watcher Unlike traditional hooking techniques (such as Windows SetWindowsHookEx) that directly intrude into the system message queue, this embodiment employs a lighter, safer, and looser-coupled "bypass listening" mode. The system pre-defines a low-level behavior logger (as a prerequisite or built-in module), which is responsible for writing window switching events to a local CSV file. The method of this invention focuses on real-time monitoring and analysis of this CSV file. In terms of technical implementation, the Node.js chokidar library is used to monitor the target CSV file. chokidar encapsulates the operating system kernel-level file change notification mechanism. ● Use inotify on Linux. ● Use FSEvents on macOS. ● Use ReadDirectoryChangesW on Windows. This mechanism allows applications to detect file changes without polling. When a CSV file experiences a change or add event, the operating system proactively notifies the main process, triggering subsequent processing. This design significantly reduces CPU usage and achieves "non-intrusive" detection. 5.2.2 Capture and Streaming of Heterogeneous Data The captured data stream contains the following key fields, forming a heterogeneous triple structure: ● Timestamp: Accurate to the second, formatted as YYYY-MM-DD HH:MM:SS. This is the foundation of time series analysis. ● Window Title: The title bar text of the active window in the foreground. For example, behavioral insights based on digital twins – Word, WeChat, App.tsx – BehaviorInsight – Visual Studio Code. This is the primary source of semantic analysis, containing rich information such as application status, filename, and webpage title. ● Window Class Name: The underlying class identifier of the application. For example, OpusApp (Word), WeChatMainWndForPC (WeChat), Chrome_WidgetWin_1 (Chrome). This is used to help identify the application type and eliminate ambiguity in titles with the same name (e.g., distinguishing between the web version of WeChat and the client version of WeChat). The data acquisition module uses a stream writing mechanism to append the above data to the local CSV file in real time. 5.3 Step S2: Construction and Evolution of Recursive Digital Twin State Vector This is the core step in achieving "profile evolution" and "contextual memory" in this invention. The system does not analyze massive amounts of historical data from scratch every time, but rather performs incremental updates based on the "current state," which aligns with the concept of "full lifecycle management" in digital twins. 5.3.1 Incremental Data Slicing When the listener detects a change in the CSV file and the number of newly added rows reaches a preset threshold (e.g., 180 rows), the update logic is triggered. Since CSV files can grow very large (hundreds of MB) over time, directly reading the entire file would cause a memory overflow. This embodiment uses a read-last-lines or similar streaming algorithm (such as using fs.createReadStream with a reverse pointer) to read only the latest N records from the end of the file. ● Selection of parameter N: The default value is 180 records. This value is an optimized choice based on experience: 180 records usually cover the user's operations in the last 1-3 hours (assuming an average window switching rate of 1-3 times per minute), which contains enough contextual information without exceeding the LLM's token limit (usually 180 rows of data consume about 2000-3000 tokens). ●Data Cleaning: After reading, the system will perform simple data cleaning to remove garbled characters, blank lines, or system records (such as the window records of "Behavior Insight Genie"), ensuring the signal-to-noise ratio of the input model. 5.3.2 State Backtracking and Context Fusion When constructing the Prompt to be sent to the LLM, the system performs the following fusion operations (such as...). Figure 2 (as shown) ● Read System Prompt ($P_{sys}$): Loads a pre-defined instruction template for a specific dimension (such as "mindset"). ● Read Historical State ($S_{t-1}$): Check if the last generated HTML file for this dimension exists in the userdata / persona / directory. ● If it exists: Read its complete text content as "early user profiling context". ● If not: If the status is blank, the corresponding section in the Prompt will be automatically adjusted to "Build based on initial data". ●Composition: Combines the three into a complete string sequence. Logical formula: Prompt structure example: You are a professional behavioral analyst. Please analyze user behavior... [Historical Context] This is the user's last generated "mindset" profile (HTML format): These are the user's most recent 180 operation logs: 2025-11-30 21:56:31, WeChat, Qt51514QWindowIcon... Please combine historical profiles and new data to generate a new, more comprehensive "mindset" profile that is compatible with previous ones. Output the HTML code directly. This recursive structure ensures that: ● Continuity: The newly generated image contains information from the old image, avoiding the loss of information. ● Corrective: LLM can correct erroneous inferences in old portraits based on new data. ● Efficiency: No matter how long the historical data has been accumulated, the amount of input processed each time remains basically constant, and the computational burden will not increase over time. 5.4 Step S3: Semantic Reasoning and State Mapping Based on a Large Language Model This step utilizes the generalization reasoning capabilities of the Large Language Model (LLM) to map low-semantic behavioral data into high-semantic personality traits. 5.4.1 Semantic Analysis Engine Configuration and API Interaction The system supports the configuration and switching of multiple model backends to adapt to different network environments and privacy requirements: ● Cloud-based models: OpenAI (GPT-4o), Google (Gemini 1.5), Moonshot (Kimi), DeepSeek (V3). Cloud-based models typically have stronger reasoning capabilities. ● Local Model: Ollama (Llama 3, Qwen 2.5). The local model runs on the user's device, and the data never leaves the domain, offering the highest level of privacy and security. In the "System Settings" interface, users need to configure the Base URL, API Key, and Model Name. The engine sends JSON-formatted requests via standard HTTP / HTTPS protocols. To improve user experience, requests are received in a Server-SentEvents (SSE) streaming manner, meaning the interface updates in real time as each token generated by the LLM reduces user anxiety during waits. 5.4.2 Specific Implementation of the 18 Orthogonal Analysis Dimensions This system deconstructs the "digital twin" into at least 18 orthogonal analytical dimensions, each with its own independent Prompt template and storage file. This "divide and conquer" strategy significantly improves the granularity of the profile. For each dimension, the system-generated HTML contains specific Tailwind CSS style classes. For example, the system instructs LLM to use the bg-blue-50 class to render the rational analysis section and the bg-red-50 class to render the risk warning section, thereby achieving visual semantic differentiation. 5.5 Step S4: Dynamic rendering and feedback of the enhanced perception interface This step transforms the "code" output by the LLM into a user-perceptible "interface," thus completing the visualization of the digital twin. 5.5.1 Security Cleaning and Format Verification Due to the inherent uncertainty of LLM output, the returned content may contain Markdown tags (such as `html`), incomplete tags, or redundant explanatory text. Therefore, the system requires rigorous preprocessing before rendering. ●Strip Markdown: Uses regular expressions to remove code block markers, keeping only the correct ones. The internal contents. ●Sanitization: Use DOMPurify or similar libraries to filter potential malicious scripts (XSS attacks). Although the data primarily originates locally, defensive programming is necessary when dealing with externally injected prompts or future expansion of networking functionality. ●Integrity check: Checks whether HTML tags are closed to prevent rendering from breaking the application layout. 5.5.2 Webview Container Rendering and Automatic Archiving The system embeds a [something] in the right-hand main content area of ​​Electron. <webview>Component (or secure iframe). ●Injection logic: Inject the cleaned HTML string into the container via webview.executeJavaScript or the srcdoc property. ●Style isolation: Thanks to the use of Tailwind CSS (Utility-first CSS), the generated HTML comes with built-in style definitions (such as class="p-4 rounded shadow"), allowing for the creation of beautiful cards, progress bars, tag clouds, and other effects without the need for external CSS files. ●Automatic Archiving: Upon successful rendering, the main process saves the HTML string as a separate file (e.g., persona_base_20251130_2200.html) in the userdata directory. This allows users to view their past status at any time in the "History," just like flipping through a diary. 5.6 Implementation of Enhanced Functional Modules 5.6.1 Efficiency Growth Analysis Report This module employs a Chain-of-Thought (CoT) prompting strategy to guide the model to perform in-depth reasoning in stages, rather than directly generating conclusions. ● Phase One (Diagnosis): ●Prompt command: Please scan the logs and plot the user's main workflow. Mark all frequently repetitive operations. This includes time-consuming tasks that are unresponsive for extended periods and operational breakpoints that frequently interrupt work. ●Model Behavior: Identifies users repeatedly copying and pasting between Excel and the ERP system, or checking WeChat every 5 minutes. ●Phase Two (Mining): ●Prompt command: Based on the diagnostic results, identify which operations are low-value, repetitive tasks? Which processes contain [unclear - possibly related to labor or processes]? Is automation a possibility? Calculate the current percentage of 'pseudo-working time'. ●Model behavior: It is pointed out that manual copy and paste can be replaced by Python scripts; it is pointed out that frequent message viewing leads to a 30% efficiency loss. ●Phase Three (Integration): ●Prompt command: Output a structured HTML report. Include key findings and priority-ordered changes. The measures and expected productivity improvement indicators (such as 'expected savings of 45 minutes') are included. ● Output results: Generates an efficiency scorecard with red / green / yellow indicator lights, a list of improvement suggestions, and a chart of expected benefits. 5.6.2 Context-Injected Agent Interaction (FIA) This feature is located under the "Character Dialogue" primary navigation, and it represents a leap from "General Dialogue" to "Context-Aware Dialogue". Role instantiation: The system provides several default System Prompt templates. ● The Godfather of Corporate CEOs: "You are a business guru, skilled at guiding companies to go public... Please guide this user in running a software company..." ● Postgraduate entrance exam tutor: "You are a tutor skilled in guiding students preparing for postgraduate entrance exams... Please monitor their learning progress, be strict but responsible..." Dynamic injection process (e.g.) Figure 4 (as shown) 1. T1 User Input: The user enters the message in the dialog box: "I feel very unproductive and anxious today, what should I do?" 2. T2-T3 Data Extraction: The system intercepts this message and simultaneously calls the file reading module to extract the latest 120 rows of records from the CSV file. 3. T6 Context Construction: Construct a composite Prompt: [Context: User's most recent 120 lines of operation log] (Log shows: In the past 2 hours, mainly browsing Bilibili, Steam, and Weibo, only 10 minutes in Word) [UserMessage: "I feel very inefficient today and anxious, what should I do?"] 4. T8 AI Generation and Response: The model combines log facts with user self-reports to make inferences. ● Reply content: "Hey, I noticed you've spent almost the last two hours switching between video websites and gaming platforms, and you only opened your Word document for 10 minutes. Your anxiety stems from 'avoidance.' Your priority now is to close Steam, put your phone away, and memorize 50 English words. Get moving, and your anxiety will disappear!" 5. This response has a stronger contextual penetration and persuasiveness than the common "Take a deep breath and try the Pomodoro Technique". 5.6.3 Automated Trigger Mechanism To achieve seamless digital twin synchronization, the system incorporates a built-in counter trigger. ●In-Memory Counter: The main process maintains a variable currentLineCount = 0. ● Listen for callbacks: Each time chokidar captures a file change event, it reads the file metadata or uses a stream pointer to calculate the number of new lines, delta. ● Threshold determination: currentLineCount += delta. If currentLineCount >= UserSetting.Threshold (e.g., 180), then: 1. Automatically reset currentLineCount to 0. 2. Trigger background task generation (call the LLM API to execute the S2-S4 process). 3. Once generated, a system notification will pop up via Electron's Notification API: "Your latest behavioral profile has been updated. Click to view." 6. Creative concept statement compared with existing technology This invention embodies significant substantive features and advancements in the following aspects: Comparison Dimensions Existing technology (traditional behavior logging software) Embodiments of the Invention (Digital Twin Insight System) Technological advantages and creativity Data processing logic Statistical analysis calculates Sum, Average, and Count. For example: "Chrome usage time: 2 hours". Semantic reasoning uses LLM to infer intent. For example: "The user is performing intensive front-end debugging work." It bridges the semantic gap, enabling the understanding of the logical connection between "writing code" and "searching for documentation," rather than simply a matter of time spent together. Image building mechanism Static snapshots are analyzed independently each time and have no historical memory. Recursive Evolution: $S_t = f(S_{t-1}, \Delta D)$ It enables the dynamic growth of the profile over time, with the old state serving as the gene for the new state, which conforms to the life cycle definition of "twin". Interaction mode Passive query users need to view the reports themselves. Context-Aware Agent virtual experts can "see" users' real-time behavior. It provides evidence-based, precise interventions, rather than general advice. Output format Fixed-format charts such as bar charts and pie charts have limited expressive power. Generative HTML / UILLM generates rich text web pages. The interface layout and content are dynamically generated by LLM, offering unlimited expressive flexibility (color, typography, metaphor) and personalization. Privacy Architecture Relying on the cloud often requires uploading data to a server for analysis. Local closed-loop + API forwarding supports the local model (Ollama). The original logs never leave the local machine; only the anonymized Prompt fragments are sent to the API, or the system runs completely offline, ensuring extremely high security. 7. Conclusion This invention creatively combines operating system-level file system monitoring technology with recursive reasoning technology from large language models to propose a complete solution for behavioral insight and perception enhancement. This system not only records "what was done," but also understands "why it was done" and "what it means," providing perception enhancement services in a human-like expert role. This invention effectively solves four major problems in the field of traditional behavioral analysis: lack of semantic data, fragmented and stagnant user profiles, passive interaction, and difficulty in protecting privacy. By transforming low-value raw data streams into high-value semantic intelligence, this invention constructs a realistic, dynamic, and interactive digital twin for users, demonstrating significant practical value and broad market prospects in areas such as personal performance enhancement, educational assistance, mental health monitoring, and enterprise management.< / webview>

Claims

1. A method for behavioral insight and perception enhancement based on digital twins, characterized in that, The method includes the following steps: ● Step S1: Non-intrusive acquisition and structured processing of heterogeneous behavioral data streams. A file system watcher is deployed at the operating system level of the user terminal to capture the user's interaction event stream in the graphical user interface (GUI) in real time. The interaction event stream includes at least a timestamp sequence, the foreground active window title, and the window class name. Through a stream writing mechanism, the heterogeneous interaction events are appended in real time to the local behavior log repository (CSV Repository), forming a time-series-based raw behavioral data stream. ● Step S2: Construction and Evolution of the Recursive Digital Twin State Vector. In response to change events in the behavior log repository, the latest N behavior records within a preset time window are read as the current behavior slice. A system prompt template containing system-level role settings, task instructions, and output constraints is constructed. The previous-time digital twin state file (Persona State File) is read from the local state storage. The system prompt, the content of the previous-time digital twin state file, and the current behavior slice are contextually fused to construct a composite input sequence. ● Step S3: Semantic Reasoning and State Mapping Based on Large Language Model. The composite input sequence is sent to the pre-trained Large Language Model (LLM) inference interface. Utilizing the semantic understanding and logical reasoning capabilities of the LLM, the user intent, thought patterns, and behavioral characteristics implicit in the behavior slices are analyzed. Combined with the state information from the previous moment, updated digital twin model data is generated. The digital twin model data is represented in Structured Markup Language (HTML / JSON) format, covering a multi-dimensional feature space encompassing basic information, core identity tags, thought patterns, and potential needs. ●Step S4: Dynamic rendering and feedback of the enhanced perception interface. Perform security cleaning and format validation on the structured markup language code returned by the large language model. Inject the cleaned code into the application's front-end rendering container (Webview) to dynamically update the user profile interface, target list, or analysis report, thereby intuitively presenting the current state of the user's digital twin on the user's terminal and achieving digital enhancement of the user's self-perception.

2. The method according to claim 1, characterized in that, The recursive digital twin state vector construction in step S2 specifically includes the following sub-steps: ● S2.1 State Regression: Check if a historical profile file exists locally. If it exists, read its complete HTML text content as the "previous state context"; if it does not exist, initialize it to an empty state. ● S2.2 Incremental Awareness: Set a data reading threshold (e.g., 180 lines) to read only incremental data at the end of the behavior log file, avoiding redundant consumption of computing resources caused by full data processing. ● S2.3 Prompt Engineering: Dynamically assemble Prompts, with a clear instruction model: "Based on the latest user computer operation logs and previous user profiles, analyze their behavior and develop a more comprehensive user profile that is compatible with previous ones." ●S2.4 State Iteration: The new HTML content output by the model is overwritten and saved as the current state file, and a historical version file with a timestamp is generated at the same time, forming the full life cycle evolution trajectory of the digital twin.

3. The method according to claim 1, characterized in that, The digital twin model is constructed covering at least 18 orthogonal analytical dimensions, including but not limited to: basic information, core identity tags, behavioral characteristics, behavioral patterns, thinking patterns, mentality, typical scenarios, professional fields and corresponding levels, personality (based on the Enneagram model) and habits, potential and talents, life cycle stage, personal development prediction, tool preferences, potential needs and pain points, risk points, environment, recommended items, and comparison with the behavior of world-renowned figures. For each dimension, the system maintains independent prompt word templates and status files to achieve fine-grained twin characterization.

4. The method according to claim 1, characterized in that, The method also includes real-time interactive intervention steps based on specific roles: ● S4.1 Role Instantiation: In response to the user's mode selection command in the dialogue interface, instantiate a preset virtual expert agent, which includes a companion motivator, a postgraduate entrance examination tutor, a corporate CEO godfather, a crowdfunding think tank, or an influence building mentor. ● S4.2 Context Dynamic Injection: When a user sends a natural language message, the system automatically extracts the previous M (e.g., 120) behavior log records at the current moment. ● S4.3 Behavior-Aware Response: User messages and captured behavior logs are used together as context input to the large language model, enabling the virtual expert agent to provide context-aware feedback or suggestions based on the user's real-time behavior status (such as whether they are entertaining or working).

5. The method according to claim 1, characterized in that, The enhanced perception also includes generating an "efficiency growth analysis report," the processing logic of which includes: ● Phase 1 (Diagnosis): Use LLM to scan behavior logs, draw workflow and time consumption graphs, and mark high-frequency operations and long-time tasks; ● Phase Two (Mining): Identify breakpoints in operations that take too long, are repeated frequently, or can be automated, and pinpoint efficiency bottlenecks; ● Phase 3 (Integration): Based on the diagnostic and mining results, output a structured HTML report that includes key findings, priority-ordered improvement measures, expected productivity improvement indicators, and target achievement predictions; ● The report generation process employs a chain-of-thought prompting strategy to guide the model through step-by-step reasoning and output the final result.

6. The method according to claim 1, characterized in that, The method also includes an automated triggering mechanism based on file system monitoring: ● Monitor write operations to behavior log files using operating system kernel-level file change notification mechanisms (such as inotify, FSEvents, ReadDirectoryChangesW). ● Maintain a memory counter to count the number of new log lines since the last generation. ● When the counter reaches a preset threshold (e.g., 180 rows), the background task is automatically triggered to execute steps S2 to S4, achieving near real-time synchronization of the digital twin state without manual intervention.

7. A behavior insight and perception enhancement system based on digital twins, characterized in that, include: ● Data acquisition module: Used to monitor and record user window operation behavior in the operating system in real time, and generate a behavior log stream in CSV format. ● Semantic Analysis Engine: Integrates a large language model API client, responsible for the dynamic construction of prompts, sending API requests and receiving response data, and supports switching and configuration of multiple model backends (OpenAI, Gemini, DeepSeek, Ollama). ● Digital Twin State Library: A storage structure built on the local file system for persistently storing multi-dimensional HTML profile files and their historical versions. ● Rendering and Display Module: An embedded browser container based on Webview technology, used to parse and render HTML code generated by the semantic analysis engine, providing visualized personality profiles, goal lists, and analysis reports. ● Interactive Control Module: Provides a graphical user interface, including a primary function navigation, a secondary dimension list, and a main content interaction area. It supports user customization of data slice size, API parameters, and system prompts.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the method as described in any one of claims 1 to 6.