Digital avatar task processing method based on Web3.0 and multi-modal data fusion
By using multimodal data fusion and blockchain technology, a personalized digital agent model is generated, which solves the problem of insufficient adaptability of personalized automated service systems in different scenarios, realizes accurate user behavior simulation and task execution, and improves the intelligence and security of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-03
AI Technical Summary
Existing personalized automated service systems struggle to fully capture user behavior characteristics and preferences across different scenarios, resulting in a lack of flexibility and adaptability when executing tasks. This makes it difficult to accurately determine users' true needs, often leading to slow or erroneous decision-making.
By collecting browser operation data, screen interaction data, and wearable device perception data, multimodal data fusion is performed to generate a digital agent model. This model is then deployed to the blockchain via smart contracts to dynamically adjust data weights, generate personalized decision-making rules, simulate user operations to complete tasks, and achieve task notarization and profit distribution.
It enhances the intelligence and personalization of task execution, provides an efficient and convenient automated service experience, and ensures the security and reliability of data and revenue, while adapting to user intent judgment in diverse scenarios.
Smart Images

Figure CN121785463A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of blockchain digital technology, and in particular to a digital avatar task processing method based on Web3.0 and multimodal data fusion. Background Technology
[0002] In the field of digital services, personalized automation technology is becoming a crucial pillar for improving user experience and efficiency. This field, by simulating user behavior to complete various tasks, greatly liberates manpower while providing users with convenient services. However, its importance is accompanied by complex technical challenges. How to enable systems to truly understand and adapt to users' personalized needs is a key area that urgently needs breakthrough. Existing methods often struggle to comprehensively capture users' behavioral characteristics and preferences across different scenarios when addressing user needs. Many solutions focus only on operation records in a single scenario, ignoring the correlation of user behavior across multiple scenarios and devices, resulting in a lack of flexibility and adaptability when facing complex tasks. This limitation makes automated services often appear rigid in practical applications, failing to truly match the user's true intentions.
[0003] How can we effectively integrate user information from different sources and dynamically adjust its priority based on the task scenario? User actions in the browser, on-screen interactions, and the states perceived by wearable devices all originate from diverse sources and differ in time and space. If the importance of this information cannot be properly balanced, the system cannot accurately determine the user's true needs at a specific moment. For example, when a user is browsing the web, the system may need to pay more attention to their click and search habits; while in a sports scenario, heart rate or cadence data from the wearable device may be more valuable. If this information is not properly balanced, the system may make misjudgments when executing tasks. For instance, when a user needs to quickly complete a web task, excessive reference to irrelevant sports data may lead to delayed or incorrect decisions.
[0004] Therefore, how to dynamically adjust the weight of various information types in different scenarios so that the system can accurately understand user intent and complete tasks has become a critical issue that urgently needs to be addressed. Solving this problem not only concerns the level of intelligence in automated services but also directly affects user trust and satisfaction during use. Summary of the Invention
[0005] This invention provides a digital avatar task processing method based on Web3.0 and multimodal data fusion, which aims to improve the intelligence and personalization of task execution, provide users with an efficient and convenient automated service experience, and at the same time ensure the security and reliability of data and revenue.
[0006] This invention provides a digital avatar task processing method based on Web3.0 and multimodal data fusion, mainly including:
[0007] The system collects browser operation data, screen interaction data, and wearable device perception data. The browser operation data includes tab creation and switching and content browsing characteristics. The screen interaction data includes code editing trajectory and interface switching habits. The wearable device perception data includes gaze point trajectory and gesture actions.
[0008] A digital agent model is generated by fusing the browser operation data, the screen interaction data, and the wearable device perception data. The fusing process includes spatiotemporal alignment and dynamic weight adjustment. The digital agent model includes personalized decision rules.
[0009] The digital agent model is deployed to the blockchain via a smart contract, which is used for task execution notarization and settlement.
[0010] Simulate user browser operations according to the personalized decision-making rules, complete web page tasks, and store the results of the web page tasks on the blockchain for evidence.
[0011] The rewards will be distributed to the user's wallet based on the results of the web page tasks.
[0012] The task template of the digital agent model is encapsulated as an NFT and traded on a blockchain marketplace. The NFT transaction triggers a copyright fee allocation to the task template owner's wallet.
[0013] Furthermore, the collection of browser operation data, screen interaction data, and wearable device perception data includes:
[0014] The browser operation data is obtained by capturing tab dwell time and content keywords in real time through a browser plugin, and the content keywords are extracted through natural language processing.
[0015] The screen interaction data is obtained by collecting the debugging operation frequency through computer vision technology and IDE hook functions, and the debugging operation frequency is associated with the operation mode;
[0016] Voice commands are acquired through the wearable device's visual and inertial sensors to obtain the wearable device's perception data, and the voice commands are associated with the environmental scene.
[0017] Furthermore, the step of fusing the browser operation data, the screen interaction data, and the wearable device perception data to generate a digital agent model includes:
[0018] The browser operation data, the screen interaction data, and the wearable device perception data are spatiotemporally aligned to construct a three-dimensional behavior mapping model, which represents the scene behavior decision relationship.
[0019] The weights of each modality data are dynamically adjusted through an attention mechanism to obtain a fusion result, which is then used to train the digital agent model.
[0020] The task execution order and priority adjustment are generated based on the personalized decision-making rules.
[0021] Furthermore, the step of obtaining the browser operation data by capturing tab dwell time and content keywords in real time through a browser plugin includes:
[0022] URL features and tab switching order are extracted from the browser operation data, and the tab switching order is used to simulate the login path;
[0023] A user preference graph is constructed using the content keywords, and the user preference graph is associated with a task priority model.
[0024] Furthermore, the step of spatiotemporally aligning the browser operation data, the screen interaction data, and the wearable device perception data to construct a three-dimensional behavior mapping model includes:
[0025] The browser operation data is aligned with the scene information to obtain the first mapping data;
[0026] The screen interaction data and the wearable device perception data are aligned based on behavioral characteristics to obtain second mapping data;
[0027] The first mapping data and the second mapping data are fused to generate the three-dimensional behavior mapping model, which is used for decision rule iteration.
[0028] The scene adaptation strategy is obtained from the three-dimensional behavior mapping model.
[0029] Furthermore, deploying the digital agent model to the blockchain via a smart contract includes:
[0030] The parameters of the digital agent model are trained through federated learning, and the parameters of the digital agent model are stored on the blockchain for evidence.
[0031] The personalized decision-making rules are encapsulated for task splitting and execution.
[0032] Furthermore, the digital agent model simulates user browser operations based on the personalized decision-making rules to complete web page tasks, including:
[0033] The digital agent model automatically completes account login based on the tab switching order;
[0034] The digital proxy model generates content based on the code editing trajectory, and the content is stored on the blockchain as a hash value.
[0035] The smart contract invokes an acceptance tool to check the results of the web page task.
[0036] Furthermore, the step of encapsulating the task template of the digital agent model into an NFT for trading on the blockchain marketplace includes:
[0037] The task execution template is extracted from the digital agent model and encapsulated into the NFT;
[0038] When the NFT is traded, the smart contract deducts copyright fees, which are then distributed to the user's wallet.
[0039] The technical solutions provided by the embodiments of the present invention have the following beneficial effects:
[0040] This invention discloses a personalized digital avatar system based on multimodal data fusion. By collecting browser operation data, screen interaction data, and wearable device perception data, it constructs a user behavior habit model to solve the adaptation problem of personalized decision-making and automated operation in web page task execution. This invention dynamically adjusts the weights of each modality of data through spatiotemporal alignment and attention mechanisms to achieve efficient data fusion. Based on this, a digital agent model is trained to generate personalized task decision rules. Simultaneously, blockchain smart contracts are used to complete task notarization and profit distribution, ensuring transparency and security. Of particular note is that this invention integrates user preference graphs, operation trajectories, and scenario adaptation strategies into model training, enabling the digital agent model to accurately simulate user behavior to complete tasks and realize value transfer through blockchain market transactions of task templates. Ultimately, this invention significantly improves the intelligence and personalization of task execution, providing users with an efficient and convenient automated service experience while ensuring the security and reliability of data and profits. Attached Figure Description
[0041] Figure 1 This is a flowchart of a digital avatar task processing method based on Web3.0 and multimodal data fusion according to the present invention.
[0042] Figure 2 This is a schematic diagram of a digital avatar task processing method based on Web3.0 and multimodal data fusion according to the present invention.
[0043] Figure 3 This is another schematic diagram of a digital avatar task processing method based on Web3.0 and multimodal data fusion according to the present invention.
[0044] Figure 4 This is another schematic diagram of a digital avatar task processing method based on Web3.0 and multimodal data fusion according to the present invention.
[0045] Figure 5 This is another schematic diagram of a digital avatar task processing method based on Web3.0 and multimodal data fusion according to the present invention.
[0046] Figure 6 This is another schematic diagram of a digital avatar task processing method based on Web3.0 and multimodal data fusion according to the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0048] like Figure 1-6 As shown, this embodiment of the invention provides a digital avatar task processing method based on Web3.0 and multimodal data fusion, which specifically includes:
[0049] S1, Collect browser operation data, including tab creation, switching, and content browsing features. The browser operation data is used to extract user behavior habits.
[0050] User operation data is acquired from the browser, including tab creation time, switching history, and specific information about browsed content. A lightweight plugin records these operation traces in real time, forming an initial operation log dataset. Feature extraction is performed on this initial dataset, including tab dwell time, switching frequency, and keyword information in the content. These features are then organized into a preliminary user operation habit profile using pre-established classification rules. Further patterns in user operation habits are extracted from this preliminary profile, including user preferences for certain content types and their operational order tendencies within specific time periods. By comparing profile data from different time periods, a dynamic trend of user operation habits is constructed. Continuous data updates are performed based on these dynamic trends. The update process involves the plugin acquiring new operation log datasets in real time and comparing and integrating them with existing trend data to ensure that the final user operation habit profile accurately reflects the true characteristics of browser operation data.
[0051] In one embodiment, obtaining user action-related data from the browser can be achieved through a lightweight plugin installed on the browser. This plugin can immediately record the creation time when a user opens a new tab. For example, if the plugin detects that a user created a tab to search for programming tutorials at 9:00 AM, it can also capture switching records, such as switching from that tab to an email page, as well as specific information about the content viewed, such as code snippets and keywords appearing on the page. The advantage of doing this is that it can accumulate raw data in real time, avoid data loss, and ensure the basic integrity of subsequent processing.
[0052] Specifically, this acquisition process is similar to the working principle of a logger, which captures traces by listening to browser API events to form an initial operation log dataset. This dataset is like a time-series archive, which helps to track the continuity of user behavior, thereby providing reliable input for habit analysis. Its beneficial effects are to improve the timeliness and accuracy of data and reduce errors caused by human intervention.
[0053] In one embodiment, when extracting features from the initial operation log dataset, a simple statistical method can be used to calculate the dwell time of tabs. For example, if a tab takes 5 minutes to open and close, this dwell time is extracted as a feature. At the same time, the switching frequency is calculated, such as 10 times per hour, and keyword information in the content is identified by a text parsing tool, such as terms like "JavaScript". Then, these features are organized into a preliminary profile of user operation habits through pre-established classification rules. These classification rules are based on threshold grouping of common behavioral patterns. For example, tabs with a dwell time of more than 3 minutes are classified as high-interest content. The benefit of doing this is that it can transform scattered data into a structured profile, which is convenient for further extracting patterns and improving the system's analysis efficiency.
[0054] Specifically, the principle behind this extraction process lies in data dimensionality reduction and induction. It transforms massive logs into a manageable feature set, and its beneficial effects include enhancing the targeting of user profiles. For example, it can identify users' preferences for front-end technologies in development scenarios, thereby supporting personalized recommendations.
[0055] In one embodiment, further refining the patterns of user operating habits from the initial profile can be achieved through time-segmented comparisons. For example, comparing profile data from the morning and afternoon reveals that users have a higher preference for technical document types in the morning, as well as a tendency in the order of operations, such as always browsing first and then switching to the editor. This constructs a dynamic trend, which is like a behavior curve reflecting the evolution of habits. Its beneficial effect is to capture the dynamism of user behavior and avoid the limitations of static profiles.
[0056] Specifically, this extraction principle is based on pattern matching and comparative analysis. It groups profile data and calculates the differences to form trend data, which helps predict future behavior. For example, in task automation, it can predict the user's next action and improve the accuracy of execution.
[0057] In one embodiment, continuous data updates for dynamic trends are accomplished by the plugin acquiring new operation log datasets in real time. For example, when a user opens a new tab, the plugin immediately records it and compares it with existing trend data. If the new data shows an increase in dwell time, the trend is merged and updated to ensure that the final user operation habit profile accurately reflects the true characteristics of browser operation data. The beneficial effect of doing so is to maintain the real-time nature and adaptability of the profile, adapting to the natural changes in user habits.
[0058] Specifically, the principle of this update process is similar to incremental learning. It integrates new and old data through comparison and fusion mechanisms. Its beneficial effects include improving the robustness of the system, such as continuously optimizing the profile in long-term use, and ultimately achieving accurate extraction and utilization of user behavior habits.
[0059] S2, collect screen interaction data, including code editing trajectory and interface switching habits, and associate and integrate the screen interaction data with the browser operation data.
[0060] This study acquires code editing trajectory data from screen interaction behavior, recording the user's key press sequences, text modification paths, and cursor movement trajectories within the editing environment. Simultaneously, it collects interface switching habit data, recording the user's jump order and dwell time between different windows or tools. The code editing trajectory data and interface switching habit data are initially organized, categorizing key press sequences and text modification paths as editing behavior features, and jump order and dwell time as interaction preference features, forming a structured behavior dataset. Tab switching order and access duration data are obtained from browser operation logs, organized into webpage behavior features, and correlated and matched with the aforementioned editing behavior features and interaction preference features. These are then merged into a unified set of user operation behaviors using timestamp alignment. Feature extraction is performed on the merged set of user operation behaviors, extracting recurring behavior patterns across scenarios to ensure the correlation between screen interaction data and browser operation data is preserved for accurate mapping in subsequent behavior simulations.
[0061] For example, when obtaining code editing trajectory data from screen interaction behavior, user input can be captured in real time through hook functions integrated into the editing environment. For instance, the path of code block indentation adjustment after the user presses the Enter key, as well as the continuous trajectory of the cursor moving from the variable declaration to the function call, can be recorded. This can lead to more accurate capture of user habits, which is beneficial to the realism of subsequent behavior simulation. At the same time, when collecting interface switching habit data, the order in which the user jumps from the code window to the debugging tool can be monitored, such as switching to the console to view errors before returning to the editing area, and the dwell time of each interface can be recorded. This can reveal the user's task flow preferences and is beneficial to building a coherent operation model.
[0062] In one possible implementation, when initially organizing code editing trajectory data and interface switching habit data, key press sequences can be categorized as continuous input features. For example, paths involving multiple deletions and rewrites can be considered as editing iteration patterns, while text modification paths can be classified as structural adjustment features. This simplifies data processing and facilitates the efficient formation of structured behavioral datasets. When categorizing jump order as interaction preference features, dwell time can be analyzed to identify high-frequency switching paths. For example, users spending a longer time on the test interface indicates a preference for debugging. This enhances the organization of the dataset and improves the accuracy of subsequent fusion.
[0063] For example, when obtaining tab switching order and access duration data from browser operation records, the sequence of users jumping from the homepage tab to the document page can be tracked. For example, the timestamps before and after the switch can be recorded and organized into webpage behavior features. This can capture web interaction habits and is beneficial for correlation with screen data. When correlating and matching with the aforementioned editing behavior features and interaction preference features, timestamp alignment methods can be used for fusion. For example, the time period of browser access to the document can be correlated with the time period of code editing on the screen to form a unified set of user operation behaviors. This can achieve cross-source data unification and is beneficial for comprehensive behavior analysis.
[0064] In one possible implementation, when extracting features from the fused set of user actions, recurring behavior patterns can be extracted. For example, identifying cross-scene loops where code is modified on the screen immediately after viewing a reference in the browser can highlight key patterns and help preserve the correlation between screen interaction data and browser operation data for accurate mapping in subsequent behavior simulations. For example, these patterns can be directly applied to generate operation sequences that conform to user logic during simulation tasks. This can improve the adaptability of the simulation and benefit the overall system reliability.
[0065] S3, Collect wearable device perception data, including gaze trajectory and gestures, and align the wearable device perception data with the browser operation data and the screen interaction data in time and space.
[0066] Perceptual data, including gaze trajectory and gesture data, is acquired from wearable devices. This data is recorded using timestamps and stored as an initial perceptual dataset. Simultaneously, user click and scrolling behavior data on web pages is extracted from browser operation records to form an initial operation dataset. Spatiotemporal alignment processing is performed on both the initial perceptual and initial operation datasets, matching them based on timestamps to ensure that the gaze trajectory and gesture data correspond to the web page click and scrolling behavior data in the time dimension, forming an aligned fused dataset. Based on the fused dataset, the changes in user gaze trajectory and the frequency features of gestures within a specific time period are extracted. Combined with the contextual information of web page operations, a user behavior association mapping table is constructed for subsequent judgment of user intent and operating habits. The correspondence between wearable device perceptual data and browser operation data is updated in real time for the user behavior association mapping table to ensure consistency between gaze trajectory and gestures and screen interaction behavior in different scenarios, achieving precise coordination between wearable device perceptual data and browser operation data.
[0067] In one possible implementation, when acquiring sensory data from wearable devices, the user's gaze trajectory can be captured in real time using built-in sensors. For example, when a user is browsing a webpage while wearing smart glasses, the device records the path data of the eye's focus moving from the top to the bottom of the page, while simultaneously capturing hand gestures such as swiping or tapping the air. This data is stored as an initial sensory dataset in the form of timestamps. The advantage of this approach is that it captures the user's natural behavior and avoids data loss. At the same time, click and scrolling behavior data can be extracted from the browser's operation records, such as when a user clicks a product link or scrolls down to view reviews on an e-commerce page, forming an initial operation dataset. This provides a complete record of the behavioral chain, which is beneficial for subsequent fusion analysis.
[0068] When performing spatiotemporal alignment processing on the initial perception dataset and the initial operation dataset, the two sets of data can be matched based on the same timestamp. For example, if the gaze trajectory shows that the user gazes at a button at 10:00:05, and the browser data records the click behavior at the same time, then they can be matched to form an aligned fused dataset. The benefit is to ensure data synchronization and avoid analysis errors caused by time deviation. This processing can bring more accurate inference of user intent.
[0069] In one possible implementation, when extracting gaze trajectory changes and gesture frequency features based on the fused dataset, it is possible to analyze the changes in gaze from static to rapid movement within a specific time period and calculate the number of gestures. Combined with the webpage operation context, such as page loading status, a user behavior association mapping table can be constructed. For example, the mapping table can associate gaze trajectory with click habits to determine whether the user is interested in the content. The beneficial effect is to improve the accuracy of behavior simulation, which can bring reliability to automated task processing.
[0070] When updating the user behavior association mapping table in real time, the consistency between the gaze trajectory and screen interaction can be adjusted in different scenarios, such as mobile browsing or static reading. For example, the mapping can be updated when the frequency of gesture actions increases, ensuring that the perception data and operation data are coordinated to achieve precise coordination. The beneficial effect is that it can adapt to diverse scenarios and improve the task execution efficiency of the digital clone. Doing so can bring a transparent benefit distribution mechanism.
[0071] S4. A digital agent model is trained based on the fusion result of the browser operation data, the screen interaction data, and the wearable device perception data. The digital agent model generates personalized task decision rules.
[0072] Data on user tab switching frequency and dwell time are obtained from browser operation logs. Simultaneously, user operation sequence and preference habits during screen interactions are collected. Combined with environmental interaction actions perceived by wearable devices, a multi-source behavior dataset is formed. This dataset contains user operation tendencies and habitual features in different scenarios, used for subsequent personalized rule construction. For this dataset, a pre-established feature extraction tool is used to classify and label the data, generating user behavior preference vectors. These vectors encompass user habitual patterns in web browsing, screen operation, and environmental interaction, providing foundational data support for subsequent rule generation through vectorization. The behavior preference vectors are input into a pre-set machine learning tool, and a decision rule base for the digital avatar is constructed through multiple iterations of training. This decision rule base contains personalized execution logic for different task scenarios, covering user preference responses in various operating environments, ensuring rule adaptability. The decision rule base is periodically updated by extracting new content from the multi-source behavior dataset, dynamically adjusting the rule base. This adjustment process optimizes the personalized task decision rules of the digital agent model by comparing the differences between new and old behavior preference vectors, maintaining consistency with actual user habits.
[0073] Data on user tab switching frequency and dwell time are obtained from browser operation logs. At the same time, user operation sequence and preference habits during screen interaction are collected. Combined with environmental interaction actions perceived by wearable devices, a multi-source behavior dataset is formed.
[0074] Specifically, this acquisition process involves real-time monitoring of browser log files. The tab switching frequency is obtained by counting the number of switches per unit time, while the dwell time is calculated by recording the opening and closing timestamps of each tab. This approach can capture the user's level of interest in specific content and helps improve the accuracy of the dataset.
[0075] In one embodiment, when a user frequently switches to the development documentation tab, the system records the switching frequency as twice per minute and combines this with the order in which the user clicks the code editor on the screen to form an operational tendency characteristic.
[0076] For example, users often open the console before editing code during screen interactions. This habit is collected as a preference sequence and combined with gestures detected by wearable devices such as AI glasses. For instance, a user's gesture to zoom in on an image during a meeting is converted into environmental interaction data, thus forming a complete multi-source behavioral dataset encompassing various scenarios. This dataset is beneficial for subsequent processing because it integrates static and dynamic elements, ensuring the comprehensiveness of behavioral characteristics.
[0077] For the aforementioned multi-source behavior dataset, a pre-established feature extraction tool is used to classify and label the data, generating user behavior preference vectors. This feature extraction tool is a software component based on a standard data processing library. It first classifies the operational data in the dataset into three categories: browser, screen, and wearable devices. Then, it labels each category with tags such as "high-frequency switching" or "sequence preference," and transforms these labels into numerical arrays using vector representation. For example, the switching frequency of browser data is labeled as the first dimension in the vector, the screen operation sequence as the second dimension, and the frequency of wearable actions such as gestures as the third dimension. The resulting vectors cover users' dwell preferences during web browsing, their debugging habits during screen operations, and their visualization tendencies in environmental interactions.
[0078] In one embodiment, if the dataset shows that users spend a long time in the browser looking at the front-end document, the tool will label it as "development preference" and quantize it as [0.8, 0.5, 0.3]. This vectorization is beneficial for quantifying habitual patterns and provides a computable foundation for subsequent rule generation.
[0079] For example, from multiple perspectives, when screen data is labeled "log check first", the vector integrates the "gesture priority" label of wearable data to form a complementary preference representation. This helps to improve the robustness of the vector because the support of different data sources ensures the comprehensiveness and consistency of the vector.
[0080] The behavioral preference vector is input into a preset machine learning tool, and a decision rule library for the digital clone is constructed through multiple iterations of training.
[0081] Specifically, machine learning tools are a standard training framework that uses gradient descent optimization. It takes a vector as input and adjusts the parameters during the iteration process to minimize the prediction error. For example, in training, the vector [0.8, 0.5, 0.3] is used to learn rules for task scenarios such as code debugging, generating logic such as "if the switching frequency is high, prioritize the development task". The rule base built in this way covers content generation in the browser environment, debugging execution under screen interaction, and reporting mode under wearable perception.
[0082] In one embodiment, iterative training involves multiple loops, each time updating the rules to match vector differences, such as adjusting from an initial rule to a personalized version. This benefits the adaptability of the rule base because it ensures that preference responses are covered across a variety of operating environments.
[0083] For example, from the perspective of task automation, when vectors represent a high preference for visualization, the rule base will generate the logic of "reporting using a 3D interface", which will be combined with the "content generation" rules supported by browser data to form a consistent execution chain.
[0084] For the aforementioned decision rule base, updated content is periodically extracted from new multi-source behavioral datasets to dynamically adjust the rule base. The adjustment process first compares the newly collected dataset with the old vectors. For example, the difference between the new vector [0.7,0.6,0.4] and the old [0.8,0.5,0.3] is quantified by calculating Euclidean distance. Based on this, the rules are optimized. For instance, the "log check before debugging" rule is adjusted to a version that focuses more on gestures. This optimized rule maintains consistency with user habits, which is beneficial for the digital agent model to generate personalized task decision rules.
[0085] In one embodiment, periodic extraction involves obtaining new data from the browser and wearable device weekly, forming an update cycle. For example, when the difference comparison shows a decrease in dwell time, the rule base will be adjusted to more efficient execution logic, supporting the continuity of the model from multiple directions. The complementary updates of browser frequency, screen order, and wearable actions ensure the dynamic adaptation of rules, ultimately focusing on training the digital agent model to generate personalized task decision rules.
[0086] S5, the parameters of the digital agent model are deployed to the blockchain through a smart contract, which is used for task execution notarization and settlement.
[0087] User behavior-related data is acquired from a multimodal data layer. This data includes browser tab browsing duration, screen operation trajectories, and gesture habits collected through smart glasses. This data is initially organized to form a structured behavior record dataset. This structured behavior record dataset is then used for distributed training using a pre-established federated learning tool to generate personalized digital avatar parameters. These parameters reflect the user's operational preferences and decision-making logic in different scenarios. The digital avatar parameters are encoded and deployed in the form of smart contracts, which are uploaded to a blockchain network to record operation logs and hash values of relevant data during task execution, ensuring data immutability. Within the blockchain network, the smart contracts are further configured as the settlement basis after task execution. The revenue distribution logic is automatically triggered based on task completion status, and relevant records are associated and stored with the digital avatar parameters to support subsequent task execution notarization and settlement needs.
[0088] For example, when acquiring user behavior-related data from the multimodal data layer, browser plugins can be used to capture tab access durations in real time. For instance, the time a user spends viewing programming documentation can be recorded as a behavioral indicator. This ensures the real-time nature of data capture, thereby providing accurate basic information for subsequent modeling and improving the personalization of the digital avatar.
[0089] In one possible implementation, the initial organization of screen operation trajectories involves classifying scattered mouse movement and click events into sequence patterns, such as organizing frequent code debugging paths into a structured dataset. This organization helps reduce data noise, makes behavior records easier to process, thereby improving training efficiency and resulting in more reliable proxy parameter generation.
[0090] Specifically, gesture habits such as zooming actions collected through smart glasses can be incorporated into the dataset and fused with browser data to form a comprehensive view of user interaction. This enhances the model's adaptability to the physical environment and improves the accuracy of decision-making in diverse scenarios for digital clones.
[0091] In one possible implementation, distributed training is performed using a pre-established federated learning tool, which is a mechanism that allows multiple devices to jointly optimize the model without sharing the original data. For example, a behavior prediction sub-model can be trained locally on a user device, and then only gradient updates can be uploaded to a central server for aggregation. This can protect privacy and reduce data transmission overhead, and is beneficial for generating parameters that reflect operational preferences, while ensuring the security and efficiency of the training process.
[0092] Specifically, the generated personalized digital avatar parameters include decision weight vectors, which can simulate the user's log checking habits in development tasks. This makes the agent more in line with real logic, resulting in improved accuracy of automated execution.
[0093] In one possible implementation, when digital avatar parameters are encoded and deployed in the form of smart contracts, which are code protocols that are automatically executed on the blockchain, such as converting parameters into executable contract terms and uploading them, decentralized storage of parameters can be achieved, which helps prevent tampering and ensures transparency in task execution.
[0094] Specifically, recording operation logs and hash values during task execution involves hashing each proxy action, such as content generation sequence, and putting it on the blockchain. This provides undeniable evidence, strengthens the trust mechanism, and supports subsequent auditing.
[0095] In one possible implementation, when configuring smart contracts as the basis for settlement in a blockchain network, for example, automatically checking conditions and triggering token transfers after task completion such as code file generation, the revenue can be distributed instantly, which is beneficial for incentivizing user participation and maintaining the fairness of the ecosystem.
[0096] Specifically, storing relevant records in association with digital clone parameters involves embedding parameter reference links in the contract, such as binding execution logs with preference parameters. This supports the integrity of the evidence, improves the automation and accuracy of the settlement process, and thus enhances the overall trustworthy task execution capability of the system.
[0097] S6, the digital agent model simulates user operations according to the task decision rules to complete the web page task. After the web page task results are stored on the blockchain, the smart contract is triggered to execute the revenue distribution.
[0098] The digital agent model first obtains operation path data from the user's historical browser operation records. It then analyzes the login process and content editing habits layer by layer to construct a serialized template of user operation habits for automated execution of subsequent tasks. Based on this serialized template, the digital agent model simulates user behavior according to the pre-established operation path when executing web page tasks, completing repetitive actions such as login and content generation, and recording key node data during task execution to form a task execution log. The digital avatar module organizes the task result data and stores it on the blockchain through a blockchain interface to ensure data immutability. It also generates a storage identifier for use in determining trigger conditions for subsequent processes. After the storage identifier is generated, the digital avatar module reads the identifier through a preset smart contract interface. If the profit distribution conditions are met, the contract logic is automatically triggered to achieve transparent profit distribution, ensuring a direct link between web page task results and profit distribution.
[0099] For example, when the digital agent model obtains operation path data from the user's historical browser operation records, it can perform layer-by-layer parsing of the login process of the Xiaohongshu platform. For example, it can first identify the positions of fields where users commonly input usernames and passwords, and then analyze the jump logic after clicking the login button. The serialized template constructed in this way can accurately reproduce user habits, avoid error interruptions in automated execution, and help improve the reliability of task completion.
[0100] In one possible implementation, this parsing process involves decomposing the operation path data into a sequence of nodes, each node corresponding to a user action, such as a mouse click or keyboard input. The templates formed in this way not only support subsequent automation but also adapt to changes in different webpage layouts, thereby ensuring that the digital clone is highly adaptable when simulating behavior and bringing about a significant improvement in task execution efficiency.
[0101] Specifically, the analysis of content editing habits can record the order in which users insert phrases and the steps for adjusting format when editing tweets, such as adding images first and then entering text. In this way, the template can automatically generate content that matches the user's style, reducing the need for manual intervention and facilitating transparent task processing in offline mode.
[0102] In one possible implementation, when executing web page tasks based on serialized templates, the digital avatar module simulates user behavior according to a pre-established operation path. For example, after logging into Xiaohongshu, it automatically navigates to the editing page, generates a tweet about sharing life experiences, and records key nodes such as the login success time and content submission time, forming a task execution log. This simulation not only replicates the user's real operation logic but also captures potential anomalies through log recording, such as failed attempts caused by network latency, thus facilitating subsequent optimization and enhancing the system's robustness.
[0103] For example, if the task involves content generation, the module can insert preset keywords based on the template to ensure that the generated tweets are consistent with the user's historical style. This can improve the quality of the task results and thus support the fairness of the distribution of benefits.
[0104] Specifically, after organizing the task execution logs and collecting the task results data, the process of uploading the data to the blockchain for verification via a blockchain interface allows the results, such as generated tweet links and completion proofs, to be packaged into transaction blocks and uploaded to the Ethereum network. This ensures the data is immutable and generates a unique verification identifier, such as a hash value, used to trigger judgments. This verification mechanism prevents tampering through distributed ledger technology, improves data transparency, and is beneficial for building a trustworthy environment for revenue transfer.
[0105] In one possible implementation, if the output data includes multimodal elements such as text and images, the module will verify its integrity before uploading it to the blockchain. This not only ensures the accuracy of the evidence but also allows for deep integration with Web3.0 technologies, supporting automated processing of offline tasks.
[0106] For example, after the notarized token is generated, the process of reading the token through a preset smart contract interface will automatically trigger contract logic if the token matches predefined profit distribution conditions, such as achieving the task completion target. This logic might involve calling a function written in Solidity to distribute tokens to the user's account, thus achieving transparent distribution. This triggering ensures a direct link between the results of the web task and the profits, which is beneficial for incentivizing user participation and achieving the efficiency of automated settlement.
[0107] Specifically, the contract logic can include conditional judgment branches. If the evidence shows that the task results are recognized by the platform, then the transfer operation is executed. This not only makes the flow of income transparent, but also reduces the risk of human intervention and supports the credibility of the overall mechanism.
[0108] S21, the browser operation data includes tab dwell time and content keywords, the content keywords are used to construct a user preference graph through natural language processing, and the user preference graph is associated with the priority adjustment of the task decision rules.
[0109] User activity data, including tab dwell time and keywords in page content, is obtained from browser tabs. These keywords are initially extracted, and their frequency and contextual semantics are recorded to form an initial keyword set. Natural language processing (NLP) tools are then used to perform semantic analysis on this initial keyword set. Keywords with similar semantics are categorized and their associated domains are labeled, forming a structured user preference graph. This graph contains the distribution of user interests across different domains. Weight data of the interest distribution is extracted from the user preference graph, and this weight data is sorted and classified to determine the user's focus during task execution, forming a priority adjustment basis for dynamic updates to subsequent task decision rules. This priority adjustment basis is then correlated with task decision rules. Based on the matching degree between task type and the user preference graph, the order of task execution and resource allocation are adjusted to ensure that the task decision rules reflect the user's true preferences.
[0110] In one possible implementation, obtaining user action data from browser tabs can effectively capture daily browsing habits.
[0111] For example, when a user opens a programming tutorial page in a browser, the system records the duration of the tab stay, such as longer reading times. This helps to identify the depth of the user's interest in a specific topic. At the same time, it extracts keywords from the page content, such as "variable declaration" or "loop structure." By initially extracting and recording the frequency of these keywords and their contextual semantics, an initial keyword set is formed. This can bring more accurate insights into user behavior and is beneficial to the accuracy of subsequent personalized processing.
[0112] For example, in practical applications, if users frequently browse front-end development-related pages, the system will perform preliminary keyword extraction after obtaining operation data from these tabs. For instance, it will record the frequency of the keyword "JavaScript" being associated with "function optimization" in the context. This not only supports the basic accumulation of data but also provides reliable input for building a preference graph, thereby improving the overall efficiency of task adaptation.
[0113] In one possible implementation, the process of performing semantic analysis on an initial set of keywords using natural language processing tools involves classifying keywords with similar semantics.
[0114] For example, categorizing "algorithm efficiency" and "time complexity" into the computation domain and labeling their related domains such as programming optimization creates a structured user preference map. This map contains the distribution of users' interests in different domains, such as a high proportion in the programming domain. This helps the system understand users' multi-dimensional preferences, resulting in task execution that is more in line with actual needs.
[0115] For example, if a user is interested in machine learning, semantic analysis will categorize keywords such as "neural network" and "deep learning" into the AI field, forming a graph that shows the distribution of interests biased towards technological innovation. This processing makes the graph more structured, facilitates subsequent weight extraction, and ensures that the decision-making rules are more targeted.
[0116] In one possible implementation, weighted data of interest distributions are extracted from the user preference graph, and then sorted and classified to determine the focus of attention during task execution.
[0117] For example, sorting and classifying weighted data in the programming domain into high-priority categories forms the basis for priority adjustment and is used for dynamic updates of task decision rules. This enables real-time optimization of rules and is beneficial for improving the responsiveness of automated tasks and user satisfaction.
[0118] For example, in a task management scenario, if the graph shows that users have a high weight in their interest in data analysis, the system will sort and classify these weighted data to determine the focus, such as the use of analysis tools, thus forming the basis for adjustments. This supports the logical chain of rule updates and avoids deviations in task execution.
[0119] In one possible implementation, when prioritization is linked to task decision rules, the execution order and resource allocation are adjusted based on the degree of matching between task type and user preference graph.
[0120] For example, for development tasks, resources are allocated preferentially if the matching degree is high, ensuring that the rules reflect the user's true preferences. This brings about an intelligent improvement in task processing and strengthens the overall system's adaptability.
[0121] For example, in content generation tasks, the system adjusts the order based on the degree of graph matching, such as prioritizing front-end development tasks and allocating more computing resources. This ensures that the decision-making rules are consistent with preferences, resulting in more efficient execution and improved user experience.
[0122] S22, the screen interaction data includes the frequency of debugging operations and the editing trajectory collected by the IDE tool hook, and the editing trajectory is converted into the task execution order rules of the digital agent model.
[0123] The debugging operation frequency and code editing trajectory are obtained from screen interaction data. The debugging operation frequency is determined by recording the user's operation intervals and frequency in the integrated development environment (IDE). The code editing trajectory is captured by hook functions of the IDE tools, capturing the user's cursor movement path and text modification records during the editing process, forming initial behavior data records. These initial behavior data records are then structured, organizing the debugging operation frequency and code editing trajectory into behavior sequence data in chronological order. This behavior sequence data contains the user's operational habits and order under specific tasks, forming an ordered chain of operational behaviors. Priority rules for task execution are extracted from this ordered chain of behaviors. These priority rules are determined by analyzing the sequential relationship and frequency of operations in the behavior sequence data, such as the order of editing actions before and after debugging, forming a preliminary mapping of task execution order rules. This preliminary mapping of task execution order rules is embedded into a digital agent model. This mapping serves as the behavioral guideline for the digital agent model when executing tasks, ensuring that the digital agent model can execute tasks in the order that the user is accustomed to, thus completing the transformation from screen interaction data to task execution order rules.
[0124] In one possible implementation, the process of obtaining debugging operation frequency and code editing trajectory from screen interaction data can be achieved through the logging mechanism within the integrated development environment. The debugging operation frequency specifically refers to the occurrence rate of actions such as setting breakpoints or checking variables during the code debugging phase. This rate is obtained by dividing the cumulative number of operations by the time period. The code editing trajectory uses hook functions to track cursor position changes and text insertion and deletion actions in real time, forming initial behavior data records. The advantage of doing this is that it can accurately capture user operation habits, provide a reliable foundation for subsequent processing, and avoid model deviations caused by data omissions.
[0125] Specifically, when structuring initial behavioral data records, the frequency of debugging operations and code editing trajectories can be organized into behavioral sequence data in chronological order using a timestamp sorting algorithm. This ensures that each operation point is associated with the preceding and following actions, forming an ordered chain of operational behaviors. The beneficial effect of this processing is that it transforms scattered data into a coherent sequence, making it easier to identify the user's habitual patterns in tasks, such as the tendency to edit immediately after frequent debugging, thereby improving the accuracy of the digital agent model in simulating real behavior.
[0126] For example, priority rules for task execution can be extracted from an ordered chain of operations by analyzing the order and frequency of operations. For instance, one can examine the proportion of debugging actions that usually occur before editing in the action sequence data. If the proportion is high, debugging rules can be set first. The beneficial effect of this extraction is that it generates a preliminary mapping of rules that conforms to user logic, helps the digital agent model avoid the inefficiency caused by random execution, and ensures that the task flow is closer to actual operating habits.
[0127] In one possible implementation, the initial mapping of task execution order rules can be embedded into the digital agent model through parameter injection. The mapping serves as a behavioral guideline, ensuring that the model follows the user's order during task execution. The benefit of this approach is that it completes the transformation of screen interaction data into task execution order rules, enabling the digital agent model to efficiently simulate user behavior and improve the reliability and consistency of automated tasks.
[0128] S23, the wearable device perception data includes environmental information collected by a visual sensor and gesture data collected by an inertial sensor, and the gesture data is mapped to the scene adaptation strategy of the digital agent model.
[0129] Environmental information collected by a visual sensor and gesture data collected by an inertial sensor are acquired from wearable devices. The environmental information includes scene features around the user, and the gesture data records the trajectory and direction of the user's hand movements. The acquired environmental information and gesture data undergo preliminary cleaning to remove noise and invalid records, forming a refined set of environmental features and a set of gesture trajectories. For these refined sets, a behavior mapping relationship for the user in a specific scenario is constructed. The direction and amplitude of movements in the gesture trajectory set are associated with scene elements in the environmental feature set, generating a scenario-based behavior matching table for subsequent adaptation processing. Based on the scenario-based behavior matching table and a pre-established user habit database, priority correspondence rules between gesture data and scene elements are determined. If the correlation between a gesture trajectory and a certain scene element is higher than a preset threshold, the gesture data is mapped to a scene adaptation instruction for the digital avatar, forming a preliminary adaptation strategy draft. This preliminary adaptation strategy draft is dynamically adjusted based on the real-time acquired environmental feature set to ensure that the instructions mapped from the gesture data remain consistent with the current scene, ultimately generating the scene adaptation strategy for the digital agent model.
[0130] Specifically, in the process of acquiring environmental information collected by visual sensors and gesture data collected by inertial sensors from wearable devices, the environmental information can be understood as the scene features around the user, such as the indoor layout or object positions captured by AI glasses. These features help to identify the specific type of environment in which the user is currently located, thereby providing basic data support for subsequent processing. This can improve the accuracy of data processing because the cleaned set of environmental features and gesture trajectory formed after the initial cleaning to remove noise and invalid records will reduce interference factors, ensure higher data quality, and benefit the reliability of the overall simulation.
[0131] In one possible implementation, when constructing a user behavior mapping relationship in a specific scenario based on the organized set of environmental features and the set of gesture trajectories, the direction and amplitude of the action in the set of gesture trajectories are associated with scene elements in the set of environmental features. For example, the direction of the user's waving action is associated with the position of the door in the scene, generating a scenario-based behavior matching table. This matching table is actually a structured correspondence list used to record how gestures respond to environmental elements. This approach can bring stronger adaptability because it transforms isolated sensor data into coherent behavioral logic, which is beneficial for the accurate response of the digital clone when simulating user habits.
[0132] For example, when determining the priority correspondence rules between gesture data and scene elements based on a scenario-based behavior matching table and a pre-established user habit database, if the correlation between the gesture trajectory and a certain scene element is higher than a preset threshold, the gesture data is mapped to the scene adaptation command of the digital clone, forming a preliminary adaptation strategy draft. For example, the user habit database stores past gesture preference data, such as a quick wave corresponding to an emergency task switch. The priority correspondence rule here is a rule that sorts by comparing correlation. The calculation process involves evaluating the matching score between the gesture amplitude and the scene element. If the score exceeds the threshold, it is mapped first. This can optimize the efficiency of command generation and help ensure that the adaptation strategy draft is more in line with the user's real operation logic and avoid interference from low correlation.
[0133] In one possible implementation, when dynamically adjusting the initial draft adaptation strategy based on a set of environmental features acquired in real time, it ensures that the instructions mapped from the gesture data remain consistent with the current scene, ultimately generating a scene adaptation strategy for the digital agent model. For example, if the real-time environmental features show that the scene changes from indoors to outdoors, the adjustment will update the instructions to adapt to the new elements, such as changing the way gestures trigger tasks. This dynamic adjustment process is achieved by continuously comparing the differences between real-time data and the draft, first identifying inconsistencies, and then modifying the instruction parameters. This approach brings higher real-time performance and helps the digital avatar maintain efficient simulation in changing environments, ultimately adhering to the goal of mapping gesture data to a scene adaptation strategy for the digital agent model.
[0134] S31, the fusion result dynamically adjusts the weights of each modality data through an attention mechanism, and the weights are used for real-time iterative updates of the digital agent model.
[0135] This system acquires multi-modal information about users from multi-source behavioral data, including data on action trajectories, voice tone, and interaction frequency. This data undergoes initial cleaning and format standardization to ensure seamless integration during subsequent processing. The cleaned multi-source behavioral data is input into a pre-established attention mechanism module. This module dynamically evaluates the contribution of each modality, generating a corresponding weight allocation scheme to ensure the influence of each data type can be adjusted according to the real-time scenario. The generated weight allocation scheme is used to weight and fuse the multi-source behavioral data, forming a comprehensive behavioral feature representation. This feature representation captures subtle changes in user habits, ensuring the accuracy of subsequent processing. The comprehensive behavioral feature representation is applied to the real-time iterative updates of the digital avatar, continuously adjusting the weight ratios of each modality based on the fusion results to ensure the digital agent model can dynamically adapt to the needs of complex task scenarios.
[0136] In one possible implementation, the process of acquiring multiple modal information of users from multi-source behavioral data can be achieved by collecting users' daily interaction data in real time. For example, recording users' movement trajectories such as gesture movement paths, voice intonation such as changes in tone when speaking, and interaction frequency such as the number of clicks per day on smart devices can comprehensively capture the multi-dimensional characteristics of user behavior, which is beneficial for subsequent accurate simulation. If this information is not acquired, a complete user profile cannot be formed, leading to increased simulation bias.
[0137] It should be noted that the specific methods for preliminary cleaning and format standardization of these data include removing noisy data such as abnormal trajectory points and converting data of different formats into a standard structure, such as unifying trajectory data into coordinate sequences. The purpose of this step is to ensure data quality, avoid interference from messy information in the fusion process, and thus improve the accuracy of the overall simulation.
[0138] For example, when a user uses a virtual assistant, if the motion trajectory data contains invalid jitter points, the cleaned data can better reflect real habits, and the resulting technical effect is that the model can more reliably predict user preferences.
[0139] It should be noted that the cleaned multi-source behavioral data is input into a pre-established attention mechanism module. This module is a neural network-based structure used to dynamically evaluate the contribution of each modality of data. The process is to determine the importance by calculating the attention score of each modality. For example, assigning a higher score to a motion trajectory indicates that it is more critical in the current scenario. This generates a corresponding weight allocation scheme to ensure that the influence of each type of data is adjusted according to the real-time scenario, which is beneficial for adapting to changing user environments. If the weights are fixed, they cannot cope with sudden changes in behavior, resulting in an inflexible simulation.
[0140] For example, in a game scenario, if a user's voice tone suddenly becomes excited, the attention mechanism module will increase its weight, thus making the fusion more focused on emotional factors, which can bring a more realistic digital clone response.
[0141] It should be noted that the specific process of weighted fusion of multi-source behavioral data using the generated weight allocation scheme is to multiply the data of each modality by its weight and then sum them to form a comprehensive behavioral feature representation. For example, the action trajectory is multiplied by 0.4, the voice tone is multiplied by 0.3, and the interaction frequency is multiplied by 0.3 before being integrated. The purpose of this step is to capture subtle changes in user habits and ensure the accuracy of subsequent processing, because the fused features can integrate information from multiple aspects and avoid the limitations of a single modality.
[0142] For example, in office tasks, this fusion can detect a decrease in interaction frequency when the user is fatigued, and then adjust the behavior of the clone to provide rest reminders. The resulting technical effect is to improve the continuity of the user experience.
[0143] It should be noted that the process of applying the comprehensive behavioral feature representation to the real-time iterative update of the digital avatar and continuously adjusting the weight ratio of each modality data based on the fusion result is achieved through a feedback loop. For example, the weights are updated in reverse based on the error of the previous simulation, ensuring that the digital agent model can dynamically adapt to the needs of complex task scenarios. This is beneficial for long-term habit simulation because it allows the model to be gradually optimized and reduces accumulated errors.
[0144] S32, the smart contract encapsulates the task template of the digital agent model as an NFT, and the NFT automatically deducts fees and is allocated to the template owner's wallet when it is traded on the blockchain market.
[0145] Among them, NFT (Non-Fungible Token) is essentially a trusted digital equity certificate with unique characteristics in the blockchain network. It is a data object with multi-dimensional and complex attributes that can be recorded and processed on the blockchain. Its difference from traditional fungible tokens represented by a certain currency lies in the unique digital identifier and the verifiability and transparency of public accounts.
[0146] After the task template of the digital avatar is constructed, its core parameters and execution rules are hashed to generate a unique identifier. This unique identifier is bound to the task template, forming a tradable digital asset unit, which is registered as a non-fungible token (NFT) on the blockchain network to ensure its uniqueness and ownership. When the NFT is traded in the blockchain market, the transaction request first triggers a pre-deployed smart contract. The smart contract automatically calculates the fees to be deducted based on the transaction amount and a preset allocation ratio, and records the fee data in the on-chain ledger. After the fee data is confirmed in the on-chain ledger, the smart contract further executes the allocation logic, transferring the deducted fees directly to the template owner's digital wallet address according to preset rules, and simultaneously generating a transaction completion record to ensure that the flow of funds is transparent and tamper-proof. After the transaction completion record is broadcast in the blockchain network, the ownership of the NFT is transferred to the new holder. The smart contract updates the token ownership information and synchronizes the updated information to the market platform, ensuring that the digital asset unit can still trigger the fee deduction and allocation process in subsequent transactions, maintaining the continuous rights and interests of the template owner.
[0147] For example, after the task template of the digital avatar is constructed, its core parameters, such as behavior prediction rules and execution logic, are input into the SHA-256 hash function. This function generates a fixed-length unique identifier by performing a one-way hash calculation on the data. This process involves serializing the parameters into a byte array and then applying hash operations to ensure that any tiny change will produce a completely different identifier, thereby binding it to the task template to form a digital asset unit. This binding is verified by submitting a transaction to the node when registering on the blockchain network, preventing tampering and maintaining uniqueness, which is beneficial for protecting intellectual property rights and facilitating subsequent transaction tracking.
[0148] In one possible implementation, when non-fungible tokens are traded on the blockchain market, the transaction request invokes a pre-deployed smart contract. This contract is executable code written in the Solidity language and deployed on the Ethereum network. When a transaction occurs, the contract first reads the transaction amount, such as 0.5 ETH, and applies a preset allocation ratio, such as 10%. It then automatically calculates and deducts a fee, such as 0.05 ETH, through multiplication. This fee data is then written as part of the transaction into the on-chain ledger. This ledger is an immutable database that records all transactions in a distributed ledger, which is beneficial for achieving transparent auditing and preventing double-spending.
[0149] For example, after the fee data is confirmed, the smart contract executes the allocation logic, transferring the deducted fees directly from the buyer's wallet to the template owner's wallet address. This logic includes checking the validity of the owner's address and sending funds using a transfer function (such as transfer), while generating a transaction completion record as an on-chain event log. This helps provide verifiable proof of fund flow, ensuring that the owner receives continuous benefits and enhancing trust.
[0150] In one possible implementation, after the transaction completion record is broadcast to the blockchain network, the ownership of the non-fungible token is transferred to the new holder by updating the token's metadata, such as the owner field. The smart contract calls the update function to synchronize the information to the market platform, which is a decentralized application interface that supports querying the token status. This is beneficial for automatically triggering the same fee deduction and distribution process in subsequent transactions, thereby maintaining the rights and interests of the template owner and promoting the ecological cycle.
[0151] For example, from another perspective, in the scenario of art creation templates, the core parameters of the task template, such as painting style rules, are generated by hashing and registered as NFTs. When an artist sells this NFT, the smart contract deducts a fee of 5% and directly distributes it to the original author's wallet. This mechanism supports the automatic distribution of royalties from secondary sales, which is beneficial for incentivizing content creation and forming a sustainable economic model. It complements the aforementioned process and ensures a complete chain from construction to multiple rounds of transactions.
[0152] In one possible implementation, the framework is extended to game development templates. During a transaction, the contract calculates fees and records them in the ledger before allocating funds. The generated record is broadcast to ensure the transfer of ownership. This aligns with the support of artistic scenarios and further strengthens the reliability of automated settlement through unique identifiers and smart contracts. This is beneficial for cross-domain applications and enhances the liquidity of digital assets.
[0153] S211, the collection of browser operation data is achieved by capturing URL features and tab switching order in real time through a lightweight plugin, and the switching order simulates the login path rules of the digital proxy model.
[0154] Browser operation data is collected in real time using a lightweight plugin. This data includes tab URL characteristics and switching order. For each tab access behavior, the timestamps of opening and closing, as well as the relationships between the tabs, are recorded to form an initial browsing path record. This browsing path record is then structured, and path rule templates are constructed based on frequently occurring tab jump relationships in the switching order. These templates describe a fixed order of jumping from one tab to another, forming a preliminary login path framework. Based on this login path framework and considering the user's operating habits on specific tabs, the jump priorities in the path rule templates are adjusted to form the final login path rules. These rules guide the digital proxy model's operation order during simulated login. The login path rules are embedded into the execution logic of the digital proxy model. During simulated browser operation, the target tabs are accessed sequentially according to these rules, ensuring that the operation order is consistent with user habits and effectively reproducing browser operation data.
[0155] Specifically, when the lightweight plugin runs in the browser, it monitors when a user opens a tab on a news website, records URL characteristics such as domain name and query parameters, and captures the order in which the user switches from that tab to the email tab. This collection method ensures real-time data collection because the plugin is directly embedded in the browser kernel, avoiding delays and providing accurate basic data for subsequent path formation. This improves the simulation accuracy of the digital agent model and avoids execution failures caused by operational deviations.
[0156] In one possible implementation, the collected browsing path records are structured and organized. For example, if users frequently jump from the homepage tab to the login page and then to the personal center, this frequent jumping is identified as a fixed sequence. When building the path rule template, these sequences are described as chain sequences, forming a preliminary login path framework. This process filters out noisy data by statistically analyzing the jumping frequency, which is beneficial to the stability of the framework because it makes the framework more in line with users' real habits. This reduces unnecessary jumping attempts during simulation and improves efficiency.
[0157] For example, by taking into account users' operational habits on the login page, such as clicking verification immediately after entering a username, the jump priority in the path rule template can be adjusted. For example, the verification step can be placed with high priority to form the final login path rule. This adjustment is based on the weighted processing of habit data, which can guide the digital agent model to prioritize the execution of key sequences during simulation. The beneficial effect of doing so is to enhance the adaptability of the rules, ensure that the model can efficiently simulate user behavior in different scenarios, and avoid login interruption caused by path confusion.
[0158] In one possible implementation, login path rules are embedded into the execution logic of the digital agent model. For example, when the model starts the simulation, it accesses the homepage, then the login page, and then the central page according to the rules, ensuring that the order is consistent with user habits. This embedding is achieved through logical mapping, which is beneficial for the effective reproduction of operational data because it allows the model to operate the browser naturally like a user, reducing the risk of anomaly detection. This improves reliability and consistency in the overall task automation. For example, in content generation tasks, this reproduction allows the agent to seamlessly complete login and generate professional content.
[0159] For example, examining the data collection process from multiple perspectives, the lightweight plugin not only records timestamps such as a tab being opened at 10:00 AM and closed at 10:05 AM, but also associates the transitions before and after the tab, such as from the search page to the results page. This supports the integrity of the path recording because time and associated data mutually verify each other, avoiding isolated information and improving the accuracy of subsequent processing. For instance, in a business scenario, the order in which a user switches from the email page to the calendar page is captured, allowing the framework to more comprehensively simulate multi-task login paths.
[0160] In one possible implementation, when building path rule templates, the order of frequent jumps, such as the daily repetition of the homepage to the creation center, is described as a fixed chain. This connects with the steps of adjusting priorities, because the template provides the foundation, and adjustments optimize it. The beneficial effect is the formation of more refined rules. For example, in a development environment, users are used to checking the log page before debugging. This allows the rules to guide the model to prioritize accessing the log page, reducing invalid operations in the simulation and improving the smoothness of task execution.
[0161] For example, the process of embedding execution logic involves transforming rules into an executable sequence. For instance, when a model receives a task, it sequentially calls the jump functions in the rules to ensure that the order of accessing the target tabs is consistent. This is related to the aforementioned framework and rules in that it is a transformation of the path from static to dynamic, which is beneficial for effective reproduction because the consistency operation can simulate the real user trajectory. For example, in a Web3.0 environment, this allows digital clones to reliably log in and store data in blockchain tasks, avoiding settlement failures caused by incorrect order.
[0162] In one possible implementation, user habits are analyzed from the side. For example, the longer a user stays on a specific page, the more likely they are to have a preference. When adjusting priorities, the weight of such jumps will be increased. This supports the personalization of rules because multi-directional data, such as time and order, support each other to form a robust framework. The beneficial effect is that the model simulation is more realistic. For example, in the scenario of AI glasses linkage, such rules can enable the agent to simulate browser operations related to user gestures, ensuring the automation efficiency of the overall ecosystem.
[0163] S212, the recognition of the screen interaction data combines computer vision technology and hook functions, the hook functions record the user's operation mode in the development environment and associate it with the task decision rules.
[0164] User operation records in the development environment are obtained from screen interaction data. These records include code editing trajectories and interface switching habits. Screen content is captured using computer vision technology to generate initial operation image data for subsequent behavior pattern recognition. Based on this initial operation image data, hook functions embedded in the development tool are used to record specific operations in the development environment in real time, capturing keystroke sequences and debugging operations to form a detailed operation behavior log. This log is then correlated with the initial operation image data. By comparing interface changes in the images with operation timestamps in the logs, user behavior habits in specific tasks are determined, generating a corresponding task decision rule mapping table. This task decision rule mapping table binds user operation habits to task execution logic. This binding process uses a pre-established rule base for matching, ensuring that the generated decision rules can be directly applied to automated task execution, thus completing the association between screen interaction data recognition and task decision rules.
[0165] For example, when acquiring user operation records in a development environment from screen interaction data, computer vision technology can be used to capture images of screen content. This technology involves using a camera or software tools to capture dynamic changes on the screen in real time, such as a user moving the cursor or switching windows in a code editor, thereby generating initial operation image data. The advantage of this method is that it captures visual interaction details, making subsequent behavior pattern recognition more accurate because it extracts trajectory information directly from visual input, avoiding the limitations of relying solely on logs. This yields useful technical benefits, namely, improving the comprehensiveness of data collection and ensuring that behavior patterns are based on real visual feedback.
[0166] In one possible implementation, initial operational image data is combined with hook functions for real-time recording. A hook function is a programming interface embedded in development tools such as an integrated development environment (IDE) that intercepts and records specific events, such as a user pressing a keyboard shortcut or clicking a debug button, thereby capturing keystroke sequences and operation order to form a detailed operational behavior log. This combination bridges visual data with event-level recording, as image data provides context, while hook functions add precise timing and action details. This approach yields useful technical benefits, namely increased log richness, support for more granular behavioral analysis, and avoidance of data silos.
[0167] For example, when correlating operation behavior logs with initial operation image data, user habits during tasks are determined by comparing interface changes in the images (such as window resizing) with operation timestamps in the logs (such as click timestamps). This might involve checking variables before coding. This comparison process involves synchronizing the timeline, aligning visual changes with event records, and generating a task decision rule mapping table. The purpose of this method is to establish logical connections between data, integrating habits from scattered information into structured rules. This yields a useful technical benefit: improved reliability of rule generation, facilitating subsequent automated applications.
[0168] In one possible implementation, user operating habits are bound to task execution logic through a task decision rule mapping table. This is achieved by matching rules against a pre-established rule base—a database storing standard behavioral templates, such as matching the habit of "test before coding" to an automation script template—ensuring that decision rules are directly applied to task execution. This binding process involves item-by-item comparison and assignment, allowing rules to be seamlessly integrated into the automation process. The purpose of this approach is to achieve a closed loop from data to execution, as it transforms abstract habits into actionable logic. This yields useful technical benefits, namely improved accuracy in task automation and reduced human intervention.
[0169] For example, in practical applications, from multiple perspectives, such as in front-end development scenarios, user operation logs can capture the habit of frequently switching to the console. Computer vision can be used to capture console window pop-ups on the screen, and hook functions can be used to record the sequence of log checks. Then, rules such as "prioritize log viewing during debugging" can be generated and ultimately bound to task logic to automatically execute similar steps. This multi-faceted support serves to verify the robustness of the method, as examples at the visual, event, and rule levels corroborate each other, forming a consistent automation path. This approach yields useful technical benefits, namely enhanced overall system adaptability and support for diverse development environments.
[0170] In one possible implementation, another approach is exemplified in backend development, where image data captures changes in the database query interface, hook functions record the SQL input order, and associated rules such as "back up data before querying" are generated and bound to a rule base for automated backup tasks. This example supports the aforementioned method from a data management perspective because it demonstrates how to handle different development types and ensure the universality of screen interaction recognition. This approach yields useful technical benefits, namely improved rule applicability across tasks, promoting efficient decision-making automation.
[0171] S221, the wearable device collects perception data by acquiring voice commands and microphone data through the AI glasses interface, and the voice commands are linked with the browser operation data to form reporting script rules.
[0172] Voice command data and ambient sound data collected by the microphone are acquired from the AI glasses interface. The voice command data includes the specific instructions issued by the user, and the ambient sound data is used to help determine the background information of the user's scene, generating an initial voice content set. The initial voice content set is parsed, and the voice command data and ambient sound data are categorized and processed to extract key intent words from the voice commands, forming a command intent list. Simultaneously, scene features from the ambient sound data are preserved to form scene auxiliary information. The command intent list is linked and matched with browser operation data. User-habitual operation paths are extracted from the browser operation data, and combined with the scene auxiliary information, a preliminary reporting script rule conforming to user operation habits is generated. The preliminary reporting script rule is optimized by deeply associating the key intent words in the command intent list with the specific operation paths in the browser operation data, forming the final reporting script rule to achieve seamless linkage between voice commands and browser operation data.
[0173] Specifically, when acquiring voice command data from the AI glasses interface, the voice command data can directly capture the user's spoken commands through the built-in speech recognition module, such as the user saying "Open the browser and switch to the development page." This data includes the text conversion result of the command and timestamp information. At the same time, the microphone collects ambient sound data to capture surrounding background noise, such as keyboard typing or meeting discussions in an office. This sound data helps determine whether the user is in a working environment, thereby generating an initial set of voice content. This set integrates the command content and environmental background into a unified data packet, which is beneficial for providing contextual support during subsequent parsing and avoiding misjudgments caused by processing commands in isolation.
[0174] In one possible implementation, during the content parsing of the initial voice content set, the voice command data is classified into intent-oriented parts, such as extracting keywords like "open" and "switch" to form a list of command intents. Meanwhile, the environmental sound data is classified to extract scene features, such as recognizing keyboard sounds to indicate a programming scene, forming scene auxiliary information. This classification process ensures the complementarity between the data. For example, the intent list can be used to guide operations, while the scene information supplements environmental adaptability, thereby improving the overall usability of the data and helping to reduce ambiguity during linkage.
[0175] Specifically, when linking and matching the list of instruction intents with browser operation data, the system extracts users' habitual operation paths from the browser operation data, such as the sequence of steps users commonly use to jump from the homepage to the toolbar. Combined with scenario-related information, such as sound characteristics in a programming environment, it generates a preliminary report script rule. For example, in a development task, it automatically generates a report statement such as "Switched to the development page, loading code". This matching process ensures that the preliminary rule conforms to user habits through the correspondence between paths and intents, which helps to improve the accuracy of automated execution.
[0176] For example, when optimizing the prototype of the reporting script rules, the key intent words in the instruction intent list, such as "open", are deeply associated with the specific operation path in the browser operation data, such as "homepage-toolbar", to form the final reporting script rules. For example, the optimized rules can generate "According to your instructions, we have switched from the homepage to the toolbar and are ready to report the code debugging results". This association process strengthens the synchronization between voice and operation, which is beneficial to achieving seamless linkage and improving the consistency of user experience.
[0177] S311, the dynamic weight update algorithm processes the browser operation data, the screen interaction data and the wearable device perception data based on the three-dimensional behavior mapping model, and the three-dimensional behavior mapping model constructs the scene-behavior-decision relationship.
[0178] This system acquires user tab switching data across different time periods from browser operation logs, extracts touch frequency and swipe trajectory information from screen interaction logs, and collects environmental perception data, including position changes and body movement signals, from wearable devices. These data are then aligned according to timestamps to form a unified time-series dataset. For this aligned dataset, a mapping between scenarios and behaviors is constructed. First, user operating habits in specific scenarios are determined using tab switching data and touch frequency information from the time-series dataset. Combined with position changes and body movement signals, the system determines whether the user's environment is a development or meeting scenario, establishing a preliminary mapping between scenario and behavior. Based on this preliminary mapping, decision factors are introduced. By analyzing user operating habits in different scenarios, the logic for adjusting task priorities is determined. The correlation between scenario, behavior, and decision is integrated to construct a complete three-dimensional mapping relationship for subsequent dynamic weight adjustments. Based on this constructed three-dimensional mapping relationship, browser operation data, screen interaction data, and wearable device perception data are comprehensively processed. By continuously updating the weight ratios of each data source, the weight adjustment logic adapts to behavioral changes in different scenarios, ultimately achieving the dynamic construction of the scenario-behavior-decision relationship.
[0179] Specifically, when obtaining tab switching data from browser operation logs, consider that users frequently switch developer tool tabs on weekday mornings. This helps to capture behavioral patterns. At the same time, extract touch frequency information from screen interactions, such as the number of clicks per minute, and swipe trajectory information, such as straight or curved paths. Collect location changes from wearable devices, such as movement from the office to the meeting room, and body movement signals, such as hand gestures. Align these data with timestamps to form a time-series dataset. Doing so ensures data synchronization and is beneficial for subsequent analysis of the continuity of user habits.
[0180] In one possible implementation, a scene-behavior correspondence is constructed for the aligned time-series dataset. Operation habits are determined by label switching data and touch frequency information, such as frequently switching code editing labels in a development scenario. The environment is determined by location changes, such as a static office area or a dynamic meeting area, as well as body movement signals, such as calm tapping or active discussion gestures, to form a preliminary mapping relationship. The beneficial effect of this is to improve the accuracy of user environment identification and avoid misjudgments caused by isolated data.
[0181] For example, decision factors are introduced on the basis of the initial mapping relationship. By analyzing operating habits, the task priority adjustment logic is determined, such as reducing the priority of development tasks in a meeting scenario. Scenarios such as meeting environment, behaviors such as tag minimization, and decisions such as priority notification processing are integrated into a three-dimensional mapping relationship for dynamic weight adjustment. This makes the system more adaptable to the user's decision-making process and helps to improve the efficiency and accuracy of task processing.
[0182] In one possible implementation, browser operation data such as page dwell time, screen interaction data such as multi-touch, and wearable device perception data such as heart rate changes are comprehensively processed based on three-dimensional mapping relationships. By updating the weight ratio, for example, increasing the weight of environmental data, the adjustment logic adapts to behavioral changes, such as switching from development to meeting. Ultimately, the dynamic construction of the scenario-behavior-decision relationship is achieved. The beneficial effect of this is to enhance the adaptability of the algorithm and the accuracy of simulating the user's real logic.
[0183] The above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the invention. Those skilled in the art will understand that implementing all or part of the above-described embodiments and making equivalent changes in accordance with the claims of the present invention are still within the scope of the invention.
Claims
1. A digital avatar task processing method based on Web3.0 and multimodal data fusion, characterized in that, include: The system collects browser operation data, screen interaction data, and wearable device perception data. The browser operation data includes tab creation and switching and content browsing characteristics. The screen interaction data includes code editing trajectory and interface switching habits. The wearable device perception data includes gaze point trajectory and gesture actions. A digital agent model is generated by fusing the browser operation data, the screen interaction data, and the wearable device perception data. The fusing process includes spatiotemporal alignment and dynamic weight adjustment. The digital agent model includes personalized decision rules. The digital agent model is deployed to the blockchain via a smart contract, which is used for task execution notarization and settlement. Simulate user browser operations according to the personalized decision-making rules, complete web page tasks, and store the results of the web page tasks on the blockchain for evidence. The rewards will be distributed to the user's wallet based on the results of the web page tasks. The task template of the digital agent model is encapsulated as an NFT and traded on the blockchain marketplace. The NFT transaction triggers the allocation of copyright fees to the wallet of the task template owner.
2. The method as described in claim 1, characterized in that, The collection of browser operation data, screen interaction data, and wearable device perception data includes: The browser operation data is obtained by capturing tab dwell time and content keywords in real time through a browser plugin, and the content keywords are extracted through natural language processing. The screen interaction data is obtained by collecting the debugging operation frequency through computer vision technology and IDE hook functions, and the debugging operation frequency is associated with the operation mode; Voice commands are acquired through the wearable device's visual and inertial sensors to obtain the wearable device's perception data, and the voice commands are associated with the environmental scene.
3. The method as described in claim 1, characterized in that, The process of fusing the browser operation data, the screen interaction data, and the wearable device perception data to generate a digital agent model includes: The browser operation data, the screen interaction data, and the wearable device perception data are spatiotemporally aligned to construct a three-dimensional behavior mapping model, which represents the scene behavior decision relationship. The weights of each modality data are dynamically adjusted through an attention mechanism to obtain a fusion result, which is then used to train the digital agent model. The task execution order and priority adjustment are generated based on the personalized decision-making rules.
4. The method as described in claim 2, characterized in that, The process of obtaining browser operation data by capturing tab dwell time and content keywords in real time through a browser plugin includes: URL features and tab switching order are extracted from the browser operation data, and the tab switching order is used to simulate the login path; A user preference graph is constructed using the content keywords, and the user preference graph is associated with a task priority model.
5. The method as described in claim 3, characterized in that, The step of spatiotemporally aligning the browser operation data, the screen interaction data, and the wearable device perception data to construct a three-dimensional behavior mapping model includes: The browser operation data is aligned with the scene information to obtain the first mapping data; The screen interaction data and the wearable device perception data are aligned based on behavioral characteristics to obtain second mapping data; The first mapping data and the second mapping data are fused to generate the three-dimensional behavior mapping model, which is used for decision rule iteration. The scene adaptation strategy is obtained from the three-dimensional behavior mapping model.
6. The method as described in claim 1, characterized in that, The step of deploying the digital agent model to the blockchain via a smart contract includes: The parameters of the digital agent model are trained through federated learning, and the parameters of the digital agent model are stored on the blockchain for evidence. The personalized decision-making rules are encapsulated for task splitting and execution.
7. The method as described in claim 1, characterized in that, The digital agent model simulates user browser operations based on the personalized decision rules to complete web page tasks, including: The digital agent model automatically completes account login based on the tab switching order; The digital proxy model generates content based on the code editing trajectory, and the content is stored on the blockchain as a hash value. The smart contract invokes an acceptance tool to check the results of the web page task.
8. The method as described in claim 1, characterized in that, The step of encapsulating the task template of the digital agent model into an NFT for trading on the blockchain marketplace includes: The task execution template is extracted from the digital agent model and encapsulated into the NFT; When the NFT is traded, the smart contract deducts copyright fees, which are then distributed to the user's wallet.