An ai browser intelligent interaction and automation operation system
Patent Information
- Application Number
- CN202610948831.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
然而,传统的浏览器交互方式严重依赖于用户手动操作,如点击、输入、滚动等,这在面对复杂、重复性的网页任务(如数据填报、电商比价、批量下载等)时,存在显著的低效性和易出错的风险
[0025]1)智能化的网页任务执行:本发明利用多模态感知技术(视觉理解、结构分析和语言理解)综合解析用户指令,使浏览器能够自动识别并执行复杂的网页任务,如点击按钮、填写表单、搜索信息等。系统不仅能够理解用户的自然语言指令,还能够通过视觉数据解析网页元素,识别用户意图,从而提供更精准、更智能的自动化执行。
Smart Images

Figure CN122817577A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of browser automation technology, and in particular to an AI browser intelligent interaction and automation operating system. Background Technology
[0002] With the rapid development of the internet and computer technology, browsers have become one of the core tools for people to obtain information, handle affairs, and enjoy entertainment. However, traditional browser interaction methods rely heavily on manual user operations, such as clicking, typing, and scrolling. This presents significant inefficiencies and a high risk of errors when faced with complex and repetitive web page tasks (such as data entry, e-commerce price comparison, and batch downloads).
[0003] Existing browser automation technologies, such as DOM-based automation scripts (like Selenium) and macro recording tools, while capable of automating some simple tasks, suffer from the following drawbacks: Vulnerability to changes in webpage structure: These tools rely on the underlying DOM structure of webpages; any subtle changes in the webpage layout can cause the automation script to malfunction. Lack of semantic understanding: Existing tools typically only handle simple DOM selectors and cannot understand the semantics of visual elements or user instructions on webpages. Therefore, they struggle with tasks requiring cognitive judgment (such as "finding the red confirmation button" or "ignoring pop-up ads"). High operational and programming requirements: Many automation tools require users to have a certain level of programming ability, which presents a significant barrier to entry for ordinary users without development experience. Furthermore, multimodal understanding technologies (combining visual, linguistic, and auditory information) have made significant progress in the field of artificial intelligence in recent years. For example, deep learning-based visual recognition and natural language processing (NLP) technologies have already achieved intelligent understanding and execution of complex tasks in certain scenarios. However, existing browser automation systems have not fully utilized these advanced technologies, failing to achieve a comprehensive understanding of webpage elements and user instructions, and lacking sufficient intelligence and adaptability. Meanwhile, with the increasing popularity of cross-device and multi-platform applications, modern users increasingly need to seamlessly switch tasks between different devices (such as mobile phones, tablets, and computers). Traditional browser automation tools usually cannot support cross-device and cross-platform task synchronization and scheduling, further increasing the operational burden on users when switching between multiple devices. Summary of the Invention
[0004] To address the aforementioned technical challenges, this invention proposes an AI-powered intelligent interactive and automated operating system for browsers. This system integrates advanced technologies such as multimodal data fusion, sentiment analysis, cross-platform task synchronization, and AR / VR interaction. It automatically understands user natural language commands, identifies visual elements on web pages, and intelligently schedules tasks based on real-time context. This significantly improves browser efficiency and intelligence, reduces user complexity, and provides a more intuitive and convenient automated operating experience. The innovation of this invention lies in its deep integration of multimodal visual understanding and natural language processing, combined with sentiment analysis and adaptive task scheduling capabilities, to provide a cross-platform, cross-device intelligent browser automation system, thereby greatly enhancing user operating efficiency and interactive experience.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows:
[0006] An AI browser intelligent interaction and automation operating system, comprising:
[0007] The multimodal perception and understanding module is used to extract and fuse information from web page data and user input data to generate a unified contextual representation. The user input data includes voice data and text data, and the web page data includes visual data, structural data, and user historical behavior data. Based on the unified contextual representation and user input data, a structured sequence of operation instructions is generated.
[0008] The cross-platform task scheduling module is used to schedule task execution based on task complexity, device status, and webpage status, and supports cross-platform task synchronization.
[0009] The sentiment analysis and intelligent feedback module is used to identify the emotional state of user input data in real time and adjust the system's feedback content and tone according to the user's emotional changes.
[0010] The adaptive task scheduling and dynamic decision-making module is used to dynamically optimize task execution strategies based on the feedback content of task execution and environmental changes using reinforcement learning.
[0011] The privacy protection and security module is used to securely process transmitted data using the AES-256 encryption algorithm and differential privacy noise during task execution.
[0012] Preferably, information is extracted from webpage data and user input data and fused to generate a unified contextual representation, including:
[0013] Speech recognition is performed on the speech data to obtain text data;
[0014] Semantic understanding of text data is performed using BERT or GPT models to extract text features;
[0015] Image features are obtained by extracting features from visual data using the Vision Transformer model.
[0016] Graph neural networks are used to parse structural data and generate structural features.
[0017] Cross-modal attention is used to fuse text features, image features, and structural features to obtain a unified contextual representation.
[0018] Preferably, the language recognition uses the Wav2Vec2 model or the Google Speech API model.
[0019] Preferably, the adaptive task scheduling and dynamic decision-making module incorporates near-end strategy optimization for reinforcement learning.
[0020] Preferably, an LSTM or BiLSTM model is used to identify the emotional state of the user input data.
[0021] Preferably, the privacy protection and security module is also used to monitor security risks during task execution in real time through anomaly detection algorithms to prevent malicious attacks or data leaks.
[0022] Preferably, it also includes an augmented reality / virtual reality interaction module, used to acquire user behavior data and spatial perception information in the AR / VR environment and convert them into control commands for the browser.
[0023] Preferably, it also includes a personalized learning and recommendation module, which uses collaborative filtering algorithms or deep learning models to deeply mine users' historical behavior data, predict tasks or operation processes that users are interested in, and provide recommendations.
[0024] Based on the above technical solution, the beneficial effects of the present invention are:
[0025] 1) Intelligent Webpage Task Execution: This invention utilizes multimodal perception technology (visual understanding, structural analysis, and language understanding) to comprehensively analyze user commands, enabling the browser to automatically identify and execute complex webpage tasks, such as clicking buttons, filling out forms, and searching for information. The system not only understands users' natural language commands but also analyzes webpage elements through visual data to identify user intent, thereby providing more accurate and intelligent automated execution.
[0026] 2) Dynamic Task Scheduling and Optimization: By employing reinforcement learning (such as the PPO algorithm), the system can dynamically optimize task execution strategies based on real-time feedback and environmental changes. When faced with issues such as changes in webpage elements, task failures, or loading delays, the system can intelligently select alternative paths or adjust the execution order to ensure efficient and successful task execution. This adaptive decision-making capability is difficult to achieve in traditional automation tools.
[0027] 3) Cross-device and cross-platform synchronization: This invention supports cross-platform and cross-device task execution. Whether on a desktop, mobile device, or other device, the system can intelligently select the appropriate device to execute the task based on the device's status (such as battery level and network quality), ensuring seamless task synchronization across multiple devices. This feature significantly improves the user experience, especially suitable for scenarios requiring switching between different devices.
[0028] 4) Emotion-Driven Personalized Feedback: Through sentiment analysis technology, the system can identify users' emotional states in real time and adjust the content and tone of feedback according to changes in user emotions. For example, when the system detects that a user is anxious, it will provide a gentler and more comforting response; when a user is satisfied, the system will provide timely encouragement and confirmation. This emotion-driven interaction method greatly enhances the user experience, making it more personalized and human-centered.
[0029] 5) Immersive Interaction in Augmented Reality (AR) / Virtual Reality (VR) Environments: This invention supports browser operation within AR / VR environments. Users can directly interact with the browser through head-mounted devices, gesture control, and eye tracking, enhancing the interactivity and immersion between the user and the browser. This innovative interaction method provides a completely new experience for the automated execution of complex tasks and adapts to future technological development trends.
[0030] 6) Highly Efficient Data Security and Privacy Protection: This invention incorporates in-depth design for data security and privacy protection, employing AES-256 encryption algorithm and Differential Privacy technology to ensure the privacy and security of user data during task execution. All sensitive user data undergoes de-identification processing and is stored and transmitted encrypted during task execution, minimizing the risk of data leakage or misuse.
[0031] 7) Simplified Operation: Because the system can execute tasks via natural language commands, users do not need programming or technical backgrounds to achieve efficient automation. This lowers the barrier to entry for users of automation tools, enabling even ordinary users to easily perform complex web page operations and promoting the popularization of automation technology.
[0032] 8) Wide Range of Application Scenarios: This invention is applicable to multiple fields such as e-commerce, office automation, data collection, online learning, and entertainment, and can provide intelligent automation services for various complex web page tasks. Whether it's automatically filling out forms, automatically purchasing goods, or complex data scraping and processing, this invention can complete tasks efficiently and accurately, saving users time and effort. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the structure of an AI browser intelligent interaction and automation operating system in one embodiment;
[0034] Figure 2 This is a schematic diagram of the processing flow of the multimodal perception and understanding module in one embodiment;
[0035] Figure 3 This is a schematic diagram of the processing flow of the cross-platform task scheduling module in one embodiment;
[0036] Figure 4 This is a schematic diagram of the processing flow of the privacy protection and security module in one embodiment. Detailed Implementation
[0037] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0038] like Figure 1 As shown, this embodiment provides an AI browser intelligent interaction and automation operating system, including: a multimodal perception and understanding module, a cross-platform task scheduling module, a sentiment analysis and intelligent feedback module, an adaptive task scheduling and dynamic decision-making module, an augmented reality (AR) / virtual reality (VR) interaction module, a personalized learning and recommendation module, and a privacy protection and security module. This invention can automatically understand users' natural language commands, identify visual elements in web pages, and intelligently schedule tasks based on real-time context, thereby significantly improving the efficiency and intelligence of browser operations, reducing user operational complexity, and providing a more intuitive and convenient automated operation experience. The following is a description:
[0039] 1. Multimodal perception and understanding module
[0040] In this embodiment, the core task of the multimodal perception and understanding module is to extract information from visual, structural, and linguistic data, understand webpage content, and generate executable task instructions based on the user's natural language commands, such as... Figure 2 .
[0041] Users input commands via voice or text. For example, "Find the cheapest wireless mouse and add it to your cart." The system converts voice data into text data using speech recognition technology (such as using Wav2Vec2 or the Google Speech API). Webpage data collection: The system obtains data from the current webpage through browser plugins or APIs, including:
[0042] Visual data: A screenshot or video stream of the current browser viewport;
[0043] Structural data: the DOM tree and accessibility tree of a webpage;
[0044] User historical behavior data: user clicks, search history, etc.
[0045] 1) Visual data processing
[0046] The system processes webpage screenshots, extracts image features using Vision Transformer (ViT), and identifies interactive elements on the page (such as buttons, images, input boxes, etc.).
[0047] Image patches are embedded using Vision Transformer (ViT).
[0048]
[0049] in It is an input image patch, which is then encoded by a Vision Transformer and output as the output. , representing the image feature representation.
[0050] 2) Structured data parsing
[0051] By using graph neural networks (GNNs) to parse the DOM tree of a webpage, the hierarchical structure and relationships of webpage elements can be extracted to understand the semantics of the elements.
[0052] Graph Neural Networks (GNNs) are used to parse the DOM tree and generate structural features.
[0053]
[0054] in It is a node In the Layer feature representation, It is a node The set of neighboring nodes, It is a weight matrix. It's a bias. It is an activation function (such as ReLU).
[0055] 3) Language data processing
[0056] The user's natural language commands are semantically understood using models such as BERT or GPT to extract the operation goals and task requirements.
[0057] Use BERT or GPT to parse natural language instructions.
[0058]
[0059] in, It is the input text command. It is the text feature representation output by the BERT model.
[0060] 4) Data fusion
[0061] By using cross-modal attention technology, visual, structural, and linguistic data are fused to generate a unified contextual representation, ultimately parsing out the specific task steps.
[0062] Cross-modal attention is used to fuse visual, structural, and linguistic data, as follows:
[0063]
[0064] in, , and It is a feature representation of visual, textual, and structural data.
[0065] 2. Cross-platform task scheduling module
[0066] In this embodiment, the cross-platform task scheduling module is responsible for scheduling task execution based on task priority, device performance, and the real-time status of the webpage, supporting cross-platform task synchronization. The system decomposes high-level user commands (e.g., "find the cheapest wireless mouse") into a series of specific operations, such as "search for wireless mice," "sort results," and "select the cheapest mouse," and generates execution paths, such as... Figure 3 .
[0067] 1) Task scheduling algorithm
[0068] Based on the complexity of the task, the status of the webpage (such as whether it has finished loading), and the status of the device (such as battery level, network speed, etc.), the system sorts the tasks using an algorithm and determines the optimal execution order.
[0069]
[0070] in, It is a node The total cost, From the starting point to the node The actual cost, From node Heuristic cost estimation for reaching the target.
[0071] 2) Equipment selection and cross-platform synchronization
[0072] If a task needs to be executed on multiple devices, the system will select the appropriate device for operation and ensure that the task status is consistent through cloud synchronization.
[0073]
[0074] in, It is a set of equipment. It is equipment The delay It is equipment Energy consumption.
[0075] 3. Sentiment Analysis and Intelligent Feedback Module
[0076] In this embodiment, the sentiment analysis and intelligent feedback module adjusts the feedback strategy according to the user's emotional state to improve the user's interactive experience.
[0077] 1) Sentiment Analysis
[0078] LSTM or BiLSTM models are used to analyze users' emotional states. Sentiment analysis is based on users' text or voice input to identify emotions such as anxiety, anger, and satisfaction.
[0079]
[0080] in, Is it LSTM at time step The hidden state, This is the input for the current time step.
[0081] 2) Emotion-driven feedback optimization
[0082] Based on the results of sentiment analysis, the system adjusts the tone of its feedback. For example, if a user exhibits anxiety, the system provides gentler, more comforting feedback; if a user is angry, the system offers quick solutions or bug fixes.
[0083]
[0084] in, It refers to the current emotional state. It is the output of sentiment analysis. It is a strategy of adjusting feedback based on emotional state.
[0085] 4. Adaptive Task Scheduling and Dynamic Decision-Making Module
[0086] In this embodiment, the adaptive task scheduling and dynamic decision-making module dynamically adjusts the execution order of tasks based on real-time feedback to ensure efficient task completion. The system simulates user operations based on parsed instructions, automatically clicking buttons, entering text, submitting forms, etc. Operations are performed through browser APIs or extensions, directly interacting with webpage elements.
[0087] 1) Task Decision-Making and Optimization
[0088] If problems are encountered during task execution (such as element loading failure, page structure changes, etc.), the system dynamically adjusts the execution path or tries alternative solutions through real-time monitoring (such as PPO reinforcement learning algorithm) to ensure that the task can be completed smoothly.
[0089]
[0090] in, It is a policy function. It is the dominant function. It is the shearing threshold.
[0091] 2) Real-time decision-making and adjustment
[0092] The system provides feedback based on the task execution result (success or failure). If the task succeeds, the result is returned; if the task fails, the system automatically adjusts its strategy based on the error message. Based on the real-time task status and user input, the system dynamically selects the optimal task execution strategy using Proximal Policy Optimization (PPO).
[0093] 5. Augmented Reality (AR) / Virtual Reality (VR) Interaction Module
[0094] In this embodiment, users can interact with the browser in an AR / VR environment and control the browser through input methods such as eye tracking and gestures.
[0095] 1) Gesture and eye tracking recognition
[0096] In AR or VR environments, the system uses Media Pipe or Open Pose technology for gesture recognition and translates them into browser control commands (such as clicks and drags). Through Gaze Tracking technology, the system can automatically execute browser operations based on the user's gaze.
[0097]
[0098] in, It is user behavior data (such as gestures or eye movement data). This is the recognition result.
[0099] 2) AR / VR environment adaptation
[0100] The system uses SLAM technology to sense the user's location in real time and adjusts the display of task execution in the virtual environment to make the task operation conform to the user's needs in the virtual environment.
[0101]
[0102] in, This is the user's current location. It's the user's speed.
[0103] 6. Personalized Learning and Recommendation Module
[0104] In this embodiment, the personalized learning and recommendation module provides personalized task execution recommendations and intelligent operation path optimization based on the user's historical behavior data.
[0105] 1) Behavioral Analysis and Modeling
[0106] User behavior is modeled based on collaborative filtering or deep learning.
[0107]
[0108] in, User For the project Predicted score and They are users and projects The embedding vector.
[0109] 2) Intelligent Recommendation
[0110] Based on the user's historical behavior, the system predicts tasks or operation processes that the user may be interested in and provides recommendations.
[0111] 7. Privacy Protection and Security Module
[0112] In this embodiment, to ensure that user privacy and data security are protected during the execution of automated tasks, such as... Figure 4 .
[0113] 1) Data encryption and protection
[0114] The system uses the AES-256 encryption algorithm to ensure the security of user data transmission and storage, preventing information leakage. User data is encrypted using AES-256 as follows:
[0115]
[0116] in, It is plain text. It is ciphertext. It is an encryption algorithm.
[0117] 2) De-identification processing
[0118] Sensitive user data is de-identified and differential privacy noise is used to protect user privacy and prevent data leakage.
[0119] Protect user data using Differential Privacy.
[0120]
[0121] in, These are the query results. and They are adjacent datasets. It's a privacy budget.
[0122] 3) Anomaly detection
[0123] The system uses anomaly detection algorithms (such as Isolation Forest) to monitor security risks during task execution in real time, preventing malicious attacks or data leaks.
[0124] Detect potential security threats using IsolationForest.
[0125]
[0126] in, Data points Path length in the tree It is a constant. This refers to the size of the dataset.
[0127] In the above embodiments, the various modules of an AI browser intelligent interaction and automation operating system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0128] The above are merely preferred embodiments of the present application and are not intended to limit the embodiments of the present application. For those skilled in the art, the embodiments of the present application can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of the present application should be included within the protection scope of the embodiments of the present application.
Claims
1. An AI browser intelligent interaction and automation operating system, characterized in that, include: The multimodal perception and understanding module is used to extract and fuse information from web page data and user input data to generate a unified contextual representation. The user input data includes voice data and text data, and the web page data includes visual data, structural data, and user historical behavior data. Based on the unified contextual representation and user input data, a structured sequence of operation instructions is generated. The cross-platform task scheduling module is used to schedule task execution based on task complexity, device status, and webpage status, and supports cross-platform task synchronization. The sentiment analysis and intelligent feedback module is used to identify the emotional state of user input data in real time and adjust the system's feedback content and tone according to the user's emotional changes. The adaptive task scheduling and dynamic decision-making module is used to dynamically optimize task execution strategies based on the feedback content of task execution and environmental changes using reinforcement learning. The privacy protection and security module is used to securely process transmitted data using the AES-256 encryption algorithm and differential privacy noise during task execution.
2. The AI browser intelligent interaction and automation operating system according to claim 1, characterized in that, Information is extracted from and fused from webpage data and user input data to generate a unified contextual representation, including: Speech recognition is performed on the speech data to obtain text data; Semantic understanding of text data is performed using BERT or GPT models to extract text features; Image features are obtained by extracting features from visual data using the Vision Transformer model. Graph neural networks are used to parse structural data and generate structural features. Cross-modal attention is used to fuse text features, image features, and structural features to obtain a unified contextual representation.
3. The AI browser intelligent interaction and automation operating system according to claim 2, characterized in that, The language recognition uses either the Wav2Vec2 model or the Google Speech API model.
4. The AI browser intelligent interaction and automation operating system according to claim 1, characterized in that, The adaptive task scheduling and dynamic decision-making module incorporates near-end strategy optimization for reinforcement learning.
5. The AI browser intelligent interaction and automation operating system according to claim 1, characterized in that, Use LSTM or BiLSTM models to identify the emotional state of user input data.
6. The AI browser intelligent interaction and automation operating system according to claim 1, characterized in that, The privacy protection and security module is also used to monitor security risks during task execution in real time through anomaly detection algorithms to prevent malicious attacks or data leaks.
7. The AI browser intelligent interaction and automation operating system according to claim 1, characterized in that, It also includes an augmented reality / virtual reality interaction module, which is used to acquire user behavior data and spatial perception information in the AR / VR environment and convert them into control commands for the browser.
8. The AI browser intelligent interaction and automation operating system according to claim 1, characterized in that, It also includes a personalized learning and recommendation module, which uses collaborative filtering algorithms or deep learning models to deeply mine users' historical behavior data, predict the tasks or operation processes that users are interested in, and provide recommendations.