Data processing method and device

By encapsulating the data processing path in the target interface layer, the problem of browser process isolation limitation is solved, enabling efficient cross-framework communication and document object model operations, thus improving task execution efficiency.

CN121858321APending Publication Date: 2026-04-14UC MOBILE CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UC MOBILE CHINA CO LTD
Filing Date
2025-11-13
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, launching a separate browser process within the target browser to handle the target task makes it difficult for automation tools to run directly, making them susceptible to being blocked by security policies. Furthermore, different automation tools require different browser configurations, resulting in low execution efficiency.

Method used

By encapsulating the data processing path in the target interface layer, selecting the appropriate target data processing path, and shielding the browser process isolation restrictions, efficient cross-framework communication and document object model operations are achieved, generating page snapshots.

Benefits of technology

It improves the execution efficiency of the target task, realizes efficient cross-framework communication and document object model operations, avoids the isolation limitations between browser processes, and improves the success rate and efficiency of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858321A_ABST
    Figure CN121858321A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device. The data processing method is applied to a target browser and comprises the steps that a to-be-processed target task is acquired, and a to-be-processed instruction is determined based on the to-be-processed target task; based on the instruction to be processed, a target data processing path is selected in a target interface layer, and at least one data processing path is packaged in the target interface layer; accessing a target page based on the target data processing path, and extracting page content of the target page; and generating a page snapshot based on the page content, and determining a target result based on the page snapshot. Through the data processing method provided by the invention, the processing efficiency of the to-be-processed target task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a data processing method and apparatus. Background Technology

[0002] With the continuous development of computer technology, when a target task is executed in the target browser using an automation tool, it is usually necessary to start an independent browser process other than the target browser and execute the target task according to the independent browser process.

[0003] The above methods avoid impacting the target browser. Furthermore, since target browsers typically have security policies, running automation tools directly in the target browser is likely to be blocked by these policies. Moreover, different automation tools often require different browser configurations; therefore, starting a separate browser process can avoid related conflicts.

[0004] However, launching a separate browser process to handle the target task makes it difficult for automation tools to run directly in the target browser, resulting in low execution efficiency of the target task. Summary of the Invention

[0005] In view of the above, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0006] According to a first aspect of the embodiments of this specification, a data processing method is provided, applied to a target browser, comprising: Obtain the target task to be processed, and determine the instruction to be processed based on the target task; Based on the instruction to be processed, a target data processing path is selected in the target interface layer, wherein the target interface layer encapsulates at least one data processing path; Access the target page based on the target data processing path and extract the page content of the target page; A page snapshot is generated based on the page content, and the target result is determined based on the page snapshot.

[0007] According to a second aspect of the embodiments of this specification, a data processing method is provided, applied to a cloud-side device, comprising: The receiving end device acquires the target task to be processed and determines the instruction to be processed based on the target task to be processed; Based on the instruction to be processed, a target data processing path is selected in the target interface layer, wherein the target interface layer encapsulates at least one data processing path; Access the target page based on the target data processing path and extract the page content of the target page; A page snapshot is generated based on the page content, and the target result is determined based on the page snapshot; The target result is sent to the end-side device.

[0008] According to a third aspect of the embodiments of this specification, a task platform is provided, including a request interface and a response unit; The request interface is used to obtain the target task to be processed and to determine the instruction to be processed based on the target task to be processed; The response unit is configured to select a target data processing path in the target interface layer based on the instruction to be processed, wherein the target interface layer encapsulates at least one data processing path; access a target page based on the target data processing path and extract the page content of the target page; generate a page snapshot based on the page content and determine the target result based on the page snapshot.

[0009] According to a fourth aspect of the embodiments of this specification, a data processing apparatus is provided, applied to a target browser, comprising: The acquisition unit is configured to acquire the target task to be processed and determine the instruction to be processed based on the target task to be processed; The selection unit is configured to select a target data processing path in the target interface layer based on the instruction to be processed, wherein the target interface layer encapsulates at least one data processing path. The extraction unit is configured to access the target page based on the target data processing path and extract the page content of the target page; The determining unit is configured to generate a page snapshot based on the page content and determine the target result based on the page snapshot.

[0010] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions, which, when executed by the processor, implement the steps of the above method.

[0011] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program or instructions which, when executed by a processor, implement the steps of the above-described method.

[0012] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0013] According to one embodiment of this specification, a target task to be processed in the target browser is obtained and converted into a processing instruction. Since at least one data processing path is encapsulated in the target interface layer, security policy interception within the target browser process is blocked. Therefore, control within the target browser process and between target browser processes can be achieved based on the at least one data processing path encapsulated in the target interface layer. Based on this, a suitable target data processing path can be selected in the target interface layer according to the processing instruction, so as to subsequently access the target page according to the target data processing path and extract the page content of the target page, thereby determining the target result. Furthermore, by encapsulating at least one data processing path in the target interface layer, the method of obtaining the target data processing path can improve the execution efficiency of the target task. Attached Figure Description

[0014] Figure 1 A flowchart of a data processing method provided according to an embodiment of this specification is shown; Figure 2 An architecture diagram for generating page snapshots according to one embodiment of this specification is shown; Figure 3 A system architecture diagram according to one embodiment of this specification is shown; Figure 4 A schematic diagram of a data processing method according to an embodiment of this specification is shown; Figure 5 This specification shows a schematic diagram of the structure of a data processing apparatus according to one embodiment; Figure 6 This specification shows an architecture diagram of a data processing system provided in one embodiment; Figure 7 A block diagram of a computing device structure according to an embodiment of this application is shown. Detailed Implementation

[0015] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0016] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0017] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0018] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0019] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0020] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0021] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0022] The Chrome DevTools Protocol (CDP) can be understood as the Chrome DevTools Protocol, which controls the browser's communication protocols.

[0023] An inline frame element (iframe element) can be understood as an inline frame element in a webpage that can embed another page within the current page.

[0024] The Model Context Protocol (MCP) can be understood as a standardized communication protocol used between artificial intelligence models and external tools.

[0025] A Document Object Model (DOM) snapshot can be understood as a structured representation of the web page's document object model, containing accessibility information and interactive state of elements.

[0026] The protocol mapping interface (IFrameDebugSession) can be understood as the core interface used in this specification to provide CDP protocol adaptation in the iframe environment.

[0027] In the target browser, the execution of the target task typically involves launching a separate browser process to handle the task, making it difficult for automation tools to run directly within the target browser. However, this approach struggles to overcome the process isolation limitations of embedded frames. Furthermore, it not only hinders efficient cross-frame communication but also prevents smooth cross-frame document object model operations and the integration of artificial intelligence capabilities with browser operations, ultimately resulting in low execution efficiency of the target task.

[0028] In view of this, this specification provides a data processing method that encapsulates at least one data processing path through a target interface layer, thereby achieving the shielding of isolation limitations between different browser processes through a unified interface layer. Based on this, according to the instruction to be processed, a suitable target data processing path is selected from the target interface layer encapsulating at least one data processing path, thereby achieving efficient cross-frame communication. This allows access to the target page based on the target data processing path, generating a page snapshot. The data processing method provided in this specification completes cross-frame document object model operations, improving the execution efficiency of the target task. This specification provides a data processing method, and also relates to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.

[0029] See Figure 1 , Figure 1 A flowchart of a data processing method according to an embodiment of this specification is shown, specifically including steps 102-108.

[0030] In one specific embodiment provided in this specification, the data processing method described herein is executed on a system where a target browser is deployed. It can be executed on a user client where the target browser is deployed, or on a corresponding server. The target browser can be understood as the primary browser environment that performs the data processing task, such as a locally installed browser application. Based on this, the data processing method is explained and described in the following manner.

[0031] Step 102: Obtain the target task to be processed, and determine the instruction to be processed based on the target task to be processed.

[0032] The target task to be processed can be understood as a semantic task description that the target user proposes in the browser or in the AI ​​agent that accompanies the browser and that the target browser needs to execute. It is usually expressed in natural language.

[0033] The instructions to be processed can be understood as a structured, executable sequence of operation instructions obtained after parsing the target task. These instructions conform to the standard format of the target browser's automated operations.

[0034] An AI agent is an executor that controls a browser to automate web page creation. An AI agent includes an intelligent processing unit and an instruction parsing model.

[0035] To facilitate understanding, this manual uses the target task to be processed as an example, which is a semantic task description proposed by the target user in the AI ​​agent that comes with the browser and needs to be executed by the target browser, to explain the data processing method.

[0036] In one specific embodiment provided in this specification, obtaining the target task to be processed and determining the instruction to be processed based on the target task includes: Obtain task information sent by the target user; The task information is analyzed by the intelligent processing unit to obtain the target task to be processed. The target task to be processed is parsed based on the instruction parsing model to obtain the initial instruction to be processed; Based on the model context interaction protocol, the initial instruction to be processed is encapsulated into a standardized protocol message; The standardized protocol message is identified as a pending instruction, wherein the pending instruction is used to guide automated operations on the target page based on the target data processing path.

[0037] In this context, the target user can be understood as the end user using the target browser, which can be an individual user or a system user with specific identity and permission settings. Task information can be understood as the original description of the requirements entered by the target user, usually expressed in natural language, containing user intent but lacking structured information.

[0038] The intelligent processing unit can be understood as a module deployed within an AI agent, used to receive task information and possessing natural language understanding and task analysis capabilities. This AI agent is an intelligent processing module integrated into a system with a target browser. The target task to be processed can be understood as a task description that has undergone preliminary parsing and standardization, retaining the user's original intent but pre-processed in a structured manner.

[0039] Instruction parsing models can be understood as specialized parsers built upon Large Language Models (LLMs) that can transform semantic tasks into executable sequences of operations. They collaborate with intelligent processing units deployed within the AI ​​agent to process user-input task information.

[0040] The initial instructions to be processed can be understood as the preliminary sequence of operation instructions output by the model, containing specific browser-based automated web page operation steps. This preliminary sequence of operation instructions can be further standardized in subsequent steps.

[0041] The Model Context Interaction Protocol (MCP) can be understood as a standardized communication protocol defined by the system to ensure data format consistency between different components. For example, the MCP eliminates interface incompatibility issues between the AI ​​agent deployed in the target browser and the browser's content. Standardized protocol messages can be understood as structured data packets conforming to a predefined format, containing complete execution context information. Pending instructions can be understood as finalized, standardized instructions that can be sent to the execution engine. The target page can be understood as the webpage from which automated operations need to be performed.

[0042] In one specific embodiment provided in this specification, the AI ​​agent receives a user's task request via a user interface or API interface. The AI ​​agent's intelligent processing unit then cleanses, normalizes, and identifies the intent of the target task input by the user. Following this process, the AI ​​agent's instruction parsing model generates specific operation steps based on the target task context to obtain initial instructions for processing. These initial instructions are then encapsulated according to protocol specifications to obtain the final instructions for processing, ensuring their integrity and executability. The final confirmation and distribution preparation of these instructions are then completed, ensuring that they can be correctly executed by the selected target data processing path.

[0043] As described above in this specification, the Model Context Interaction Protocol (MINTERP) provides a unified communication language for all system components. Therefore, during the process of acquiring the target task from the intelligent processing unit, the MINTERP can eliminate the interface incompatibility issue between the AI ​​agent and the target browser kernel. Furthermore, the use of standardized protocol message formats for the instructions ensures the integrity of the instructions and the consistency of the execution environment's context.

[0044] If a processing instruction is obtained according to this specification, a target data processing path is further selected in the target interface layer based on the processing instruction.

[0045] Step 104: Based on the instruction to be processed, select a target data processing path in the target interface layer, wherein the target interface layer encapsulates at least one data processing path.

[0046] The target interface layer can be understood as an abstract interface layer deployed in the system architecture of the target browser. It provides a unified data processing path control interface for intra-process communication or inter-process communication of the browser, thereby shielding the differences between the different implementation methods of intra-process communication and inter-process communication of the browser and encapsulating the project logic for selecting the target data processing path.

[0047] The data processing path can be understood as the specific automated implementation method of the target browser, such as the in-process control path of the target browser based on iframe, the inter-process control path of the browser based on the CDP protocol, etc.

[0048] The target data processing path can be understood as the execution path selected from multiple data processing paths based on specific conditions (such as user login status, task complexity, etc.) to execute the pending instructions corresponding to those conditions. For example, when the user is logged into the target browser, the pending instructions can be understood as instructions used to select and execute the embedded processing path within the browser process. When the user is not logged into the target browser, the pending instructions can be understood as instructions used to select and execute the remote cloud path between browser processes. Furthermore, when the user is logged into the target browser, but the complexity of the target task is high, the pending instructions can be understood as instructions used to select and execute both the embedded processing path within the browser process and the remote path between browser processes.

[0049] To facilitate understanding, the process of selecting the target data processing path is explained in the following manner in this manual.

[0050] In one specific embodiment provided in this specification, selecting a target data processing path in the target interface layer based on the instruction to be processed includes: Identify the target user corresponding to the target task to be processed, and determine at least one login status corresponding to the target user in the target browser; Based on the login status, select the target data processing path in the target interface layer.

[0051] The login status can be understood as the authentication status of the target user in the target browser, which reflects the target user's permission level and access permissions in the current browsing session.

[0052] The target data processing path can be understood as the specific technical execution path selected based on the login state decision results.

[0053] As described above, when the system architecture detects that a target user is logged in, it prioritizes using resources within the target browser process to reduce network overhead and server load. When the user is not logged in, it uses cloud resources (i.e., inter-process resources of the target browser). Path selection based on the user's login status enables reasonable scheduling and load balancing of computing resources.

[0054] It should be understood that the login state mentioned in this specification includes both logged-in and unlogged-in states, and the target interface layer encapsulates at least the embedded processing path and the cloud remote path.

[0055] The embedded processing path can be understood as a localized processing solution based on the iframe embedded framework, which executes tasks directly within the user's current browser environment. Specifically, the embedded processing path can be understood as the in-process control path of the target browser, that is, a sandboxed control mechanism within the target browser process.

[0056] It's important to understand that when using an embedded processing path, because the target browser's target interface layer encapsulates the browser's intra-process communication processes, the embedded page in the target browser and the user interface (main page) of the AI ​​agent containing the task to be processed share the same target browser process. This is because the logical isolation is achieved through the target browser's security sandbox, not process isolation. Therefore, the embedded page and the user interface in the AI ​​agent can share the same CPU and memory resources within the same target browser process.

[0057] A cloud-based remote path can be understood as a cloud-based processing solution based on a remote browser instance, executing tasks in an isolated cloud environment. Specifically, because the target browser's target interface layer encapsulates the browser's inter-process communication process, a cloud-based remote path can be understood as a cross-process remote control mechanism between the target browser's process and the target browser's process. It is independent of the target browser's process and is a separate browser process.

[0058] It's important to understand that when using a cloud-based remote path, there is process isolation between the cloud-based remote path and the target browser. Furthermore, communication between the cloud-based remote path and the target browser uses a network protocol (WebSocket).

[0059] Furthermore, in a specific embodiment provided in this specification, the embedded processing path is constructed in the following manner: Create an embedded frame element in the target browser; The target bridging script is injected into the embedded page corresponding to the embedded frame element, wherein the target bridging script includes a protocol mapping interface; Based on the instruction to be processed, control the embedded frame element to navigate to the target page.

[0060] In this context, the embedded frame element can be understood as an iframe element in the target browser's webpage, used to embed another independent document environment within the current page. The target bridging script can be understood as the core JavaScript code that implements protocol conversion and communication bridging. The embedded page can be understood as the specific webpage content loaded by the embedded frame element. The protocol mapping interface can be understood as a conversion interface that translates standardized instructions into native operations of the target page.

[0061] In one specific implementation provided in this specification, an independent embedded page is created, isolated from the main page environment. Permissions are controlled through the sandbox attribute of the target browser's security sandbox. Based on this, standardized pending instructions (e.g., CDP instructions) are received and converted into native operation instructions executable on the target page. According to the converted native operation instructions, the embedded frame element is controlled to navigate to the target page.

[0062] As described above, an isolated execution environment is created using iframes to avoid interference or contamination of the main page. In-process communication is performed using a secure sandbox based on the embedded frame element. The bridging script converts the standard CDP protocol into native operations of the target page. Furthermore, an independent execution environment is achieved based on the embedded page provided by the embedded frame element, thereby preventing the stability of the main system architecture from being affected by abnormalities in the target page.

[0063] Furthermore, in a specific embodiment provided in this specification, the cloud-based remote path is constructed in the following manner: Start a remote browser instance, wherein the remote browser instance is used to enable the developer tools protocol debugging port; Obtain the Uniform Resource Locator (URL) for the network socket provided by the remote browser instance; Based on the network socket Uniform Resource Locator, establish a network socket connection with the remote browser instance; Based on the network socket connection, a developer tools protocol session is created and attached to the remote browser instance.

[0064] In this context, a remote browser instance can be understood as a complete browser process launched in a cloud server environment, typically running in headless mode. The Developer Tools Protocol (DP) debug port can be understood as the CDP service listening port, providing a programmatic control interface to the browser. The Network Socket Uniform Resource Locator (URL) can be understood as a WebSocket debug URL, used to establish a bidirectional communication connection with the browser instance. A Network Socket connection can be understood as a full-duplex communication channel based on the WebSocket protocol. The Developer Tools Protocol (DP) session can be understood as the CDP session context, maintaining the browser state and operation sequence.

[0065] In one specific implementation provided in this specification, a browser instance is initialized and a debugging interface is configured in a cloud environment: the browser runs in headless mode. A CDP debugging service is started on a specified CDP port, and necessary browser parameters are configured to ensure stable operation and secure isolation. A separate user data directory is created to maintain session isolation. In the remote browser instance, a WebSocket debugging URL is provided to establish a connection with the target browser. A WebSocket connection is created and the communication link is managed to implement asynchronous message processing of the CDP protocol. A CDP session is created and bound to the browser target, thus binding the CDP session to the target page of the remote browser instance.

[0066] As described above in this manual, by enabling a remote browser instance, the target browser's execution environment and control logic are physically separated. This supports centralized management and resource pooling of browser instances, improving resource utilization. It provides a technical foundation for distributed deployment and load balancing. Creating CDP sessions ensures the accuracy and reliability of control. Furthermore, using a remote browser instance supports rich browser operation capabilities. It avoids compatibility and maintenance cost issues related to the target browser's proprietary protocols. On this basis, the remote browser instance runs in an independent environment with a different process than the target browser, avoiding interference between tasks. Furthermore, the WebSocket connection mentioned in this manual can be understood as a long-lived WebSocket connection; furthermore, long-lived WebSocket connections can reduce the overhead of frequent connection establishment.

[0067] Furthermore, in the case where the target interface layer provided above encapsulates an embedded processing path and a cloud remote path, in a specific embodiment provided in this specification, based on the login state, selecting the target data processing path in the target interface layer includes: When the user is logged in, the target data processing path is determined to be an embedded processing path in the target interface layer; and / or When the login state is not logged in, the target data processing path is determined to be a remote path in the cloud in the target interface layer.

[0068] In one specific implementation provided in this specification, an embedded path is used when logged in, and based on this, secure sessions existing in related technologies are used to prevent the leakage of sensitive information. When not logged in, a cloud path is used to ensure that potentially risky operations are executed in an isolated environment, protecting the target user's environment security within the browser process. Furthermore, an automated secure path selection mechanism reduces security vulnerabilities caused by human error, and the system architecture intelligently adapts to the target user's current login state. Logged-in users enjoy the low latency and high response speed of local processing. Not logged-in users benefit from the convenience and full functionality of cloud processing. Local resources are prioritized when logged in, reducing network transmission overhead and cloud computing costs. Cloud resources are utilized efficiently when not logged in. The overall system resource utilization is optimized based on the user's login state, establishing an automated path selection model based on login state to reduce manual decision-making costs. Dynamic adjustment and optimization of the path selection strategy are supported.

[0069] Based on the methods described above for using embedded paths while logged in and using remote cloud paths when not logged in, this specification employs the following methods to access target pages according to different target data processing paths.

[0070] Step 106: Access the target page based on the target data processing path and extract the page content of the target page.

[0071] The target page can be understood as the webpage from which automated operations need to be performed, such as airline websites, e-commerce websites, and email login pages.

[0072] The process of accessing a target page can be understood as a series of browser behaviors, including page navigation, element positioning, and interactive operations.

[0073] Extracting page content can be understood as obtaining relevant text, images, links, and other information from a target page.

[0074] For ease of understanding, this specification provides an exemplary description of the process of accessing a target page using an embedded path while logged in, using the following methods.

[0075] In one specific embodiment provided in this specification, accessing a target page based on the target data processing path includes: Send the instruction to be processed to the embedded frame element; The instruction to be processed is mapped to the native operation instruction of the target page through the protocol mapping interface in the target bridging script; Within the embedded frame element, the native operation instructions are executed to interact with the target page.

[0076] The embedded frame element can be understood as the created iframe sandbox environment, serving as an isolated container for command execution. The target bridging script can be understood as the core protocol conversion component running inside the iframe.

[0077] In one specific embodiment provided in this specification, a message broker mechanism can be used to send the instruction to be processed to the embedded frame element, thereby achieving cross-domain communication within the same process provided by the target browser. Based on this, an instruction transmission channel is established from the main page where the task to be processed resides to the inside of the iframe, using the HTML5 standard postMessage API to achieve secure cross-domain communication. The instruction to be processed is mapped to the native operation instructions in the target page according to the core interface adapted by the CDP protocol in the iframe environment. Various standard operation types (input, click, scroll, data extraction, etc.) are converted within the iframe. As described above, the protocol mapping interface ensures accurate conversion from abstract instructions to concrete operations. Native DOM manipulation simulates real user behavior. All operations are confined to the iframe sandbox, ensuring no impact on the main page's stability. Cross-domain communication is achieved through the standard postMessage method. Furthermore, by establishing a sandbox within the target browser to handle operations, abnormal operations are isolated within the sandbox environment, preventing them from propagating to the main system. Operation instructions based on the CDP protocol are transformed into native operation instructions for the target page within the iframe environment, ensuring compatibility with various front-end frameworks. Moreover, the target page's native operation instructions directly invoke the browser engine, resulting in high execution efficiency.

[0078] For ease of understanding, this specification provides an exemplary description of the process of accessing a target page using a remote cloud path while logged in, using the following methods.

[0079] In one specific embodiment provided in this specification, accessing a target page based on the target data processing path includes: The pending instruction is sent to the remote browser instance via the network socket connection; Based on the instruction to be processed, the remote browser instance is controlled to load and process the target page.

[0080] In this context, a network socket connection can be understood as a persistent, full-duplex communication channel based on the WebSocket protocol, connecting the control end and a remote browser instance. A remote browser instance can be understood as a browser process running on a cloud server, typically executing in headless mode. The instructions to be processed can be understood as a standardized, encapsulated sequence of browser operation commands. Control can be understood as programmatically operating the remote browser via the CDP protocol.

[0081] Furthermore, it should be understood that the loading and processing of remote browser instances mentioned in this specification can be understood as including the complete process of page navigation, content rendering, interactive operations, and data extraction for the target page. The target page can be understood as the specific webpage for which automated operations need to be performed.

[0082] In one specific implementation provided in this specification, standardized CDP commands are sent to a remote browser via a WebSocket connection. The remote browser instance receives and executes the complete process of the CDP commands, for example, implementing page loading status detection based on CDP event listeners; supporting complex interaction sequences, including form filling, element clicking, and scrolling operations; implementing an intelligent waiting mechanism to handle dynamically loaded content and asynchronous JavaScript; and providing complete result capture and error recovery capabilities.

[0083] As described above, WebSocket full-duplex communication reduces network latency and improves command transmission efficiency. The CDP protocol natively supports browser operations, resulting in a shorter execution path and faster performance. WebSocket persistent connections avoid the overhead and uncertainty of frequent connection establishment. The CDP protocol provides standardized error codes and exception handling mechanisms, enabling automatic retries and failover to ensure eventual consistency of task execution. Each remote browser instance runs in an independent environment, ensuring complete isolation between tasks. The cloud execution environment is isolated from the target browser process, enhancing the security of target browser information. Furthermore, as a browser industry standard, the CDP protocol ensures compatibility with different browser versions. WebSocket, as an IETF standard protocol, is supported by all modern programming languages ​​and platforms.

[0084] Step 108: Generate a page snapshot based on the page content, and determine the target result based on the page snapshot.

[0085] A page snapshot can be understood as a data representation of page content after it has been structured and semantically processed, containing structured data such as semantic information of page elements and interaction states.

[0086] The target result can be understood as the final processing result returned to the user or upper-layer application, which corresponds to the target task to be processed.

[0087] In one specific embodiment provided in this specification, generating a page snapshot based on the page content includes: Perform element analysis on the elements in the target page, and calculate the accessible name and interaction state of one or more elements in the target page; Based on the accessible names and interaction states, target interactive elements and information elements are identified and filtered out. The target interactive elements and information elements are structured to obtain structured information of the target interactive elements and information elements, and the structured information is serialized into a page snapshot in a predetermined format.

[0088] Element analysis can be understood as the process of systematically parsing and extracting features from HTML elements in the Document Object Model (DOM). Accessible names can be understood as the semantic identifiers of elements within an accessible context, calculated by combining multiple attributes. Interaction state can be understood as the operability and status indicators of elements in the current page environment. Target interactive elements can be understood as key interactive controls that users may need to operate on. Information elements can be understood as display elements such as text and images that carry important content information. Structured information can be understood as element attribute data organized according to a unified schema. Serialization can be understood as converting data structures in memory into a storable and transmittable format. A predefined format page snapshot can be understood as a standardized semantic representation format of the page.

[0089] Furthermore, the identification and filtering involved in this specification can be understood as selecting valuable page elements based on preset rules and algorithms.

[0090] In one specific implementation provided in this specification, accessible names are comprehensively calculated based on accessibility standards. Multi-dimensional detection of element visibility, usability, and focus status is performed. Dynamic state capture supports real-time state tracking for single-page applications. Repetitive boilerplate content such as navigation bars and footers is excluded. Key information in the main content area is prioritized. Adaptive filtering rules based on page type (e-commerce, news, forms, etc.) are supported. Multiple output formats are supported (YAML, JSON, Protocol Buffers). The hierarchical relationships and semantic connections of elements in the original page are maintained. Rich metadata is included to support subsequent AI understanding and processing. Efficient data compression is achieved.

[0091] For ease of understanding, this instruction manual combines... Figure 2 The content shown illustrates the process of generating a page snapshot.

[0092] Figure 2 An architecture diagram for generating page snapshots is shown according to one embodiment of this specification.

[0093] In one specific embodiment provided in this specification, a Document Object Model (DOM) tree traversal is performed on the target page. For example... Figure 2 As shown, it includes an element visibility detection module, an accessible name calculation model, an element role recognition module, an interactivity analysis model, and a structured snapshot output module.

[0094] In the element visibility detection module, recursively traversing the Document Object Model (DOM) nodes can be understood as a depth-first traversal of the entire DOM to ensure no nested elements are missed. Shadow DOM support can be understood as penetrating the boundaries of private DOM subtrees (Shadow Root) to access the internal structure of web components. Embedded frame element (iFrame) content traversal: cross-frame content collection, achieving full-page coverage.

[0095] Shadow Root can be understood as the core part of the Shadow DOM, a private DOM subtree isolated from the main document DOM tree.

[0096] In the accessibility model calculation module, style calculation caching can be understood as caching calculated style information to avoid redundant calculations and improve performance. Visible area detection can be understood as determining the actual visibility of an element based on its viewport position. Element position analysis can be understood as recording the relative position and coordinates of elements within the page layout.

[0097] In the element role recognition module, accessibility parsing (aria-label parsing) can be understood as directly reading ARIA tag attributes. Here, aria-label can be understood as an attribute used to improve webpage accessibility. Element tag association (aria-labelledby association) can be understood as referencing associated tag elements by ID. The aria-labelledby attribute is used to identify other elements with the current element tag. Native webpage (HTML) semantics: based on the default semantic roles of HTML tags. Content text extraction is used to extract meaningful text from element content.

[0098] In the interactivity analysis module, click event detection can be understood as analyzing whether an element has a click event handler bound to it. Cursor style analysis can be understood as determining interactivity through the CSS cursor property. Form element recognition can be understood as recognizing all form input controls. State attribute detection can be understood as checking state attributes such as disabled and readonly.

[0099] Based on the above content in this specification, it utilizes accessibility standards to accurately understand the functional semantics of elements. It goes beyond surface text to deeply understand the interactive intent and meaning of elements. It provides AI with rich contextual information, significantly improving task execution accuracy. It filters redundant content, and structured serialization reduces data volume, improving transmission and processing efficiency. It supports incremental updates, transmitting only key elements that have changed. It comprehensively identifies truly interactive elements based on multiple features, avoiding misjudgments and omissions. It accurately captures element states, ensuring that operation instructions are executed at the correct time. It supports complex interaction scenarios, such as form filling and multi-step operations. The structured snapshot format facilitates understanding and processing by the intelligent agent processing unit.

[0100] Based on the data processing methods described above in this specification, the unified management and scheduling of target browser automation technology is achieved through the abstract design of the target interface layer. Within the target browser, the target execution path can be intelligently selected based on real-time conditions according to the target task, improving the success rate and efficiency of task execution. Furthermore, the target interface layer ensures that each data processing path does not require modification of the upper-layer logic, and the encapsulated architecture guarantees independent control over each data processing path.

[0101] For ease of understanding, this instruction manual combines... Figure 3 The system architecture diagram shown above, combined with the search instances of the target user for the task to be processed, explains the above data processing method.

[0102] Figure 3 A system architecture diagram according to one embodiment of this specification is shown.

[0103] In one specific embodiment provided in this specification, the target user inputs a task to be processed in the user page of the AI ​​agent in the target browser. The target browser integrates the user page of the AI ​​agent and controlled embedded pages.

[0104] Furthermore, the AI ​​agent's user page integrates the target browser's built-in debugging interface (JSDebugger API). This built-in debugging interface is used to monitor and control the AI ​​agent's user page behavior in real time. The AI ​​agent's user page also integrates an automation tool (Playwright), which provides the AI ​​agent (e.g., an AI assistant) with the ability to perform page operations.

[0105] Embedded pages integrate computing environments (e.g., C++ Native environments) to handle complex tasks. They also integrate content extraction units (content extraction SDKs) to extract specific information from the target page. Furthermore, embedded pages integrate a Document Object Model (DOM) to provide a structure tree of the target page, representing how the webpage content is organized.

[0106] It is important to understand that inter-process communication (IPC) is used between the user page and the embedded page of the AI ​​agent.

[0107] exist Figure 3 The system architecture diagram shown also includes a server-side browser (independent of the target browser process, determined via a remote path in the cloud). The server-side browser integrates controlled pages and a remote browser (headless).

[0108] The controlled page and the embedded page also integrate a computing environment, content extraction unit, and document object model, which will not be elaborated further in this specification. A debugging interface is integrated into the remote browser to allow remote control of the intelligent agent unit.

[0109] It's important to understand that in a remote browser, the controlled page and the remote browser also communicate through inter-process communication.

[0110] exist Figure 3 The system architecture diagram shown also includes an AI agent, which integrates an intelligent processing unit (Agent Server) and an automation tool model context protocol (Playwright MCP). The model context protocol is used to coordinate automated tasks. Furthermore, the AI ​​agent also integrates a large language model (LLM).

[0111] The intelligent processing unit communicates with the embedded page in the target browser via a proprietary protocol. It also communicates with the remote browser via the CDP protocol.

[0112] based on Figure 3 The system architecture described above is explained in this specification using a search example of the target user's task to be processed to illustrate the data processing method.

[0113] For example, consider an embedded processing path. The user inputs a task to be processed through an AI assistant. In the intelligent processing unit, a large language model understands the task and generates an automated operation sequence (i.e., the instructions to be processed). The task is coordinated via Playwright MCP. In the user's page of the AI ​​agent in the target browser, the instructions to be processed are transmitted to the embedded page via IPC, where the content of the target page is extracted within the embedded page's computing environment.

[0114] For example, consider a remote path in the cloud. The user inputs a task to be processed through an AI assistant. In the intelligent processing unit, a large language model understands the task and generates an automated operation sequence (i.e., instructions to be processed). Communication with a remote browser is established via the CDP protocol, and the automated processing is executed on a controlled page within the remote browser. The specific operation details can be found in the embedded page's operation process to obtain the target result. The remote browser then transmits the target result to the user's page on the AI ​​agent via a proprietary protocol.

[0115] Furthermore, in this specification, in conjunction with Figure 4 The data processing diagram shown is for... Figure 3 The processing logic of the system architecture shown is illustrated with examples.

[0116] Figure 4 A schematic diagram of a data processing method provided according to one embodiment of this specification is shown.

[0117] exist Figure 4 In this process, tasks to be processed are retrieved from the user page of the AI ​​agent. Interaction between the intelligent processing unit and the target browser kernel is established via the MCP protocol. The target browser's toolkit is then obtained at the target browser's tool layer. Finally, a unified interface for the embedded processing path and the cloud remote path is encapsulated at the target interface layer.

[0118] Building upon this, an embedded frame element is created within the target browser's kernel layer for the embedded processing path. The target bridging script is then injected into the embedded page corresponding to the embedded frame element. In the protocol adaptation layer, the protocol mapping interface is adapted, mapping the instructions to be processed to the native operation instructions of the target page.

[0119] The embedded page integrates an accessible element calculation module, an element state analysis module, and a page snapshot generation module.

[0120] Furthermore, the target browser's kernel layer integrates a Developer Tools Protocol (CDP) endpoint for the cloud-based remote path, providing a CDP protocol and remote access interface. The cloud-based remote path also integrates a WebSocket connection to facilitate real-time communication between the target page and the intelligent processing unit.

[0121] Furthermore, a protocol adaptation layer is integrated into the browser kernel layer to facilitate the maintenance of a stable connection between the intelligent processing unit and the target browser.

[0122] The protocol adaptation layer integrates path interface adaptation and developer tool protocol conversion. Path interface adaptation includes interface adaptation for embedded processing paths and interface adaptation for remote cloud paths. Developer tool protocol conversion (CDP protocol conversion) can be understood as a process used to convert the CDP protocol into a format used within the system architecture.

[0123] Furthermore, a page control layer is integrated into the browser kernel layer, and a snapshot generator is integrated into the page control layer. Document Object Model (DOM) manipulation tools are also integrated into the page control layer. For details on the snapshot generator's functionality, please refer to the above-mentioned sections of this manual. Figure 2 The contents shown are not repeated here.

[0124] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 5 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 5 As shown, the device includes: The acquisition unit 502 is configured to acquire the target task to be processed and determine the instruction to be processed based on the target task to be processed; Selection unit 504 is configured to select a target data processing path in the target interface layer based on the instruction to be processed, wherein the target interface layer encapsulates at least one data processing path; Extraction unit 506 is configured to access the target page based on the target data processing path and extract the page content of the target page; The determining unit 508 is configured to generate a page snapshot based on the page content and determine the target result based on the page snapshot.

[0125] Furthermore, the target interface layer encapsulates at least an embedded processing path and a cloud remote path.

[0126] Furthermore, the embedded processing path is constructed in the following manner: Create an embedded frame element in the target browser; The target bridging script is injected into the embedded page corresponding to the embedded frame element, wherein the target bridging script includes a protocol mapping interface; Based on the instruction to be processed, control the embedded frame element to navigate to the target page.

[0127] Furthermore, selection unit 504 is also configured as follows: Send the instruction to be processed to the embedded frame element; The instruction to be processed is mapped to the native operation instruction of the target page through the protocol mapping interface in the target bridging script; Within the embedded frame element, the native operation instructions are executed to interact with the target page.

[0128] Furthermore, the cloud-based remote path is constructed in the following manner: Start a remote browser instance, wherein the remote browser instance is used to enable the developer tools protocol debugging port; Obtain the Uniform Resource Locator (URL) for the network socket provided by the remote browser instance; Based on the network socket Uniform Resource Locator, establish a network socket connection with the remote browser instance; Based on the network socket connection, a developer tools protocol session is created and attached to the remote browser instance.

[0129] Furthermore, selection unit 504 is also configured as follows: The pending instruction is sent to the remote browser instance via the network socket connection; Based on the instruction to be processed, the remote browser instance is controlled to load and process the target page.

[0130] Furthermore, selection unit 504 is also configured as follows: Identify the target user corresponding to the target task to be processed, and determine at least one login status corresponding to the target user in the target browser; Based on the login status, select the target data processing path in the target interface layer; The login status includes a logged-in status and a logged-out status; Selection unit 504 is also configured as follows: When the user is logged in, the target data processing path is determined to be an embedded processing path in the target interface layer; and / or When the login state is not logged in, the target data processing path is determined to be a remote path in the cloud in the target interface layer.

[0131] Furthermore, unit 508 is also configured as follows: Perform element analysis on the elements in the target page, and calculate the accessible name and interaction state of one or more elements in the target page; Based on the accessible names and interaction states, target interactive elements and information elements are identified and filtered out. The structured information of the target interactive elements and information elements is serialized into a page snapshot in a predetermined format.

[0132] Furthermore, the acquisition unit 502 is also configured as follows: Obtain task information sent by the target user; The task information is analyzed by the intelligent processing unit to obtain the target task to be processed. The target task to be processed is parsed based on the instruction parsing model to obtain the initial instruction to be processed; Based on the model context interaction protocol, the initial instruction to be processed is encapsulated into a standardized protocol message; The standardized protocol message is identified as a pending instruction, wherein the pending instruction is used to guide the target data processing path to operate on the target page.

[0133] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.

[0134] See Figure 6 , Figure 6 This specification illustrates an architecture diagram of a data processing system according to one embodiment of the present specification. The data processing system may include a client 100 and a server 200. Client 100 is used to send a target task to be processed to server 200 and determine a processing instruction based on the target task; Server 200 is configured to: select a target data processing path in the target interface layer based on the instruction to be processed, wherein the target interface layer encapsulates at least one data processing path; access a target page based on the target data processing path and extract the page content of the target page; generate a page snapshot based on the page content and determine the target result based on the page snapshot; and send the target result to client 100. Client 100 is also used to receive the target result sent by server 200.

[0135] Using the solutions described in the embodiments of this specification, the data processing system may include multiple clients 100 and a server 200. The clients 100 may be referred to as end-side devices, and the server 200 may be referred to as cloud-side devices. Multiple clients 100 can establish communication connections through the server 200. In the data processing scenario, the server 200 is used to provide data processing services between the multiple clients 100. Each client 100 can act as a sender or receiver, communicating through the server 200.

[0136] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In data processing scenarios, users can publish data streams to server 200 through client 100, server 200 can generate target results based on the data stream, and push the target results to other clients that have established communication.

[0137] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.

[0138] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0139] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0140] It should be understood that the data processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the data processing methods provided in the embodiments of this specification. In other embodiments, the data processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0141] Figure 7 A block diagram of a computing device structure according to an embodiment of this application is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.

[0142] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0143] In one embodiment of this application, the aforementioned components of the computing device 700 and Figure 7 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 7 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0144] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.

[0145] The processor 720 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described data processing method.

[0146] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the data processing method described above.

[0147] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0148] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.

[0149] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0150] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the data processing method described above.

[0151] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0152] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0153] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0154] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0155] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, applied to a target browser, comprising: Obtain the target task to be processed, and determine the instruction to be processed based on the target task; Based on the instruction to be processed, a target data processing path is selected in the target interface layer, wherein the target interface layer encapsulates at least one data processing path; Access the target page based on the target data processing path and extract the page content of the target page; A page snapshot is generated based on the page content, and the target result is determined based on the page snapshot.

2. The method as described in claim 1, wherein the target interface layer encapsulates at least an embedded processing path and a cloud remote path.

3. The method of claim 2, wherein the embedded processing path is constructed in the following manner: Create an embedded frame element in the target browser; The target bridging script is injected into the embedded page corresponding to the embedded frame element, wherein... The target bridging script includes a protocol mapping interface; Based on the instruction to be processed, control the embedded frame element to navigate to the target page.

4. The method as described in claim 3, wherein accessing the target page based on the target data processing path includes: Send the instruction to be processed to the embedded frame element; The instruction to be processed is mapped to the native operation instruction of the target page through the protocol mapping interface in the target bridging script; Within the embedded frame element, the native operation instructions are executed to interact with the target page.

5. The method as described in claim 2, wherein the cloud-based remote path is constructed in the following manner: Start a remote browser instance, where, The remote browser instance is used to enable the developer tools protocol debugging port; Obtain the Uniform Resource Locator (URL) for the network socket provided by the remote browser instance; Based on the network socket Uniform Resource Locator, establish a network socket connection with the remote browser instance; Based on the network socket connection, a developer tools protocol session is created and attached to the remote browser instance.

6. The method of claim 5, wherein accessing the target page based on the target data processing path includes: The pending instruction is sent to the remote browser instance via the network socket connection; Based on the instruction to be processed, the remote browser instance is controlled to load and process the target page.

7. The method as described in claim 2, wherein selecting a target data processing path in the target interface layer based on the instruction to be processed includes: Identify the target user corresponding to the target task to be processed, and determine at least one login status corresponding to the target user in the target browser; Select the target data processing path at the target interface layer based on the login status; The login status includes a logged-in status and a logged-out status; The selection of the target data processing path at the target interface layer based on the login status includes: When the user is logged in, the target data processing path is determined to be an embedded processing path in the target interface layer; and / or When the login state is not logged in, the target data processing path is determined to be a remote path in the cloud in the target interface layer.

8. The method according to any one of claims 1 to 7, wherein generating a page snapshot based on the page content comprises: Perform element analysis on the elements in the target page, and calculate the accessible name and interaction state of one or more elements in the target page; Based on the accessible names and interaction states, target interactive elements and information elements are identified and filtered out. The structured information of the target interactive elements and information elements is serialized into a page snapshot in a predetermined format.

9. The method of claim 1, wherein obtaining a target task to be processed and determining a processing instruction based on the target task to be processed, comprises: Obtain task information sent by the target user; The task information is analyzed by the intelligent processing unit to obtain the target task to be processed. The target task to be processed is parsed based on the instruction parsing model to obtain the initial instruction to be processed; Based on the model context interaction protocol, the initial instruction to be processed is encapsulated into a standardized protocol message; The standardized protocol message is identified as a pending instruction, wherein the pending instruction is used to guide the target data processing path to operate on the target page.

10. A data processing method, applied to a cloud-side device, comprising: The receiving end device acquires the target task to be processed and determines the instruction to be processed based on the target task to be processed; Based on the instruction to be processed, a target data processing path is selected in the target interface layer, wherein the target interface layer encapsulates at least one data processing path; Access the target page based on the target data processing path and extract the page content of the target page; A page snapshot is generated based on the page content, and the target result is determined based on the page snapshot; The target result is sent to the end-side device.

11. A task platform, comprising a request interface and a response unit; The request interface is used to obtain the target task to be processed and to determine the instruction to be processed based on the target task to be processed; The response unit is used to select a target data processing path in the target interface layer based on the instruction to be processed, wherein the target interface layer encapsulates at least one data processing path. Access the target page based on the target data processing path and extract the page content of the target page; generate a page snapshot based on the page content, and determine the target result based on the page snapshot.

12. A data processing apparatus, applied to a target browser, comprising: The acquisition unit is configured to acquire the target task to be processed and determine the instruction to be processed based on the target task to be processed; The selection unit is configured to select a target data processing path in the target interface layer based on the instruction to be processed, wherein the target interface layer encapsulates at least one data processing path. The extraction unit is configured to access the target page based on the target data processing path and extract the page content of the target page; The determining unit is configured to generate a page snapshot based on the page content and determine the target result based on the page snapshot.

13. A computing device, comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9.

14. A computer-readable storage medium storing a computer program or instructions that, when executed by a processor, implement the steps of the method of any one of claims 1 to 9.

15. A computer program product comprising a computer program or instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.