Task execution method and device, electronic equipment and medium
By using image recognition and non-intrusive interaction technologies, the system simulates human vision to identify and perform operations in target applications, solving the problem that existing technologies cannot adapt to complex operations across platforms and achieving efficient and stable automated process execution.
Patent Information
- Application Number
- CN202511835550.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-03
AI Technical Summary
In applications such as intelligent customer service and user outreach, existing technologies cannot implement actions through open APIs or protocols, and existing automation solutions have poor scalability and cannot adapt to complex operations on the application interface.
By using image recognition technology, it simulates human vision to identify and perform operations on the display page of the target application. It adopts a non-intrusive interactive method to achieve cross-platform automated process execution, uses preset templates and image matching algorithms to locate target operations, and adapts to different operating systems through a script library.
It achieves cross-platform, highly versatile automated process execution, improves the continuity and efficiency of operation, adapts to changes in display parameters in different terminal environments, and enhances the stability and flexibility of task execution.
Smart Images

Figure CN121597330A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computers, and more particularly to the fields of process automation and intelligent agents, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for performing tasks. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] As enterprises become increasingly digitalized, many business processes rely on manual operations through instant messaging tools, desktop clients, and browser backends. Especially in applications such as intelligent customer service and user outreach, more and more actions cannot be implemented via open APIs or protocols, requiring manual clicking, input, and judgment within the application interface. Alternatively, some automation solutions platforms have strong limitations and poor scalability.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for performing a task.
[0006] According to one aspect of this disclosure, a method for performing a task is provided, comprising: in response to receiving a trigger request for a target task, determining a target application corresponding to the target task; determining a preset template corresponding to the target task, wherein the template includes a first image of a first object identifier for performing the target task, wherein the first object identifier corresponds to a target operation to be performed in the target application; based on the first image, identifying a second object identifier corresponding to the target operation on a display page of the target application by image recognition; and performing the target operation on the display page based on the second object identifier to perform the target task.
[0007] According to another aspect of this disclosure, an apparatus for performing a task is provided, comprising: a response unit configured to determine a target application corresponding to the target task in response to receiving a trigger request for the target task; a determination unit configured to determine a preset template corresponding to the target task, wherein the template is used to execute a first image of a first object identifier of the target task, wherein the first object identifier corresponds to a target operation to be performed in the target application; an identification unit configured to identify a second object identifier corresponding to the target operation on a display page of the target application based on the first image through image recognition; and an execution unit configured to execute the target operation on the display page based on the identified second object identifier to perform the target task.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described in this disclosure.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.
[0011] According to one or more embodiments of this disclosure, non-intrusive interaction based on image recognition simulates human vision for positioning and operation, thereby achieving cross-platform, highly versatile automated process execution.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown; Figure 2 A flowchart illustrating a method for performing a task according to an embodiment of the present disclosure is shown; Figure 3 A schematic diagram of a system architecture for implementing a method for performing a task according to an embodiment of the present disclosure is shown; Figure 4 A structural block diagram of an apparatus for performing a task according to an embodiment of the present disclosure is shown; and Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0016] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0017] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0018] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0019] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0020] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of methods for performing tasks.
[0021] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.
[0022] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0023] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to trigger target tasks or receive data related to the execution of target tasks. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0024] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0025] Network 110 can be any type of network well known to those skilled in the art, and can support data communication using any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0026] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0027] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0028] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0029] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0030] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as templates, users, etc. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0031] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0032] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0033] Robotic Process Automation (RPA) is an application software technology that uses software programs to simulate human actions on a computer, automatically executing repetitive and routine tasks according to preset rules. In traditional business processes, a large number of processes rely on manual repetitive operations performed by humans in instant messaging tools, desktop clients (such as CRM customer management system clients), or browser backgrounds. This is not only inefficient but also prone to human error. Another type of existing automation solution often heavily relies on the low-level interfaces of specific software. Once the target application version is updated or the interface is not open, the automation platform becomes limited and lacks scalability.
[0034] Therefore, an embodiment of the present disclosure provides a method for performing a task. Figure 2 A flowchart illustrating a method for performing a task according to an embodiment of the present disclosure is shown, such as... Figure 2 As shown, method 200 includes: in response to receiving a trigger request for a target task, determining a target application corresponding to the target task (step 210); determining a preset template corresponding to the target task, wherein the template includes a first image of a first object identifier for executing the target task, wherein the first object identifier corresponds to a target operation to be executed in the target application (step 220); based on the first image, identifying a second object identifier corresponding to the target operation on the display page of the target application by image recognition (step 230); and executing the target operation on the display page based on the second object identifier to execute the target task (step 240).
[0035] According to embodiments of this disclosure, non-intrusive interaction based on image recognition simulates human vision for positioning and operation, thereby achieving cross-platform, highly versatile automated process execution.
[0036] In this disclosure, the target task may refer, for example, to a business process with clearly defined start and end conditions that the user expects the automation system to execute, while the target application is the software platform that hosts the execution of this business process. The trigger request refers to a signal that activates the execution of the task, typically generated by time conditions, data changes, or external system calls.
[0037] As an example, taking an internal enterprise information distribution scenario, the target task could be "sending a preset notification to a designated contact," and the corresponding target application could be a general enterprise instant messaging software. A trigger request is generated when the business system detects the completion of a specific business node (e.g., an order status update).
[0038] In this disclosure, a template can refer to a pre-configured set of data describing the operations required to complete the target task; the operations can refer to specific operations simulating user input devices (such as mouse clicks, keyboard input, dragging); and the object identifier can refer to a control entity (such as a button, input box, or label) that carries interactive functions in a graphical user interface. Therefore, the first image can be a reference screenshot representing the aforementioned controls, captured and saved in advance during the template configuration stage (the controls in the template serve as the first object identifier).
[0039] In the embodiments of this disclosure, the target operation for achieving the target task can be multiple target operations, that is, it can be a combination of a series of operations, which are executed sequentially in a predetermined execution order.
[0040] As an example, in a business supply chain management scenario, the target task could be "querying and exporting logistics order details with a specific number," and the corresponding target application could be an "Enterprise Resource Planning (ERP) system client." In this case, the preset template could define three stages such as "login," "search," and "export." The "search" stage includes an action of "clicking the search icon." To execute this action, the system can load the first image corresponding to the search icon (i.e., a pre-captured screenshot of the magnifying glass icon) as a feature to search for a matching second object identifier on the display screen. Once a match is found, the click action can be executed.
[0041] In embodiments according to this disclosure, image recognition operations between the first image and the display page of the target application can be performed using any suitable visual recognition algorithm, including but not limited to OpenCV, OpenCV-wasm, and other technologies, to achieve high-precision template matching and image processing.
[0042] In the embodiments of this disclosure, the display page of the target application can be a visual graphical user interface (GUI) rendered to the user in real time through the terminal device's display screen during the application's operation. This can include not only the static main window after the application starts, but also dynamic secondary pop-ups, floating layers, drop-down menus, or system prompts that are generated as business logic progresses. For example, in the operation of the aforementioned ERP system, the display page can be a main interface with a navigation bar; when the system performs the "export details" action, the display page may become a modal dialog box, requiring the selection of a file save path. By continuously capturing the current display page and using the aforementioned first image to perform pixel-level feature matching on these pages, the automated system ensures that regardless of page transitions, it can accurately locate the key elements required for the next operation, thereby improving the coherence and efficiency of the execution of each target operation.
[0043] In some embodiments, when multiple target tasks are obtained (these multiple target tasks may correspond to the same target application or different target applications), a corresponding priority can be preset for each target task to implement the target task based on the priority pair. For example, the system dynamically inserts tasks into the execution queue according to task type, priority, and execution conditions to ensure that high-priority tasks can be processed first, while avoiding concurrent conflicts or duplicate operations.
[0044] According to some embodiments, method 200 may further include: in response to receiving a trigger request for a target task, determining the system of the device where the first client is located; and determining the script corresponding to the system in a preset script library, so as to execute the target task based on the script corresponding to the system, wherein the script library includes multiple scripts corresponding to each system.
[0045] In the above embodiments, the use of a system adaptation mechanism based on script libraries can significantly improve the cross-platform compatibility and stability of automation methods.
[0046] In some embodiments, a system adaptation layer can be built to abstract underlying capabilities such as opening applications, controlling windows, and querying status into a unified interface, effectively shielding the differences between different operating systems in window management, screen scaling, input method status, and clipboard mechanisms. This means that the core business logic at the upper layer does not need to concern itself with the specific implementation details of the underlying platform, thereby ensuring that automated processes can have consistent execution results in different environments.
[0047] For example, the operating system of the first client device can be Windows, macOS, or Linux. After determining the device system, scripts corresponding to the system can be used to implement parameters such as opening applications, obtaining system themes (light or dark mode), system language interfaces, screen resolution, and scaling ratios. Accordingly, the preset script library is a collection of instructions that implement the same functions for different operating systems. For example, to perform the same actions such as application launch, window showing / hiding, theme reading, or process querying, the script library will call PowerShell or command-line scripts in a Windows environment, and AppleScript or Shell commands in a macOS environment. In this way, the system can automatically match and execute the correct underlying instructions based on the identified device environment, improving the stability and scalability of the target task execution.
[0048] According to some embodiments, after responding to receiving a trigger request for a target task and determining the target application corresponding to the target task, the method further includes: determining and adjusting the running state of the target application so that its running state is in a ready state capable of fulfilling the target task. The running state includes at least one of the following: the target application startup state, the target application account login state, the target application window activation state, and the target application page display state.
[0049] In the above embodiments, determining the target application corresponding to the target task involves not only locking the application object, but also adjusting and determining the running state of the target application so that it can be used to achieve the target task.
[0050] For example, if the target application is not running, its window is obscured, or the user is not logged in, it often directly leads to the inability to complete the target task. Therefore, a standardized execution environment can be established before task execution. For instance, before task execution, the system can automatically check whether the target application is running, whether the window is visible, whether the focus is active, and whether the user is logged in to determine whether subsequent actions can continue. If the target application is detected as not running, the system can automatically attempt to start the application and dynamically monitor its startup status during the startup process to ensure the task execution environment is ready. For windows that are obscured or lack focus, the system will programmatically bring the target window (i.e., the target application's window) to the foreground and activate its focus to ensure the validity of subsequent actions such as mouse and keyboard operations. In addition, the system can also detect the user's login status; if the user is not logged in, a login detection mechanism can be triggered, including automatically filling in account information or prompting the user to log in, thereby ensuring the continuity of task execution.
[0051] Through the above embodiments, system-level anomalies can be proactively identified and automatically repaired before task execution, effectively avoiding task failures caused by insufficient environmental preparation.
[0052] In some embodiments, the page display state of the target application can be predetermined. This display state can be information such as the target application's theme (e.g., light or dark mode), language, version, scaling ratio, etc. In some embodiments, when determining the running state of the target application, the "page display state" can be adjusted, which specifically includes, but is not limited to: system theme (e.g., ensuring the application is in a preset light or dark mode to avoid color differences affecting image matching confidence), language interface (ensuring the interface text is consistent with the template's preset language), scaling ratio and resolution (preventing coordinate offset due to DPI differences), and application version (ensuring the function entry layout is compatible with the template). Through comprehensive control of the above states, this embodiment can ensure that the target application is always in the "ready state" most suitable for automated execution.
[0053] Typically, users may not expect the page display state of their target application to be adjusted. Therefore, according to some embodiments, determining the preset template corresponding to the target task may include: based on the page display state of the target application, determining the preset template corresponding to the target task from the preset template library corresponding to the target application, wherein the preset template library includes multiple templates for adapting to different page display states of the target application, wherein the page display state includes at least one of the following: page resolution, page scaling ratio, the theme of the target application, the language of the target application, and the version of the target application.
[0054] In the above embodiments, the preset template corresponding to the target task is determined using a dynamic matching mechanism. That is, based on the detected page display state of the target application, a template (or set of templates) that matches the current page display state is selected from the preset template library. By constructing multiple templates corresponding to different page display states in this way, the robustness and recognition accuracy of the automation method in heterogeneous display environments can be greatly improved.
[0055] In traditional vision-based automation technologies, a single image template often struggles to handle the diversity of display parameters on terminal devices. Changes in screen pixel density or switching of interface styles can cause a mismatch between the pre-stored feature map and the actual screenshot, leading to task interruption. However, this disclosure addresses this by pre-setting multiple templates, enabling adaptive and accurate identification of interface elements in different environments.
[0056] Specifically, the preset template library encompasses adaptation schemes for various page display states. For example, firstly, the library can pre-set templates adapted to different page resolutions (e.g., standard 1080p, high-definition 2K, and ultra-high-definition 4K resolutions) and different scaling ratios (e.g., the common 100%, 125%, or 150% display scaling in Windows systems), ensuring that image size characteristics are accurately matched regardless of whether the screen is high-resolution or ordinary. Secondly, to address differences in application personalization configurations and fully accommodate the target application's theme settings and language environment, image templates with different contrast ratios can be further reserved for "light mode," "dark mode," and different language environments. Furthermore, considering the continuous updating nature of software, the template library can also include compatible templates for minor UI differences (such as minor adjustments to icon positions or changes in rounded corners) caused by different application versions. By locking these current state parameters and loading the corresponding dedicated templates before task execution, the system can eliminate visual interference and ensure pixel-level matching accuracy for subsequent image recognition.
[0057] According to some embodiments, the method is applied to a first client. The method further includes at least one of the following: receiving a trigger request sent by a target agent, the trigger request being generated and sent to the first client by the target agent when it detects trigger information for the target task during a call with a second client; receiving the trigger request generated based on a detected preset event associated with the target task; and generating the trigger request in response to detecting an operation for the second client to trigger the target task.
[0058] According to the above embodiments, trigger requests for target tasks can be obtained through a multimodal interaction monitoring mechanism, which may include at least one of intent triggering based on agent dialogue, automatic triggering based on associated events, and manual operation triggering. Adopting this diversified triggering mechanism can significantly improve the flexibility and response speed of automated systems in complex business scenarios.
[0059] In some embodiments, the second client may refer to the terminal device used by the external user who initiates the consultation or service (such as the user's mobile APP or web page), the first client refers to the terminal device used to process business and automate the task process (such as the office computer of customer service or operations personnel), and the target intelligent agent is a virtual service assistant (such as an online customer service robot, voice assistant, etc.) that can communicate with the second client.
[0060] Correspondingly, the triggering mechanism combines user inquiries with platform service interaction scenarios. For example, a second client user engages in real-time question-and-answer interaction with an online customer service robot via text or voice. For instance, a user inquires, "I want to speak with a contact to learn about your latest XXX." After understanding the statement and recognizing the intent to further communicate with the contact, the target agent generates a trigger request and sends it to the first client where that contact resides. At this point, the system immediately activates the instant messaging software on the first client, automatically performing actions such as adding a friend.
[0061] In some embodiments, the trigger request may also be triggered based on a preset event associated with the target task. For example, a "preset event" can refer to a state change in a business process that meets specific conditions. As an example, in an e-commerce scenario, when the backend system detects the event "new user registration successful" (i.e., the preset event), the server automatically pushes a task to the automated client (the first client) deployed on the operator's computer, driving it to automatically enter the new user's profile information or drive an email client to send a welcome message, thus achieving an event-driven workflow without manual intervention.
[0062] In some embodiments, the trigger request can also be initiated proactively by the first client. A typical scenario is "proactive service initiated by staff." For example, staff proactively initiate tasks through the backend system. This allows platform staff to perform complex actions for second client users in a batch and in a standardized manner with simple clicks.
[0063] According to some embodiments, based on the first image, identifying the second object identifier corresponding to the target operation on the display page of the target application by image recognition includes: obtaining a second image including the display page of the target application by performing a screenshot operation; performing a preprocessing operation on the second image to calculate the similarity between the preprocessed second image and the first image, and identifying the second object identifier corresponding to the target operation on the display page of the target application based on the similarity, wherein the preprocessing operation includes at least one of the following: grayscale processing, matrix processing, binarization processing, edge detection, and contrast enhancement.
[0064] In some embodiments, a second image containing the target application's display page can be obtained through a screenshot operation. This second image is then subjected to a series of preprocessing operations, and finally, its similarity is calculated with a first image used as a template. These preprocessing operations effectively overcome the interference caused by screen rendering differences on automated execution, significantly improving the accuracy and computational efficiency of target recognition.
[0065] Specifically, the second image can be a full-screen screenshot of the current terminal device screen, or a screenshot of a region limited to the target application window. For example, when a "Confirm" button needs to be clicked, the sample of that button is the first image, while the screenshot of the main interface of the application software containing that button is the second image. The system needs to locate the coordinate region that matches the first image within this larger search area of the second image.
[0066] To improve the accuracy of this matching process, after acquiring the second image, it can be preprocessed based on at least one of the following: grayscale conversion, which converts a three-channel color image into a single-channel grayscale image to eliminate interference from color changes and focus on the comparison of brightness and shape information; matrix conversion, which converts image data into a numerical matrix that is easy for computers to calculate quickly, accelerating subsequent convolution or correlation calculations; binarization, which forces image pixels to be divided into black and white by setting a threshold, which can greatly enhance the isolation between foreground and background when recognizing text or simple icons; edge detection (such as Canny or Sobel operators), which is used to extract contour lines with drastic brightness changes in the image, which is particularly effective for recognizing controls whose color changes with mouse hover but whose shape remains unchanged; and contrast enhancement, which stretches the histogram distribution of the image, making the features of blurred controls in low-contrast interfaces clearer and more distinguishable.
[0067] By combining the above preprocessing methods, the system can construct a feature map that is most conducive to algorithm matching, thereby ensuring the accuracy of each automatic operation.
[0068] According to some embodiments, based on the first image, identifying a second object identifier corresponding to the target operation on the display page of the target application by image recognition includes: obtaining a second image including the display page of the target application by performing a screenshot operation; calculating the similarity between the second image and the first image to identify a second object identifier corresponding to the first object identifier in the first image on the display page of the target application based on the similarity; and in response to not identifying a second object identifier corresponding to the corresponding first object identifier, performing optical character recognition on the first image and the second image of the corresponding first object identifier to determine a second object identifier corresponding to the corresponding first object identifier based on the identified characters.
[0069] In the above embodiments, a hybrid positioning mechanism that combines visual matching and semantic recognition is introduced. Specifically, one can first attempt to locate an object based on the similarity between the first image and the second image; if the matching fails (for example, the similarity does not reach the preset threshold), then one can further perform optical character recognition (OCR) on both the second image obtained by taking a screenshot and the first image used as a template, and use the recognized text information to complete the final object positioning.
[0070] In some embodiments, "similarity calculation" may refer to the process of calculating the similarity scores of the preset first image (template) at various positions in the second image (screenshot) based on computer vision algorithms. For example, one can adopt the normalized cross-correlation (NCC) or template matching algorithm, perform a sliding window scan on the screenshot, and search for the region with the pixel distribution feature closest to the template. When the calculated highest similarity score exceeds the preset confidence threshold (for example, 90%), it is determined that the recognition is successful.
[0071] However, when the object is not recognized by the similarity calculation (that is, the matching scores of all regions are lower than the threshold), the system will execute the step of "performing optical character recognition on the first image and the second image". For example, the system performs a full-screen or partial OCR scan on the current screenshot (the second image) to retrieve whether there is the same text content. Exemplarily, if the two characters "Confirm" identical to those in the first image are found on the screen, the system directly obtains the coordinate position of this text on the screen and uses it as the actual position of the target object.
[0072] Adopting this "image + text" dual-verification strategy significantly improves the fault tolerance and adaptability of the automated method when facing complex UI changes. In practical applications, the interface of the target application often undergoes fine-tuning, such as button color changes, icon style flattening, or background transparency adjustment. These changes often lead to pixel-based image matching failures. However, in this embodiment, by introducing the OCR mechanism, even if the visual style of the control changes, accurate positioning can still be achieved through character recognition. This mechanism effectively solves the problem of task interruption caused by UI fine-tuning by dynamically adjusting the matching strategy, ensuring the continuous availability of the automated script after the application version is updated.
[0073] In some embodiments, when determining the second object identifier corresponding to the corresponding first object identifier based on the recognized characters, the recognized characters can be fully matched, or the semantic similarity between the recognized characters can be further calculated to enable correct recognition even when the text content changes, significantly improving the fault tolerance of the recognition.
[0074] According to some embodiments, performing the target operation on the display page to perform the target task includes: determining the screen coordinates corresponding to the identified second object identifier; converting the screen coordinates corresponding to the identified second object identifier into viewport coordinates corresponding to the display page through a coordinate transformation operation; and performing the target operation for the second object identifier based on the viewport coordinates.
[0075] In the above embodiments, the screen coordinates refer to the position data in a global Cartesian coordinate system established with the physical display screen of the terminal device, such as the upper left corner as the origin (0,0). It is the raw result returned by the image recognition engine after analyzing the full-screen screenshot, representing the physical position of the control on the entire desktop. The viewport coordinates, on the other hand, refer to the position in a local coordinate system established with the actual rendering area inside the target application (i.e., the viewport, such as the Webview area of a browser or the content panel of an application) as the origin (e.g., the upper left corner).
[0076] Through the above embodiments, precise coordinate mapping logic is introduced during the execution phase. Specifically, this includes determining the absolute position of the object identifier on the physical screen (screen coordinates) and mapping it to a relative position within the execution environment (viewport coordinates) using a transformation algorithm. Then, actions are driven based on these viewport coordinates. This effectively achieves operational consistency of automated scripts across resolutions and devices, eliminating deviations caused by different resolutions or display ratios, and ensuring that input commands accurately fall within the effective response area of the target control regardless of window movement or scaling.
[0077] According to some embodiments, the target operation includes an input operation, wherein the input operation is implemented based on at least one of the following operations: mouse operation, keyboard operation, clipboard operation, and input method switching operation.
[0078] In the above embodiments, the target operation specifically encompasses an input operation, which is achieved through the coordinated use of at least one of mouse operations, keyboard operations, clipboard operations, and input method switching operations. This composite input control strategy can significantly improve the success rate and accuracy of automated text entry.
[0079] In some embodiments, before performing an input operation, it can be determined whether the focus is in the target control (such as an input box) corresponding to the input action. In actual automated processes, simply simulating keyboard tapping is often not robust enough. If the target input box happens to lose system focus when the input command is issued (e.g., due to a sudden system pop-up), the input content may be lost or incorrectly entered into another location. This embodiment ensures that every input is accurately applied to the intended text area by forcibly verifying the activation state of the target control before performing the input operation. If it is detected that the target control has not gained focus, the click activation logic is automatically triggered.
[0080] Specifically, the input operations refer to the interactive behavior of the system submitting data or instructions to the target application. "Mouse operations" can include not only simple clicks for confirmation, but also simulated long presses, dragging sliders, or hovering over specific areas to trigger drop-down menus; "keyboard operations" encompass simulated character typing, function key input (such as Enter to submit), and shortcut key combinations (such as Ctrl+S to save); "clipboard operations" are an efficient input method, which involves writing long text to the system clipboard and simulating a paste action (Ctrl+V) to replace character-by-character typing. This significantly improves speed and avoids key conflicts when processing large blocks of text or special symbols; and "input method switching operations" are specific optimizations for multilingual environments. For example, automatically switching the input mode to English before inputting English instructions avoids mis-entering code as pinyin characters due to input method errors, thus ensuring the semantic correctness of the input content.
[0081] According to some embodiments, performing the target operation on the display page to perform the target task includes: after performing the target operation, taking a screenshot of the display page of the target application to determine whether the display page has undergone a page change that matches the target operation; and re-performing the target operation in response to determining that no page change has occurred that matches the target operation.
[0082] Through the above embodiments, a closed-loop feedback control mechanism is introduced into the execution logic. This means that after executing the corresponding action, a screenshot of the target application's display page is immediately taken to determine if the page has undergone the expected changes. If no changes are found, the logic is re-executed. This "post-execution verification and automatic retry" technique effectively solves the common problems of "false execution" or lost instructions in automation, significantly improving the reliability of task execution. In complex software interactions, due to network latency, system lag, or delayed interface rendering, a click or input command may fail to elicit a successful response from the application. By implementing a "post-click verification" mechanism, it is ensured that every critical step is truly effective, rather than merely completing the instruction transmission.
[0083] During this process, "whether the display page undergoes a page change matching the target operation" refers to the visual feedback basis by which the system logically infers whether the operation has been successfully received and processed by the target application by comparing the image features before and after the target operation is executed. For example, for the operation of "clicking a checkbox", the matching page change means that the image feature of the control area changes from "unselected state (empty box)" to "selected state (checkmark)"; for the operation of "clicking the page turn or submit button", the matching page change is usually manifested as a refresh of the overall layout of the current page, the disappearance of the pop-up window, or a jump to a completely new functional interface; and for the "text input" operation, the page change means that the expected character pixels appear in the target input box area, rather than remaining blank or still displaying the default prompt placeholder. Only after recognizing these specific visual feedbacks will the system confirm the completion of the current step and proceed to the next stage; otherwise, it will initiate a retry process.
[0084] In some embodiments, when the preset template library includes multiple templates for adapting to different page display states of the target application, if it is determined that no page change matching the target operation has occurred (i.e., the target operation fails to execute), the applicable template can be automatically switched according to the current resolution, theme mode, or language version, so as to re-implement the target operation based on the switched template, thereby ensuring that the recognition process adapts to the visual differences brought about by UI changes or application upgrades.
[0085] According to some embodiments, re-executing the target operation in response to determining that no page change matching the target operation has occurred includes: re-executing the target operation in a manner different from that used in the previous execution of the target operation in response to determining that no page change matching the target operation has occurred.
[0086] In the above embodiments, the anomaly recovery strategy is further refined. Specifically, when the system determines that the previous operation failed through page change detection, it no longer simply repeats the same operation instruction, but re-executes it in a different way than the method used to execute the target operation last time, thereby greatly improving the self-repair capability in abnormal scenarios.
[0087] Taking input operations as an example: Suppose we first try to use keyboard input, that is, to write data to the target input box by simulating the character keystrokes of a physical keyboard. However, due to the target application having anti-injection protection enabled or due to high CPU load causing key events to be lost, the system finds that the input box is still empty (i.e., no matching page change has occurred) during the verification process. At this time, when a retry is triggered, it can automatically switch to clipboard pasting. Since the system interface called by clipboard operation is completely different from that of simulated keystrokes, this method can often successfully bypass the interception or interference of keyboard signals, thus successfully completing the data entry.
[0088] According to some embodiments, the method is applied to a first client. The method further includes: receiving a trigger request sent by a target agent, the trigger request being generated and sent to the first client by the target agent when it detects trigger information for the target task during a call with a second client.
[0089] According to some embodiments, the method further includes: during the execution of the target operation on the display page, obtaining execution status information corresponding to the target operation, wherein the execution status information is used by the target agent to analyze and generate corresponding feedback information, wherein the feedback information is used to provide feedback on the call with the second client.
[0090] In the above embodiments, during the execution of the target operation, the corresponding execution status information can also be obtained in real time and sent back to the target intelligent agent for analysis, thereby generating feedback information for the second client (user end).
[0091] In some embodiments, the execution status information of the target operation may include, but is not limited to, operation completion status, image recognition success rate, task execution time, operation failure type, task execution status (e.g., task execution can be driven by a state machine: managing task status through a state machine, including in progress, completed, execution failed, etc., to achieve controllability and traceability of the task flow), screenshot logs (e.g., logging through screenshots: the system automatically takes a screenshot and records the operation status each time an operation is executed or an exception occurs, facilitating task backtracking and problem localization), error codes, etc. For example, the system can automatically take a screenshot and record the operation status each time an operation is executed or an exception occurs, facilitating task backtracking and problem localization.
[0092] Therefore, this "state awareness and real-time feedback" technical mechanism significantly improves the controllability, traceability, and transparency of business transformation of automated tasks. As mentioned above, in some embodiments, by introducing a state machine-driven task management mechanism, the flow logic of a task from "in execution" to "completed" or "failed" is precisely defined, achieving refined control over the entire lifecycle of the task.
[0093] In this process, the execution status information is a set of technical metadata, which, as mentioned above, can cover things like operation completion status, image recognition confidence and success rate, task execution time, specific operation failure types, and screenshots and logs of key nodes. Feedback information, on the other hand, is a natural language response output by the target agent to the user (second client) after semantic processing based on the aforementioned technical data. For example, when the execution status information displays "Friend addition action completed" and "Execution status is Successful," the feedback information generated by the agent might be a voice or text reply: "The contact has been successfully added for you; you can start the conversation directly." Conversely, if the execution status information displays "Search button recognition failed (error code 404)," the feedback information generated by the agent might be: "The current system interface is slow; I am trying to search again; please wait." This mechanism ensures that the user can perceive the execution progress on the desktop (first client) in real time and accurately during the call, eliminating information asymmetry in human-computer collaboration.
[0094] According to some embodiments, the target application accesses the system through the following operations to achieve the target task: in response to receiving a plugin registration request for the target application, the application obtains the registration information of the target application, wherein the registration information includes the preset template and control logic for achieving the target task; and based on the registration information, the application loads the plugin corresponding to the target application to control the target application to achieve the target task through a preset interface.
[0095] In some embodiments, control logic can refer to a set of rules and instructions encapsulated within a plugin for orchestrating and driving specific business processes. Control logic can be a set of internal rules that guide how to execute these actions. By encapsulating these complex business rules as "control logic" within a plugin, the automation system does not need to understand how the specific application operates. It only needs to send signals to the plugin through a preset interface, and the plugin will autonomously direct the UI automation process according to its internal control logic, thereby achieving a high degree of encapsulation and decoupling of business details.
[0096] Through the above embodiments, a highly flexible "plug-in" application access architecture is constructed. Specifically, this method allows external target applications, or a target task within a target application, to dynamically access the automation system as a plug-in through a standardized registration process. In this process, the system responds to the registration request and obtains registration information, including a preset template, then loads the corresponding plug-in, and finally uses a standardized preset interface to control the target application.
[0097] This plug-in-based access and control mechanism achieves deep decoupling between the automation system and specific business applications. This embodiment utilizes a "plug-in pluggable mechanism," making the integration of new applications extremely fast—simply develop and register the corresponding plug-in, without modifying the core engine code. This not only supports independent release, upgrade, and rollback of plug-ins, greatly improving system maintainability, but also shields the underlying operating system or client implementation differences (such as differences between Windows and Mac) through the plug-in layer. This allows business developers to focus solely on the business logic itself, without dealing with complex visual recognition or underlying compatibility issues, thus significantly reducing the development difficulty of automated processes.
[0098] Figure 3 A schematic diagram of a system architecture for implementing a method for performing a task according to an embodiment of the present disclosure is shown. Figure 3 As shown, the system adaptation layer is used to implement cross-platform system control and capability abstraction. As mentioned above, the system adaptation layer designs corresponding scripts for different platforms to realize a script library. The core layer can provide visual recognition, target operation execution, task scheduling, and high-stability execution capabilities. For example, in addition to the template matching, target operation implementation, image processing, and recognition actions mentioned above, it can also break down automated tasks into a series of target operations, execute them in a predefined order, and support parallel or asynchronous operations. For target operations that have not responded for a long time, timeout processing can be triggered to avoid task blocking and improve the overall process stability. The plug-in layer is used to abstract the business operations (i.e., tasks) in specific applications into standardized automated action flows (such as maintaining visual templates related to business operations), while decoupling business logic from the core capabilities of the automation system. The plug-in layer is pluggable, allowing different applications to flexibly access the automation system and supporting independent upgrades and expansions. The application layer serves as the desktop runtime environment, completing functions such as task reception, task scheduling, and execution result reporting. It not only provides a user-visible desktop interface but also handles task management and execution scheduling for the local automation system, thereby enabling automated execution and real-time feedback of business actions. The framework layer provides the underlying code framework for implementing the above logic and operations.
[0099] According to embodiments of this disclosure, such as Figure 4As shown, an apparatus 400 for performing a task is also provided, comprising: a response unit 410 configured to determine a target application corresponding to the target task in response to receiving a trigger request for a target task; a determination unit 420 configured to determine a preset template corresponding to the target task, wherein the template includes a first image of a first object identifier for performing the target task, wherein the first object identifier corresponds to a target operation to be performed in the target application; an identification unit 430 configured to identify a second object identifier corresponding to the target operation on a display page of the target application based on the first image through image recognition; and an execution unit 440 configured to execute the target operation on the display page based on the second object identifier to perform the target task.
[0100] Here, the operation of each of the above-mentioned units 410 to 440 of the device 400 for performing the task is similar to the operation of steps 210 to 240 described above, and will not be repeated here.
[0101] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0102] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0103] refer to Figure 5 The present invention describes a structural block diagram of an electronic device 500 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0104] like Figure 5As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface X05 is also connected to the bus 504.
[0105] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0106] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).
[0107] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0110] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0111] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0112] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0113] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0114] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A method for performing a task, comprising: In response to receiving a trigger request for a target task, the target application corresponding to the target task is determined; A preset template corresponding to the target task is determined, wherein the template includes a first image of a first object identifier for performing the target task, wherein the first object identifier corresponds to the target operation to be performed in the target application; Based on the first image, a second object identifier corresponding to the target operation is identified on the display page of the target application through image recognition; and The target task is performed by executing the target operation on the display page based on the second object identifier.
2. The method as described in claim 1, wherein, The method further includes: In response to receiving the trigger request, the system of the device where the first client is located is determined; Determine the script corresponding to the system from the preset script library; and The target task is executed based on the script, wherein the script library includes scripts corresponding to multiple systems.
3. The method as described in claim 1, wherein, After responding to receiving a trigger request for a target task and determining the target application corresponding to the target task, the method further includes: Determine and adjust the running state of the target application to bring it into a ready state capable of accomplishing the target task. The running state includes at least one of the following: the target application startup state, the target application account login state, the target application window activation state, and the target application page display state.
4. The method of claim 1, wherein, The method is applied to a first client, and the method further includes at least one of the following: The system receives a trigger request sent by the target agent, which is generated and sent to the first client by the target agent when it detects trigger information for the target task during a call with the second client. Receive the trigger request generated based on a preset event detected that is associated with the target task; as well as In response to detecting an operation targeting the second client to trigger the target task, the trigger request is generated.
5. The method as described in claim 1 or 3, wherein, The step of determining the preset template corresponding to the target task includes: Based on the page display state of the target application, a preset template corresponding to the target task is determined from the preset template library corresponding to the target application. The preset template library includes multiple templates for adapting to different page display states of the target application. The page display state includes at least one of the following: page resolution, page scaling ratio, theme of the target application, language of the target application, and version of the target application.
6. The method of claim 1, wherein, The step of identifying the second object identifier corresponding to the target operation based on the first image includes: By performing a screenshot operation, a second image including the display page of the target application is obtained; Perform preprocessing operations on the second image; and A similarity calculation is performed between the preprocessed second image and the first image to identify the second object identifier corresponding to the target operation on the display page of the target application based on the similarity. The preprocessing operations include at least one of the following: grayscale conversion, matrix conversion, binarization, edge detection, and contrast enhancement.
7. The method as claimed in claim 1 or 6, wherein, The step of identifying the second object identifier corresponding to the target operation on the display page of the target application based on the first image through image recognition includes: By performing a screenshot operation, a second image including the display page of the target application is obtained; The second image and the first image are compared to calculate their similarity, so as to identify a second object identifier corresponding to a first object identifier in the first image on the display page of the target application based on the similarity; and In response to the failure to identify a second object identifier corresponding to the corresponding first object identifier, optical character recognition is performed on the first image and the second image of the corresponding first object identifier to determine the second object identifier corresponding to the corresponding first object identifier based on the identified characters.
8. The method of claim 1, wherein, Performing the target operation on the display page to execute the target task includes: Determine the screen coordinates corresponding to the identified second object identifier; Through coordinate transformation, the screen coordinates corresponding to the identified second object identifier are converted into the viewport coordinates corresponding to the displayed page; and Based on the viewport coordinates, perform the target operation for the second object identifier.
9. The method of claim 1, wherein, The target operation includes an input operation, wherein the input operation is implemented based on at least one of the following operations: mouse operation, keyboard operation, clipboard operation, and input method switching operation.
10. The method of claim 1 or 9, wherein, Performing the target operation on the display page to execute the target task includes: After performing the target operation, a screenshot of the target application's display page is taken to determine whether the display page undergoes a page change that matches the target operation; and In response to determining that no page change matching the target operation has occurred, the target operation is re-executed.
11. The method of claim 10, wherein, In response to determining that no page change matching the target operation has occurred, re-executing the target operation includes: In response to determining that no page change matching the target operation has occurred, the target operation is re-executed in a manner different from that used in the previous execution of the target operation.
12. The method of claim 1, wherein, The method is applied to a first client, and the method further includes: The system receives a trigger request sent by the target agent, the trigger request being generated and sent to the first client by the target agent when it detects trigger information for the target task during a call with the second client; and During the execution of the target operation on the display page, the execution status information corresponding to the target operation is obtained. The execution status information is used by the target intelligent agent to analyze and generate corresponding feedback information, wherein the feedback information is used to provide feedback on the call with the second client.
13. The method of claim 1, wherein, The target application accesses the system through the following operations to achieve the target task: In response to receiving a plugin registration request for the target application, the system obtains the registration information of the target application, wherein the registration information includes the preset template and control logic for implementing the target task; and Based on the registration information, the plugin corresponding to the target application is loaded to control the target application to perform the target task through a preset interface.
14. An apparatus for performing a task, comprising: The response unit is configured to determine the target application corresponding to the target task in response to receiving a trigger request for the target task; The determining unit is configured to determine a preset template corresponding to the target task, wherein the template is used to execute a first image of a first object identifier of the target task, wherein the first object identifier corresponds to a target operation to be executed in the target application; The identification unit is configured to, based on the first image, identify a second object identifier corresponding to the target operation on the display page of the target application through image recognition; and An execution unit is configured to perform the target operation on the display page based on the identified second object identifier, in order to perform the target task.
15. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.
17. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-13.