Page operation method and related device

By obtaining user input information, and automatically performing page operations using the intent identification model and tool set, the problem of low operation efficiency of users in web pages and applications is solved, and user experience and operation efficiency is improved.

CN120295528APending Publication Date: 2025-07-11HENAN QINWEI DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213528.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When users use web pages and applications, frequent page navigation, page jumps and form filling operations are inefficient and have poor user experience.

Method used

By obtaining user input information, using the intent identification model to identify the user's operation intention, and calling the corresponding tool set (such as page element click tool, page jump tool, page filling tool) to automatically perform operations, combining data index and text generation model to ensure the accuracy and efficiency of operations.

Benefits of technology

It realizes automatic completion of page operations, improves user experience and operation efficiency, reduces the user's manual operation needs, ensures the accuracy of operation sequence and intelligent scheduling of tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295528A_ABST
    Figure CN120295528A_ABST
Patent Text Reader

Abstract

The invention discloses a page operation method and a related device, and the method comprises the steps: obtaining the input information of a user, the input information comprises the information that the user carries out the target operation on a target page; the input information is input into an intention recognition model to output a user intention, and the user intention represents the category of the target operation; according to the user intention, a target tool is obtained from the multiple tools through matching, and the target tool represents a function module for executing the target operation; and calling the target tool to enable the target tool to execute the target operation. In the application, after the input information of the user is acquired, the target operation indicated by the input information can be automatically completed without manual operation of the user, so that the user experience and the operation efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a page operation method and related device. Background Art

[0002] With the popularization of the Internet and the wide use of mobile devices, software platforms such as web pages (Web) and applications (Application, App) have become an indispensable part of people's daily lives. These software platforms provide rich functions and content, but users may need to frequently perform page operations such as page navigation, page jumping, form filling, etc. when using them, resulting in low work efficiency and poor user experience. Therefore, a page operation method is needed to improve user experience and operation efficiency. Summary of the Invention

[0003] An embodiment of this application provides a page operation method and related device, which can automatically complete page operations and improve user experience and operation efficiency.

[0004] In a first aspect, an embodiment of this application provides a page operation method, including: obtaining input information of a user, where the input information includes information about a target operation on a target page by the user; inputting the input information into an intent recognition model to output a user intent, where the user intent represents the category of the target operation; and calling a target tool to cause the target tool to perform the target operation.

[0005] In an embodiment of this application, after obtaining the input information of the user (such as voice, text, etc.), the target operation indicated by the input information can be automatically completed without manual operation by the user, improving user experience and operation efficiency. Inputting the input information into an intent recognition model to output a user intent, and the intent recognition model ensures the accuracy of user intent recognition. According to the user intent, a target tool is matched from multiple tools, realizing intelligent scheduling and load balancing of multiple tools. Calling the target tool to cause the target tool to perform the target operation realizes automatic execution of the target operation without manual operation by the user, improving user experience and operation efficiency.

[0006] In a possible implementation manner, the multiple tools include at least one of a page element click tool, a page jump tool, and a page filling tool; the page element click tool represents a module for performing a page element click operation, the page jump tool represents a module for performing a page jump operation, and the page filling tool represents a tool for performing a page information filling operation.

[0007] In this implementation, the multiple tools include at least one of a page element click tool, a page jump tool, and a page filling tool. The page element click tool, the page jump tool, and the page filling tool constitute an automated operation tool set, enabling the server to simulate user interactions with the page and perform operations such as automated testing, automated operations, and automated data extraction. Taking "filling in information on a certain page in a browser" as an example, it is necessary to first call the page element click tool to open the page elements of the browser; then call the page jump tool to perform page jumps in the browser and reach the specified page; and then call the page filling tool to fill in information on the specified page.

[0008] In a possible implementation, the above-mentioned calling of the target tool includes: retrieving the data set according to the input information to obtain matching data, where the matching data includes parameter information related to performing the target operation; inputting the input information and the matching data into a text generation model to generate text, where the text includes target parameter information, and the target parameter information is obtained by the text generation model processing the parameter information in the matching data; and calling the target tool according to the target parameter information.

[0009] In this implementation, according to the input information, the data set is retrieved to obtain matching data. The data set includes multiple document data, and each document data includes parameter information for navigating to a certain page and performing operations on that page. The matching data includes at least one document data related to the input information. The input information and the matching data are input into a text generation model to generate text. The text generation model processes the parameter information in the matching data to obtain the target parameter information, and finally forms the text. The processing of the matching data by the text generation model includes text understanding, data screening, etc., thus ensuring the accuracy of the target parameter information. According to the target parameter information, the target tool is called to enable the target tool to perform the target operation, realizing the automated execution of the target operation without the need for manual user operation, improving the user experience and operation efficiency.

[0010] In a possible implementation, there are multiple target tools, and the target parameter information includes multiple parameter sub-informations. Each parameter sub-information represents the parameter information of one target tool among the multiple target tools. The above-mentioned calling of the target tool according to the target parameter information includes: determining the calling order of the multiple target tools according to the order of the multiple parameter sub-informations in the target parameter information; and sequentially inputting the multiple parameter sub-informations into the multiple target tools according to the calling order, so that the multiple target tools are sequentially called according to the calling order.

[0011] In this implementation, during the process of automating page operations, the order of page operations is crucial for successfully completing page operations. Taking "filling in information on a certain page in a browser" as an example, it is necessary to first call the page element click tool to open the page elements of the browser; then call the page jump tool to perform page jumps in the browser to reach the specified page; and then call the page filling tool to fill in the information on the specified page. According to the order of multiple parameter sub-information in the target parameter information, the calling order of multiple target tools is determined. The calling order can ensure the accuracy of the page operation order during the process of automating page operations. According to the calling order, multiple parameter sub-information are sequentially input into multiple target tools, so that the multiple target tools are sequentially called according to the calling order, thereby sequentially implementing multiple page operations and completing the page operation task indicated by the input information.

[0012] In a possible implementation, the above-mentioned retrieving the matching data from the data set according to the input information includes: using the data set and the input information as the input of the data indexing model, so that the data indexing model matches the data of the input information and the data set to output the matching data.

[0013] In this implementation, the data indexing model is a neural network model that supports semantic representation and retrieval tasks in more than multiple languages, ensuring the accuracy and precision of the matching data.

[0014] In a possible implementation, the data in the data set is vectorized data. The above-mentioned retrieving the matching data from the data set according to the input information includes: performing vectorization processing on the input information to obtain vectorized input information; and matching the vectorized data in the data set with the vectorized input information to obtain the matching data.

[0015] In this implementation, performing vectorization processing on the data in the data set and the input information improves the efficiency of data matching for the data set.

[0016] In a possible implementation, the data in the data set includes at least one of page path data, page element data, and form element data.

[0017] In this implementation, the page path data is an example of parameter information for page navigation, and the page element data and form element data are examples of parameter information for page operations. The examples of the data in the data set in this implementation are only for illustrative purposes and do not constitute a limitation on the data in the data set.

[0018] In a possible implementation, the method of collecting data in the data set includes at least one of collecting data through buried points, docking the application programming interface of the business database, manual entry, or collection through scripts.

[0019] In this implementation, data collection is carried out by means of data collection through logging, docking with the application programming interface of the business database, manual entry, or collection through scripts, ensuring the integrity and richness of the data in the data set and the accuracy of the target parameter information.

[0020] In one possible implementation, the above-mentioned inputting the input information and the matching data into the text generation model to generate text further includes: generating a prompt word according to the input information and the matching data; inputting the prompt word into the text generation model to generate text.

[0021] In this implementation, the prompt engineering of the text generation model is utilized, which can provide more accurate and richer prompt content for the text generation model, ensuring the accuracy of the target parameter information generated by the text generation model.

[0022] In one possible implementation, the above-mentioned matching the target tool from multiple tools according to the user intention includes: comparing the similarity between the user intention and the function information of each tool in the multiple tools, and determining the tool with a similarity exceeding the preset threshold in the multiple tools as the target tool.

[0023] In this implementation, by comparing the similarity, the tool with a similarity exceeding the preset threshold is determined as the target tool, improving the accuracy of the target tool.

[0024] In one possible implementation, multiple tools are obtained through user definition and registration.

[0025] In this implementation, multiple tools are obtained through user definition and registration, which can meet the personalized needs of different users.

[0026] In one possible implementation, the method further includes: when the data type of the input information is voice, inputting the voice data into the speech recognition model to convert the input information into text form.

[0027] In this implementation, the speech recognition model is a neural network model, which is used for multi-speech recognition, speech-to-text conversion, speech translation, speaker recognition, etc. The speech recognition model can convert the input information in voice form into text form, ensuring the accuracy of the input information.

[0028] In one possible implementation, the above-mentioned input information is obtained through multiple rounds of conversations and / or historical conversations between the user and the agent.

[0029] In this implementation, the input information of the user is determined based on multiple rounds of conversations, historical conversations, etc., which ensures the accuracy of the input information. Consequently, it guarantees that the intent recognition model can identify a more precise user intent, thereby reducing the possibility of incorrect matching of the target tool.

[0030] In one possible implementation, the method further includes: when the user intent does not meet the preset requirements, inputting the input information and the user intent into a text generation model to output a reply text, where the reply text is used to generate multiple rounds of conversations between the user and the intelligent agent. The above-mentioned inputting the input information into the intent recognition model to output the user intent includes: inputting the data of multiple rounds of conversations between the user and the intelligent agent into the intent recognition model to output the user intent.

[0031] In this implementation, when the user intent is inaccurate or has a low precision, a multiple-round conversation with the user can be constructed by combining with the text generation model. The multiple-round conversation ensures the accuracy of the input information, which in turn guarantees that the intent recognition model can identify a more precise user intent, thereby reducing the possibility of incorrect matching of the target tool.

[0032] In one possible implementation, the method further includes: obtaining the execution result of the target tool, where the execution result represents the result of the target tool performing the target operation; inputting the execution result into the text generation model to generate output information; and outputting the above output information to the user.

[0033] In this implementation, by providing the user with feedback on the execution result of the target tool, it is convenient for the user to understand the result of the target operation on the target page, thereby improving the user experience.

[0034] In a second aspect, an embodiment of the present application provides a page operation device, including: an acquisition module, configured to acquire the input information of the user, where the input information represents the information of the user performing a target operation on a target page; a processing module, configured to input the input information into an intent recognition model to output a user intent, where the user intent represents the category of the target operation; and, according to the user intent, matching a target tool from multiple tools, where the target tool represents a functional module for performing the target operation; and, invoking the target tool to enable the target tool to perform the target operation.

[0035] In one possible implementation, the multiple tools include at least one of a page element click tool, a page jump tool, and a page filling tool; the page element click tool represents a module for performing a page element click operation, the page jump tool represents a module for performing a page jump operation, and the page filling tool represents a tool for performing a page filling operation.

[0036] In a possible implementation, the above processing module is specifically configured to: retrieve a data set according to input information to obtain matching data, where the matching data includes parameter information related to performing a target operation; input the input information and the matching data into a text generation model to generate text, where the text includes target parameter information, and the target parameter information is obtained by the text generation model processing the parameter information in the matching data; and call a target tool according to the target parameter information.

[0037] In a possible implementation, there are multiple target tools, and the target parameter information includes multiple parameter sub-information, and each parameter sub-information represents the parameter information of one of the multiple target tools. The above processing module is specifically configured to: determine the call order of the multiple target tools according to the order of the multiple parameter sub-information in the target parameter information; and sequentially input the multiple parameter sub-information into the multiple target tools according to the call order, so that the multiple target tools are sequentially called according to the call order.

[0038] In a possible implementation, the above processing module is specifically configured to: use the data set and the input information as the input of a data indexing model, so that the data indexing model matches the data of the input information and the data set to output the matching data.

[0039] In a possible implementation, the data in the data set is vectorized data. The above processing module is specifically configured to: perform vectorization processing on the input information to obtain vectorized input information; and match the vectorized data in the data set with the vectorized input information to obtain the matching data.

[0040] In a possible implementation, the data in the data set includes at least one of page path data, page element data, and form element data.

[0041] In a possible implementation, the method for collecting data in the data set includes at least one of collecting data by logging, docking with the application programming interface of the business database, manual entry, or collection through a script.

[0042] In a possible implementation, the above processing module is specifically configured to: generate a prompt word according to the input information and the matching data; and input the prompt word into the text generation model to generate text.

[0043] In a possible implementation, the above processing module is specifically configured to: compare the similarity between the user intention and the function information of each of the multiple tools, and determine the tools with similarity exceeding a preset threshold among the multiple tools as the target tools.

[0044] In a possible implementation, the multiple tools are obtained through user definition and registration.

[0045] In a possible implementation, the above processing module is further configured to: when the data type of the input information is voice, input the voice data into a speech recognition model so that the input information is converted into text form.

[0046] In a possible implementation, the above input information is obtained through multiple rounds of conversations and / or historical conversations between the user and the intelligent agent.

[0047] In a possible implementation, the above processing module is further configured to: when the user intention does not meet the preset requirements, input the input information and the user intention into a text generation model to output a response text, and the response text is used to generate multiple rounds of conversations between the user and the intelligent agent. Specifically, the processing module is configured to: input the multiple rounds of conversation data between the user and the intelligent agent into an intention recognition model to output the user intention.

[0048] In a possible implementation, the above processing module is further configured to: obtain the execution result of the target tool, where the execution result represents the result of the target tool executing the target operation; input the execution result into a text generation model to generate output information; and output the above output information to the user.

[0049] In a third aspect, an embodiment of the present application provides an intelligent agent, which is deployed with an intention recognition model and a text generation model, and the intelligent agent is used to execute the method described in the first aspect and any of its possible implementations above.

[0050] In a fourth aspect, an embodiment of the present application provides a server, and the server is used to execute the method described in the first aspect and any of its possible implementations above.

[0051] In a fifth aspect, an embodiment of the present application provides a page operating system, including: a terminal device, the server in the fourth aspect above, and the terminal device and the server are communicatively connected.

[0052] In a sixth aspect, an embodiment of the present application provides a chip system, which includes a processor and a power supply circuit. The power supply circuit is used to supply power to the processor, and the processor is used to execute the method described in the first aspect and any of its possible implementations above.

[0053] In a seventh aspect, an embodiment of the present application provides a computing device, which includes a processor and a memory. The processor is used to execute the instructions stored in the memory so that the computing device executes the method described in the first aspect and any of the above possible implementations.

[0054] In an eighth aspect, an embodiment of the present application provides a computing device cluster, including at least one computing device, and each computing device includes a processor and a memory; the processors of at least one computing device are configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method described in the first aspect and any possible implementation manner thereof as described above.

[0055] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium, including computer program instructions, which, when run on a computing device cluster, cause the computing device cluster to execute the method described in the first aspect and any possible implementation manner thereof as described above, where the computing device cluster includes at least one computing device.

[0056] In a tenth aspect, an embodiment of the present application provides a computer program product containing instructions, which, when run on a computing device cluster, cause the computing device cluster to execute the method described in the first aspect and any possible implementation manner thereof as described above, where the computing device cluster includes at least one computing device.

[0057] It can be understood that for the beneficial effects of the above second aspect to tenth aspect, reference may be made to the relevant descriptions in the first aspect and any possible implementation manner thereof as described above, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The following briefly introduces the drawings required for description in the embodiments or the prior art.

[0059] Figure 1 It is a schematic diagram of the composition of a page operating system provided in an embodiment of the present application;

[0060] Figure 2 It is a flowchart of a page operation method provided in an embodiment of the present application;

[0061] Figure 3 It is a flowchart of a page operation method executed by a page operating system provided in an embodiment of the present application;

[0062] Figure 4 It is a schematic diagram of the composition of an example of a page operating system provided in an embodiment of the present application;

[0063] Figure 5 It is a schematic diagram of an example of a page provided in an embodiment of the present application;

[0064] Figure 6 It is a schematic diagram of the structure of a page operation device provided in an embodiment of the present application;

[0065] Figure 7 It is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0066] Figure 8 It is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application;

[0067] Figure 9 It is a schematic structural diagram of another computing device cluster provided by an embodiment of the present application. Detailed implementation manners

[0068] In this article, the term "and / or" is a relational description of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In this article, the symbol " / " represents that the associated objects are in an "or" relationship. For example, A / B represents A or B.

[0069] The terms "first", "second", etc. in the description and claims of this application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.

[0070] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0071] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. Exemplarily, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.

[0072] To facilitate the understanding of the solution provided by the embodiments of the present application, some terms related to this solution are briefly introduced first.

[0073] Artificial Intelligence (AI): A technology that uses computer technology and algorithms to simulate human thinking and behavior by analyzing and processing a large amount of data. It enables a computer to perform tasks that usually require human intelligence, such as learning, reasoning, perception, understanding, and creation.

[0074] Python: A high-level programming language, renowned for its concise syntax, powerful library support, and wide range of application domains. Python has a very wide range of applications in the field of artificial intelligence. From data processing, feature extraction to model training and evaluation, Python provides rich tools and libraries to support the implementation of these tasks.

[0075] Knowledge Base (KB): A database system that stores and manages professional knowledge, rules, facts, and concepts in a specific domain. In the field of artificial intelligence, the knowledge base provides the knowledge support required for intelligent systems to reason and make decisions. Python can build and manage knowledge bases through database management systems and specific libraries, supporting the knowledge acquisition and application of intelligent systems.

[0076] Retrieval-augmented Generation (RAG): A technique that combines information retrieval and text generation. In this method, the model not only relies on its own knowledge and training data to generate text but also retrieves relevant information from an external knowledge base to assist in the generation process. This technique can improve the relevance and accuracy of the generated text, especially in scenarios where specific facts, data, or knowledge need to be cited.

[0077] Voice Recognition: Voice recognition technology refers to the process of converting human speech signals into machine-readable text information. It uses technologies such as audio signal processing and natural language processing to analyze and understand human spoken content, thus realizing the conversion from speech to text.

[0078] Agent: An autonomous entity that can perceive the environment, make decisions, and execute actions. In the field of artificial intelligence, agents learn how to make optimal decisions by interacting with the environment to achieve specific goals. As a programming language, Python provides rich libraries and tools for the development and implementation of agents, especially in fields such as machine learning and deep learning, and these technologies provide strong support for the intelligent decision-making and behavior of agents.

[0079] Artificial Intelligence Generated Content (AIGC): Information, text, images, audio, or video content automatically generated using artificial intelligence technology. This technology can learn from large amounts of data and generate new content with a specific style or theme according to preset rules, algorithms, or models. In Python, the function of generating artificial intelligence content can be achieved by calling relevant machine learning and deep learning libraries or application programming interfaces (APIs).

[0080] Web Automation Tool (Selenium): An open-source tool for automated testing of web applications. It can simulate user operations in a browser, such as clicking and inputting, to detect the functionality and performance of the application. Selenium supports multiple browsers, operating systems, and programming languages and is one of the important tools in the field of web automation testing.

[0081] Large Language Model (LLM): Generally refers to a neural network model with an extremely large number of parameters (usually over one billion). Large language models demonstrate excellent capabilities in various natural language processing tasks, such as text classification, sentiment analysis, summary generation, translation, etc.

[0082] BERT (Bidirectional Encoder Representation from Transformers) Model: A pre-trained model for language representation. The BERT model adopts certain training strategies to generate a deep bidirectional language representation model.

[0083] An embodiment of this application provides a page operation method. In this method, input information is obtained, and the input information includes information related to the user's target operation on the target page; an Artificial Intelligence (AI) model is called to perform intent recognition and parameter extraction on the user's input information; according to the intent recognition result, a target tool is matched from multiple tools; the target tool is called according to the extracted target parameter information, and the target tool can automatically perform the target operation.

[0084] In an embodiment of this application, the target operation indicated by the user's input information can be any operation performed by the user on the page, including but not limited to automatic page jump, automatic form filling, generating suggestions related to user operations, generating explanations of certain functions of the page, generating predictions of possible user operations, chatting and interacting with the user through an agent, etc.

[0085] In an embodiment of this application, after obtaining the user's input information (such as voice, text, etc.), the target operation indicated by the input information can be automatically completed without the need for the user to manually operate, improving the user experience and operation efficiency.

[0086] Figure 1 It is a schematic diagram of the composition of a page operation system provided in an embodiment of this application. As Figure 1As shown in the figure, an embodiment of the present application provides a page operating system 100 (hereinafter referred to as "system 100"). The system 100 can be an intelligent agent, which mainly includes: a terminal device 110 and a server 120. Among them, the terminal device 110 deploys a page of the intelligent agent, such as an intelligent assistant, which is used to obtain user input information, which can be voice, text, etc. input by the user. The server 120 deploys an AI model, such as an intent recognition model and a text generation model. The server 120 also deploys a dataset (also known as a knowledge base) and a tool set, and the tool set includes one or more tools. In the embodiment of the present application, a tool refers to a functional module for automating page operations, which can be an Application Programming Interface (API), or a python class and function. The tool can receive input information and return the execution result after the tool is executed.

[0087] Further, the intent recognition model and the text generation model can be deployed in the same server or in different servers.

[0088] Further, the server 120 can include one or more servers ( Figure 1 illustrated by taking one server as an example), and the server 120 can provide the methods and / or devices provided in the embodiments of the present application for one or more terminal devices.

[0089] Further, a software platform related to the methods and / or devices of the present application can be installed on the terminal device 110. The above software platform can provide a page and an intelligent agent, such as an intelligent assistant. The terminal device 110 can receive input information input by the user on the intelligent assistant. The input information includes information about the user's target operation on the target page. The target page is a page in the above software platform, and the above input information is sent to the server 120. The server 120 deploys an AI model, such as an intent recognition model and a text generation model. It can use the received input information as the input of the intent recognition model, output the user intent, and match the target tool from the tool set according to the user intent; retrieve and match the matching data related to the input information from the dataset according to the received input information, use the received input information and the matching data as the input of the text generation model, and output the text including the target parameter information; call the target tool according to the target parameter information, and the target tool automatically executes the target operation, thereby completing the automatic operation of the target page.

[0090] It should be understood that in some alternative implementations, the terminal device 110 is deployed with an AI model, and the terminal device 110 can also perform the action of obtaining the execution result based on the received input information by itself, without the cooperation of the server. The embodiments of the present application do not limit this.

[0091] In the embodiments of the present application, the server 120 is deployed with an intent recognition model and a text generation model, and it can also be understood that the server 120 is configured with or has the calling ability of the intent recognition model and the text generation model.

[0092] In the embodiments of the present application, the terminal device 110 can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiments of the present application do not impose any restrictions on this. Exemplary embodiments of the terminal device 110 involved in this solution include, but are not limited to, electronic devices running iOS, Android, Windows, Harmony OS, or other operating systems. The embodiments of the present application do not specifically limit the type of the electronic device.

[0093] In the embodiments of the present application, the server 120 can be various servers, such as servers with an X86 architecture, specifically, it can be a whole cabinet server, a blade server, a high-density server, a rack server, or a high-performance server, etc. In other words, the embodiments of the present application do not specifically limit the specific category of the server. Further, it can be understood that Figure 1 The structure of the server shown does not constitute a limitation on the structure of the server. The server may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0094] Further, the server 120 can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured as a cloud server or a cloud server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The cloud server cluster is deployed in several cloud data centers; the software can be an application that implements the object control method, etc., but is not limited to the above forms.

[0095] In a possible scenario, the server 120 can serve as the cloud (a software platform that adopts Application Virtualization technology and integrates multiple functions such as software search, download, use, management, backup, etc.). When specifically used, the server 120 can deploy a cloud management platform and a data center, and the terminal device 110 and the cloud interact through the cloud management platform. Additionally, nodes can be deployed in the data center, where the nodes in the data center can be virtual machine instances, container instances, physical servers, etc.

[0096] In another possible scenario, the page operation method provided by the embodiments of the present application can be implemented through software. This software has a terminal and a server side. The terminal device 110 runs the terminal of this software, and the server 120 runs the server side of this software. During the process of the terminal device 110 running the terminal of this software, it can call the server side running on the server 120 to implement the page operation method provided by the embodiments of the present application.

[0097] That is to say, the page operation method provided by the embodiments of the present application can be applied to the above-mentioned terminal device 110 or the server 120. When specifically implemented, it can run on the terminal device 110 or the server 120 in the form of software. For example, this software can be a service or an application program. The embodiments of the present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments of the present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0098] In the embodiments of the present application, the intelligent agent includes intelligent terminals such as robots, in-vehicle devices, or drones. The intelligent agents involved may include various handheld devices, robots, in-vehicle devices, drones, wearable devices, computing devices, or other processing devices. Exemplarily, the intelligent agent may be a Mobile Station (MS), a subscriber unit, a user equipment (UE), a cellular phone, a smartphone, a wireless data card, a Personal Digital Assistant (PDA) computer, a tablet computer, a wireless modem, a handset, a laptop computer, a Machine Type Communication (MTC) terminal, etc.

[0099] The above is the introduction to the system 100 provided by the embodiments of the present application. Next, based on the above content, a page operation method provided by the embodiments of the present application will be introduced. It can be understood that the above method is proposed based on the system 100 described above, and some or all of the content in the above method can refer to the description of the system 100 above.

[0100] Figure 2 It is a schematic flowchart of a page operation method provided by the embodiments of the present application. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities, such as an intelligent agent. As Figure 2 shown, a page operation method mainly includes the following steps:

[0101] Step S210, obtain the input information of the user. The input information includes the information of the user's target operation on the target page.

[0102] In the embodiments of the present application, the input information of the user may be voice, text, etc. The target page represents any page in any software. The target operation represents any operation performed on the target page, and the target operation includes navigating to the target page, clicking on a page element in the target page, filling out a form on the target page, etc.

[0103] Optionally, when the input information of the user is voice, input the voice input by the user into a voice recognition model. The voice recognition model recognizes the voice input by the user and outputs the text data corresponding to the voice input by the user.

[0104] In the embodiments of the present application, the speech recognition model is a neural network model, which is used for multi-speech recognition, speech-to-text conversion, speech translation, speaker recognition, etc. The speech that the speech recognition model can recognize includes, but is not limited to, common mainstream languages and some relatively niche languages.

[0105] Step S220: Input the input information into the intent recognition model to output the user intent. The user intent represents the category of the target operation.

[0106] In the embodiments of the present application, the intent recognition model is a neural network model. The intent recognition model can be a BERT model or any other classification model. The intent recognition model can consider the information before and after a certain word in the text at the same time, so as to more accurately understand the meaning of the word in the context and more precisely assist in understanding the user's intent.

[0107] Optionally, the input information includes data in multi-turn conversations and / or historical conversations.

[0108] Furthermore, the user's input information can be determined according to multi-turn conversations, historical conversations, etc. In this way, a more accurate user intent can be recognized, and thus the possibility of matching errors of the target tool is reduced.

[0109] Exemplarily, multi-turn conversations are described exemplarily. The user's initial input information is, for example, "Delete the components in the shopping cart". Since it is impossible to determine which components to delete, the intent recognition model cannot recognize the accurate user intent based on the initial input information. The components to be deleted by the user can be determined through multi-turn conversations. For example, after multi-turn conversations, the user's input information is, for example, "Delete the first two components in the shopping cart".

[0110] Exemplarily, historical conversations are described exemplarily. The user's initial input information is, for example, "Delete the components in the shopping cart". Since it is impossible to determine which components to delete, the intent recognition model cannot recognize the accurate user intent based on the initial input information. The components to be deleted by the user can be determined through historical conversations. For example, in the historical conversation, when the user's input information is "Delete the components in the shopping cart", it means deleting all components. Then, according to the historical conversation, the user's input information can be determined as "Delete all components in the shopping cart".

[0111] Optionally, when the intent recognition model cannot recognize the accurate user intent based on the input information of the conversations that have occurred, the agent calls the text generation model to enable the text generation model to generate the response text for the next turn of conversation according to the input information of the conversations that have occurred, so that the user can conduct the next turn of conversation according to the response text, thereby achieving the purpose of multi-turn conversations.

[0112] In the embodiments of the present application, the text generation model is a large language model. The text generation model is trained based on a large-scale training dataset and has stronger instruction-following ability and text analysis and understanding ability. The text generation model has the ability to support long texts, powerful multilingual ability, and stronger professional domain knowledge, and thus can better understand and respond to users, generating more accurate and complete response texts.

[0113] Optionally, the agent can call an AI model to accurately identify the user's intention in the form of multiple rounds of conversations. For example, the agent calls a speech recognition model, a text generation model, and an intention recognition model, and cooperates with functions such as multiple conversations and memory of conversations to accurately identify the user's intention. The speech recognition model is used to recognize the speech input by the user and output the text data corresponding to the speech input by the user. The intention recognition model is used to recognize the user's intention. The text generation model is used to generate response texts in multiple rounds of conversations according to the user's input information and the user's intention; the form of the response text includes but is not limited to text, speech, etc.

[0114] Step S230, according to the user's intention, match a target tool from multiple tools. Among them, the target tool represents a functional module for performing a target operation.

[0115] In the embodiments of the present application, multiple tools can be understood as a tool set or a tool collection. The tools in the tool set can be existing tools or tools defined and / or registered. In the agent, the tool is a predefined functional module, and the tool can accept input parameters and return results. The user can register various tools (such as selenium automation operation tool classes, etc.) in the agent and assign a unique identifier to it. The user can also customize tools to meet the personalized needs of different users.

[0116] In the embodiments of the present application, the tools among multiple tools can be API interfaces or python classes and functions, which need to be able to receive input and return results. The tools are used to implement specific functions, such as functions of clicking, jumping, and automatically filling in page elements.

[0117] In the embodiments of the present application, once the user's intention is recognized, the agent will match the user's intention with the registered tools. The matching process may involve further parsing of the user's intention, comparison of the user's intention and tool functions, etc., to ensure the selection of the accurate tool.

[0118] Optionally, a similarity comparison is performed between the user intention and the function information of each of the multiple tools, and the tools with similarity exceeding a preset threshold among the multiple tools are determined as target tools. By performing the similarity comparison and determining the tools with similarity exceeding the preset threshold as target tools, the accuracy of the target tools is improved. For example, the preset threshold is 90%, the user intention is "page jump", the user input information is "jump to the shopping cart page", and the function information of the page jump tool is "page jump". The similarity between the function information of the page jump tool and the input information is 95%, then the tool for "page jump" is determined as the target tool.

[0119] Optionally, the multiple tools include at least one of a page element click tool, a page jump tool, and a page filling tool. The page element click tool represents a module that performs a page element click operation, the page jump tool represents a module that performs a page jump operation, and the page filling tool represents a tool that performs a page filling operation.

[0120] Furthermore, the multiple tools include at least one of a page element click tool, a page jump tool, and a page filling tool. The page element click tool, the page jump tool, and the page filling tool constitute an automated operation tool set, enabling the server to simulate user interaction with the page and perform operations such as automated testing, automated operation, and automated data extraction. Taking the page in a browser as an example, the operations that the page element click tool, the page jump tool, and the page filling tool can control the browser to perform include opening the browser, navigating to a web page, finding and operating elements on the page, performing clicks, inputting text, etc.

[0121] Step S240, call the target tool to enable the target tool to perform the target operation.

[0122] Optionally, the process of calling the target tool includes: retrieving a data set according to the input information to obtain matching data, where the matching data includes parameter information related to performing the target operation; inputting the input information and the matching data into a text generation model to generate text, where the text includes target parameter information, and the target parameter information is obtained by the text generation model processing the parameter information in the matching data; and calling the target tool according to the target parameter information.

[0123] Furthermore, inputting the input information and the matching data into a text generation model to generate text, and the text generation model processes the parameter information in the matching data to obtain target parameter information, and finally forms text; the processing of the matching data by the text generation model includes text understanding, data screening, etc., thus ensuring the accuracy of the target parameter information.

[0124] Exemplarily, taking "filling in information on a certain page in a browser" as an example, it is necessary to first call the page element click tool to open the page elements of the browser; then call the page jump tool to perform page jumps in the browser to reach the specified page; and then call the page filling tool to fill in the information on the specified page.

[0125] Optionally, there can be multiple target tools, and the target parameter information includes multiple parameter sub-information, where each parameter sub-information represents the parameter information of one of the multiple target tools. The process of calling the target tools further includes: determining the calling order of the multiple target tools according to the order of the multiple parameter sub-information in the target parameter information; and inputting the multiple parameter sub-information into the multiple target tools in sequence according to the calling order, so that the multiple target tools are called in sequence according to the calling order.

[0126] Furthermore, in the process of automatically operating the page, the operation order of the page is the key to successfully completing the page operation. Taking "filling in information on a certain page in a browser" as an example, it is necessary to first call the page element click tool to open the page elements of the browser; then call the page jump tool to perform page jumps in the browser to reach the specified page; and then call the page filling tool to fill in the information on the specified page. Determine the calling order of the multiple target tools according to the order of the multiple parameter sub-information in the target parameter information. The calling order can ensure the accuracy of the operation order of the page in the process of automatically operating the page. Input the multiple parameter sub-information into the multiple target tools in sequence according to the calling order, so that the multiple target tools are called in sequence according to the calling order. Thus, multiple page operations are realized in sequence to complete the page operation task indicated by the input information.

[0127] Optionally, the above dataset can be determined by means of data collection. For example, data can be collected through data embedding to obtain user behavior data, common behavior data, etc., to obtain the dataset. Another example is to obtain user data, business data, etc. by docking with the business database API to obtain the dataset. Another example is that page path data, page element data, form element data, etc. can be manually entered or collected through scripts.

[0128] Optionally, in necessary cases, the above dataset can be organized into data in the form of standard questions and answers. For example, the questions in the question-and-answer form are description information, and the answers in the question-and-answer form are parameter information. The description information is used to match and obtain the matching data, and the parameter information is used to obtain the target parameter information in the matching data. In this way, it is more convenient for data retrieval and matching, and it is more convenient to calculate the similarity between the user's input information and the data in the dataset.

[0129] Exemplarily, data embedding refers to a technique of collecting user behavior data by implanting statistical codes at key conversion points of products or services. The data embedding collection methods mainly include the following: Manual data embedding, where developers manually insert collection codes at the places where data needs to be collected. This method is applicable to situations where the amount of data to be collected is small or does not change frequently; Automatic data embedding, which automatically collects data by writing programs and is suitable for situations where a large amount of data needs to be collected and changes frequently; Front-end data embedding, where tracking codes are set on the page to capture user behaviors such as clicks and browsing; Back-end data embedding: Tracking is set on the server side to record back-end behaviors such as API calls and database operations; Visual data embedding: The data embedding positions are set through a visual interface to simplify technical operations; Agentless data embedding, which uses advanced tracking technologies to automatically capture user behaviors without setting each tracking point, also known as full data embedding.

[0130] Optionally, the agent can call a data indexing model to match the data in the dataset with the user's input information. The data indexing model can retrieve data matching the input information from the dataset and output the matching data.

[0131] In the embodiments of this application, the data indexing model is a neural network model. For example, the data indexing model is a semantic vector model (Embedding Model). The data indexing model supports semantic representation and retrieval tasks in more than multiple languages and has multi-language and cross-language analysis capabilities. The input of the data indexing model can be input text of a certain length, and it can implement retrieval tasks at different granularities such as sentences, paragraphs, chapters, and documents. The data indexing model integrates dense retrieval, sparse retrieval, multi-vector retrieval, etc., and can support different semantic retrieval scenarios in one stop.

[0132] In the embodiments of this application, the methods for matching the data in the dataset with the user's input information include, but are not limited to, matching methods such as similarity matching and coincidence matching.

[0133] Exemplarily, this dataset can be understood as a knowledge base in the technical field of Retrieval-augmented Generation (RAG) and is used to support the knowledge acquisition and application of the agent.

[0134] In the embodiments of this application, the processing of the matching data by the text generation model includes text understanding, data screening, etc., thus ensuring the accuracy of the target parameter information.

[0135] Optionally, according to the input information and the matching data, a prompt word is generated; the prompt word is input into the text generation model to generate the target parameter information.

[0136] Exemplarily, taking the vector retrieval method as an example, the agent vectorizes the data in the dataset and stores it in the vector database. When the user asks a question (an example of the user's input information), the data indexing model vectorizes the question, then retrieves relevant document fragments in the vector database, and then forms a prompt together with the user's question and sends it into the text generation model for text generation. The text generation model will refer to this information to generate richer and more accurate text content, which includes but is not limited to the target parameter information required to call the target tool.

[0137] It can be understood that when the agent calls the target tool, it passes the target parameter information to the target tool, and then the target tool performs the target operation indicated by the input information on the page according to the target parameter information.

[0138] Optionally, after the target tool performs the target operation indicated by the input information on the page, it can return the execution result to the agent; the agent inputs the execution result into the text generation model; the text generation model generates the final output information according to the execution result of the target tool; the agent returns the output information to the user. The output information is, for example, "The page operation task has been completed".

[0139] Exemplarily, taking page jump as the target operation, the target parameter information is, for example, the jump page text and the jump route. The target tool is the tool that performs the page jump operation after inputting the target parameter information. The jump page text represents the web page before and after the jump. The jump route represents the route from the web page before the jump to the web page after the jump. Among them, the route can be a set of predefined functions, methods, classes, and protocols; the route is used for page jump and parameter passing; by configuring the parameter information in the routing program, the target tool can jump the page (i.e., the page before the jump) to the page specified by the parameter information (i.e., the page after the jump).

[0140] In one example, the user input information is "Jump to page A". Based on the user input information, the recognized user intent is "page jump". According to the user intent, the matched target tool is the page jump tool, such as the function "<router-link to="{path:' / ***'}">". According to the user input information, the matched data retrieved from the dataset is multiple pieces of matched data that include page A. Based on the input information and the matched data, the target parameter information in the text generated by the text generation model is "page element data of page A". According to the target parameter information, the script file for invoking the target tool is "<router-link to="{path:' / page element data of page A'}">". In this way, after executing this script file, it is possible to jump to page A. Among them, in this script file, if ' / ' is set in front of the parameter of page A, it means jumping to page A starting from the root route; if ' / ' is not set, it means jumping to page A starting from the current route.

[0141] In the embodiments of the present application, the intelligent agent can automatically complete the target operation indicated by the input information (such as voice, text, etc.) by obtaining the user's input information, without the need for the user to manually operate, improving the user experience and operation efficiency, and solving many problems in the related art.

[0142] Figure 3 It is a schematic flowchart of a page operation method executed by a page operating system provided in an embodiment of the present application. It can be understood that as Figure 3 shown, in order to more clearly describe Figure 1 the process of the system 100 executing Figure 2 the page operation method, the embodiments of the present application also provide the process of the system 100 executing the page operation method, which mainly includes the following steps:

[0143] In the embodiments of the present application, the system 100 obtains the user's input information. The input information includes information about the user's target operation on the target page.

[0144] In the embodiments of the present application, the system 100 invokes the deployed intent recognition model, inputs the input information into the intent recognition model to output the user intent. The user intent represents the category of the target operation. Moreover, the system 100 invokes the deployed tool set, and based on the user intent, matches and obtains the target tool from the tool sets of multiple tools. Among them, the tool represents a functional module for performing operations on the page.

[0145] In the embodiments of the present application, system 100 calls the deployed data set, retrieves the data set according to the input information, and obtains the matching data, where the matching data includes parameter information related to performing the target operation. In addition, system 100 calls the deployed text generation model, inputs the input information and the matching data into the text generation model to generate text, and the text includes target parameter information. The target parameter information is obtained by the text generation model processing the parameter information in the matching data.

[0146] Optionally, according to the input information and the matching data, a prompt word is generated; the prompt word is input into the text generation model to generate the target parameter information.

[0147] In the embodiments of the present application, system 100 calls the target tool and transmits the target parameter information to the target tool, so that the target tool performs the target operation on the target page.

[0148] Optionally, after the target tool performs the target operation indicated by the input information on the target page, the execution result can be returned to the text generation model; the text generation model generates the final output information according to the execution result of the target tool and returns the output information to the user. Among them, the output information is, for example, "The page operation task has been completed".

[0149] Next, in combination with an example, a detailed example description will be given of system 100 and the page operation method provided in the embodiments of the present application.

[0150] Figure 4 It is a schematic diagram of the composition of an example of a page operation system provided in the embodiments of the present application. As Figure 4 shown, in one example, system 100 is the "Urban Digital Operating System". In the page of the "Urban Digital Operating System", there is a wake-up entry for the intelligent assistant (an example of an intelligent agent). The user starts the intelligent assistant through the wake-up entry and inputs information to the intelligent assistant; the intelligent agent processes the user's input information and performs the target operation indicated by the input information on the page of the "Urban Digital Operating System" to complete the user's intention. Among them, the process of the intelligent agent processing the user's input information will be introduced in detail below. It can be understood that the method of processing the user's input information can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities, such as an intelligent agent. Hereinafter, taking the intelligent agent as the execution subject, an exemplary description will be given.

[0151] Step S410, obtain the user's input information.

[0152] For example, the user's actions include: opening and logging into the "City Digital Operating System", calling out the smart assistant, and entering the core keyword "empty component shopping cart" through voice input or text. The intelligent agent can obtain the user's input information through the intelligent assistant: "empty component shopping cart". The intelligent agent's operations on the page include: automatic navigation and automatic selection of application for page elements such as automatic clicking on resource supermarkets, shopping carts, and component batch applications.

[0153] In the embodiment of the present application, the "city digital operating system" is an open smart city big data platform based on big data technology and a new generation of artificial intelligence technology, and is the digital and technical foundation of the smart city. The city digital operating system is equivalent to the operating system in a computer, which can manage various resources in the smart city and support various vertical applications such as smart transportation, smart planning, and smart energy. These applications can be understood as various application software in the operating system in the computer. It is a connector between the physical world and the digital world, a hub for urban data acquisition, management, analysis, and presentation, and a container for public components of data-related applications in the precipitation city. The city operating system is carried on the cloud computing infrastructure and under the smart city application. It manages the underlying data through the cloud infrastructure downward, and connects and schedules the hardware equipment in the city; it provides public components and interfaces upward to support the development, operation, and coordinated linkage of various smart city applications.

[0154] In the embodiments of the present application, the software platforms such as Web and app in the "City Digital Operating System" have the following characteristics: Frequent version iterations. For example, there are more and more functions and the operations are more and more complicated. If you are not an experienced user or a professional, it is difficult to understand all the functions. The system technology is outdated and the performance is poor. The entrances to some functions are not obvious or are hidden too deeply. Some functions are complicated to operate and there are too many processes. The software platform has its own defects. For example, after refreshing the page, the filled content disappears; after refreshing, you return to the initial page and need to re-operate. Some operations are frequent and repeated.

[0155] In the embodiment of the present application, in the process of realizing the automatic navigation operation of the page in the software platform, the communication barriers between AI, data and automation are broken. The embodiment of the present application designs and implements a set of standardized communication protocols and interfaces to ensure seamless connection between AI algorithms, data management systems and automated control systems, realize real-time transmission and sharing of information, realize data fusion and standardized processing, and realize the integration of AI algorithms and automated processes.

[0156] In the embodiments of the present application, perfect cooperation among multiple models, multi-dimensional data, and agents is achieved. In the embodiments of the present application, multiple AI models are integrated, such as neural network models like deep learning, machine learning, and reinforcement learning. Appropriate models or model combinations are selected according to task requirements to achieve more accurate and efficient decision-making and prediction. During the process of constructing the dataset, multi-dimensional data collection and vectorization processing are realized.

[0157] In the embodiments of the present application, during the process of constructing the toolset, a universal encapsulation technology with automation functions is adopted. Through functional modular design, multiple tools are obtained; the appropriate target tool is matched from multiple tools to achieve intelligent scheduling and load balancing; users can design and register tools by themselves, making the toolset have the characteristics of extensibility and compatibility.

[0158] Optionally, the agent can call a speech recognition model. When the input information of the user is speech, the agent inputs the speech input by the user into the speech recognition model. The speech recognition model recognizes the speech input by the user and outputs the text data corresponding to the speech input by the user.

[0159] Step S420, input the input information into an intent recognition model to output the user intent, where the user intent represents the category of the target operation.

[0160] In the embodiments of the present application, the agent can call a BERT model (an example of an intent recognition model). The agent inputs the input information into the BERT model, and the BERT model outputs the user intent. Exemplarily, the user intents output by the BERT model include: automatic jump, automatic filling, operation suggestions, function explanations, guess what you want to do, AI chat, etc.

[0161] Furthermore, when the agent recognizes the user intent of automatic jump, it can achieve automatic page jump, realizing automated page operation and improving user experience and operation efficiency.

[0162] Furthermore, when the agent recognizes the user intent of automatic filling, it can achieve automatic filling of form data, solving the inconvenience of filling forms caused by the defects of the software platform itself. For example, after the page is refreshed, the filled content disappears; after refreshing and returning to the initial page, the need to re-operate.

[0163] Furthermore, when the agent recognizes the user intent of operation suggestions, it can provide users with real-time operation tips and method guidance, helping users get started quickly and avoiding troubles caused by improper operations.

[0164] Furthermore, when the agent recognizes the user intention of function explanation, it can help the user understand the functions of the page, improving the convenience of the user's page operation and solving the problems of inconvenient operation caused by more and more functions, more and more complex operations, old system technologies, poor performance, and unclear or deeply hidden positions of some function entrances.

[0165] Furthermore, when the agent recognizes the user intention of "guess what you want to do", it can predict the user's operations, reducing the user's workload and improving the user experience.

[0166] Furthermore, when the agent recognizes AI chat, it can chat with the user in real time, solving the user's communication needs and improving the user experience.

[0167] Optionally, the agent can call a text generation model to have a multi-round conversation with the user, enabling the BERT model to recognize more accurate user intentions.

[0168] Optionally, the dataset can be determined by data collection. The data collection methods include but are not limited to: collecting data through data logging to obtain user behavior data, common behavior data, etc. to get the dataset; leveraging the support ability of the database API to obtain user data, business data, etc. by docking with the business database API to get the dataset; manually entering or collecting page path data, page element data, form element data, etc. through scripts.

[0169] Optionally, during the data matching process, the agent calls a data indexing model. The data indexing model retrieves data matching the input information from the dataset to obtain the matching data.

[0170] Step S430: According to the user intention, match the target tool from multiple tools. The target tool represents the functional module for performing the target operation.

[0171] Furthermore, once the user intention is recognized, the agent will match it with the registered tools. The matching process may involve further parsing of the user intention, comparison between the user intention and the tool functions, etc. to ensure the selection of the accurate tool.

[0172] For example, compare the similarity between the user intention and the function information of each tool among multiple tools, and determine the tool with a similarity exceeding the preset threshold among the multiple tools as the target tool. Exemplarily, the preset threshold is 90%. When the input information is "jump to the shopping cart page" and the similarity between the function information "page jump" of a tool and the input information is 95%, then the tool with "page jump" is determined as the target tool.

[0173] Step S440: Call the target tool to enable the target tool to perform the target operation.

[0174] Optionally, the process of invoking the target tool includes: retrieving a data set according to the input information to obtain matching data, where the matching data includes parameter information related to performing the target operation; inputting the input information and the matching data into a text generation model to generate text, where the text includes target parameter information, and the target parameter information is obtained by the text generation model processing the parameter information in the matching data; and invoking the target tool according to the target parameter information.

[0175] Furthermore, the processing of the matching data by the text generation model includes text understanding, data screening, etc., thus ensuring the accuracy of the target parameter information. The intelligent agent invokes the text generation model, and the text generation model generates target parameter information according to the input information by matching the data set. When invoking the target tool, the intelligent agent passes the target parameter information to the target tool, and then the target tool automatically performs the target operation indicated by the input information on the page according to the target parameter information.

[0176] Furthermore, "invoking the target tool" can be understood as an "action" step of the intelligent agent. Generally, the working principle of the intelligent agent can be understood as an iterative process of "thinking, acting, observing". In the "thinking" step, the intelligent agent performs task decomposition and action planning; for example, making a judgment and analysis of the input information, and then determining how to perform the next action. In the "action" step, the intelligent agent uses the tool to complete the task according to the result of "thinking". In the "observing" step, the intelligent agent generates output information according to the feedback information (such as the execution result of the tool) after the "action" is completed.

[0177] Optionally, in the "thinking" step, the intelligent agent generates a prompt according to the matching data and the user's input information by using the prompt engineering of the large language model. The intelligent agent invokes the text generation model, inputs the input information and the prompt of the matching data into the text generation model, and the text generation model then generates the target parameter information.

[0178] Optionally, in the "action" step, the intelligent agent invokes a tool set. The intelligent agent matches and obtains the target tool from the tool sets of multiple tools according to the user's intention; the intelligent agent invokes the target tool according to the target parameter information obtained in the "thinking" step, so that the target tool performs the target operation indicated by the input information on the page.

[0179] Optionally, in the "observing" step, the intelligent agent obtains the execution result of the target tool performing the target operation indicated by the input information on the page. The intelligent agent returns the execution result of the target tool to the text generation model; the text generation model generates the final output information according to the execution result of the target tool and returns the output information to the user. Among them, the output information is, for example, "The page operation task has been completed".

[0180] Exemplarily, the toolset is a wrapper of Selenium tools for the intelligent agent to automatically call tools.

[0181] Optionally, to complete the above work, the basic capabilities that the intelligent agent needs to possess include but are not limited to Python programming function, Selenium tool call function, RAG retrieval-augmented generation function, API interface call function, and Agent artificial intelligence framework.

[0182] In the embodiment of the present application, an AI decision engine is provided. As the intelligent core of the system, the AI decision engine is responsible for receiving input information (such as user instructions, application program status, etc.), analyzing and understanding it through AI technologies such as deep learning and natural language processing, and generating corresponding operation strategies or instructions. This engine can continuously learn and optimize to adapt to the characteristics of different application programs and user needs.

[0183] In the embodiment of the present application, a dataset for retrieval-augmented generation (RAG) is provided. This dataset can be understood as a knowledge base of RAG. The functional module for calling this dataset is assumed to be a rule automatic generation module. The rule automatic generation module, based on the output of the AI decision engine, combines preset business rules and logics, and automatically generates operation rules or scripts (examples of parameter information) for specific application programs. These rules or scripts are designed to guide tools such as Selenium to perform specific operation tasks, such as clicking buttons, inputting information, verifying results, etc.

[0184] In the embodiment of the present application, a toolset for automatically performing page operations is provided. This toolset is, for example, the Selenium automation testing framework. As the execution layer of the system, the target tool in this toolset is used to receive operation rules or scripts (examples of target parameter information) from the RAG module and perform corresponding automated operations on the page of the target application program. The Selenium framework supports multiple browsers and platforms, and can simulate the operation behaviors of real users to achieve efficient application program testing and operations.

[0185] In the embodiment of the present application, a data interaction and storage module is provided. This module is responsible for transferring data between the AI decision engine, the RAG module, the Selenium framework, and external data sources. This module supports multiple data formats and protocols to ensure accurate and efficient data transmission. At the same time, this module is also responsible for storing key information such as operation logs and test results for subsequent analysis and optimization.

[0186] In the embodiments of the present application, a user interface (UI) is provided. This UI can provide a friendly interaction interface for users, supporting functions such as user input of instructions, voice wake-up, viewing operation results, and configuring parameters. Through interaction with the AI decision-making engine, the UI can accurately understand and feedback user intentions.

[0187] In the embodiments of the present application, the user experience is improved. The intelligent agent can intelligently recognize user intentions and automatically perform operations such as page jumping and form filling, greatly reducing the user's operation burden and enhancing the usability. The intelligent agent provides real-time operation tips and method guidance to help users quickly get started and avoid troubles caused by improper operations.

[0188] In the embodiments of the present application, the intelligence level of the intelligent agent is enhanced. By combining a speech-to-text tool and an AI intention recognition model, the intelligent agent can accurately understand user voice instructions and achieve true intelligent interaction. Through continuous learning and optimization, the system can adapt to changes in user behavior and provide more personalized services.

[0189] In the embodiments of the present application, the work efficiency of users is improved. The intelligent agent executes automated script code to implement the methods of the embodiments of the present application, enabling repetitive operations to be completed quickly, significantly improving the work efficiency of users. The intelligent agent can intelligently recommend or execute relevant operations according to user habits and needs, further saving user time. The intelligent agent can intelligently recognize user intentions and automatically perform operations such as page jumping and form filling, greatly reducing the user's operation burden and enhancing the usability. The intelligent agent provides real-time operation tips and method guidance to help users quickly get started and avoid troubles caused by improper operations.

[0190] In the embodiments of the present application, technological innovation and development are promoted. In the embodiments of the present application, the integration and innovation of technologies such as automated scripts, AI intention recognition, and speech-to-text are promoted, providing new ideas and directions for the technological development of related fields. The successful application of the intelligent agent will encourage more enterprises or individuals to invest in the research and application of intelligent technologies, promoting the progress and development of the entire industry.

[0191] In the embodiments of the present application, the data management function is optimized. The intelligent agent can comprehensively collect and analyze user data, behavior data, etc., providing rich training materials for the AI model, which helps to improve the recognition accuracy and generalization ability of the model. In the embodiments of the present application, the management and utilization of data are more efficient, which helps enterprises or individuals better understand user needs and further optimize products and services.

[0192] In the embodiment of the present application, an intelligent assistant and intelligent operation are provided for the city digital operating system. In other application scenarios, such as software projects of different industries and sizes, the method of the embodiment of the present application can also be applied. The embodiment of the present application can also be further integrated with other technologies (such as machine learning, deep learning, natural language processing, etc.) to achieve richer and more intelligent functions.

[0193] Figure 5 A schematic diagram of an example of a page provided in an embodiment of the present application. Figure 4 An example of system 100, correspondingly, as Figure 5 As shown, the embodiment of the present application also provides an example of a page in the system 100. In the application scenario of this example, the page in the city digital operating system is operated by the intelligent assistant to achieve the purpose of clearing the component shopping cart. The work of the intelligent assistant includes: automatically clicking on the automatic navigation of page elements such as the resource supermarket, shopping cart, and component batch application, and automatically checking the application and other operations.

[0194] In this example, the user's input information is "clear the component shopping cart". Exemplarily, the user opens and logs in to the "city digital operating system", calls out the smart assistant, and inputs the core keyword "clear the component shopping cart" through voice input or text input.

[0195] In this example, the user's intention is "auto jump" and "auto fill". "Auto jump" means clicking on a page element to jump to the page and then find the component shopping cart. "Auto fill" means that when performing a clear operation, you need to submit a form formed by the cleared components to the smart assistant, and the smart assistant will perform a clear operation on the components in the form based on the filled form.

[0196] In this example, the target tool matched from multiple tools can be any tool that implements functions such as page element clicks, automatic page jumps, automatic form filling, etc., such as a tool created and registered by the user himself, or an automated page operation tool such as Selenium.

[0197] In this example, the data indexing model vectorizes the collected dataset and stores the dataset in a vector database. When the user inputs information, the data indexing model vectorizes the input information and then retrieves matching data related to the input information in the vector database. The user's input information is "Empty the component shopping cart", and the matching data includes page element data, routing data, etc. The page element data includes page elements such as "Resource Supermarket", "Shopping Cart", "Component Shopping Cart List", "Batch Application", etc. The routing data refers to the routing for emptying the component shopping cart, including clicking on the page element "Resource Supermarket" in sequence, clicking on the page element "Shopping Cart", clicking on the checkbox of the page element "Component Shopping Cart List", clicking on the page element "Batch Application", automatically filling out the form, clicking on the page element "Submit Application", etc. The routing data includes multiple parameter sub-information, and the order of the multiple parameter sub-information in the routing data indicates the calling order of multiple target tools.

[0198] In an embodiment of this application, an example of the dataset shown in Table 1 is also provided. This example is only an exemplary description of the dataset and does not constitute a limitation on the dataset.

[0199] Exemplarily, when the user's input information is "Select the component shopping cart", the data indexing model first vectorizes the input information and retrieves the most relevant matching data to the input information in the vector database of the dataset, obtaining matching data 1 and matching data 2 in Table 1. Then, the retrieved matching data 1 and matching data 2 are combined with the user's input information to form a prompt, which is input into the text generation model to generate text including target parameters. Among them, the text generation model will refer to the description information and parameter information of matching data 1 and matching data 2, generate target parameter information, and output text that is richer and more accurate and includes target parameter information.

[0200] It can be understood that in actual applications, the agent can input multiple matching data related to the user's input information that match into the text generation model, and the text generation model can also screen and combine the input matching data, filter out inappropriate matching data, and then output richer and more accurate text data.

[0201] Example of the dataset in Table 1

[0202]

[0203] In this example, the matching data and the user's input information are combined to form a prompt, which is sent into the text generation model for text generation. The text generation model will refer to the information in the prompt and generate richer and more accurate text content, including but not limited to the target parameter information required by the target tool.

[0204] Exemplarily, the prompt of the text generation model is, for example, "Please answer the user's question based on the provided reference information: [matching data]: [input information]. During the process of generating the answer, please think about the steps to solve the user's question and the parameter information for each step."

[0205] In this example, according to the target parameter information in the text output by the text generation model, the target tool is called so that the target tool performs the target operation indicated by the user's input information on the page.

[0206] In this example, an exemplary description is given of the detailed steps for the target tool to perform operations on the page. The data required for each of the following steps is the target parameter information required by the target tool. Among them, multiple target tools include a page element click tool and a page filling tool. The calling process of the multiple target tools includes: first calling the page element click tool; then calling the page filling tool. The process for the page element click tool to perform page element clicks: a. Recall the "Resource Supermarket" element data from the knowledge base and click; b. Recall the "Shopping Cart" element data from the knowledge base and click; c. Recall the "Component Application Batch Selection" element data from the knowledge base and click; d. Recall the "Batch Application" element data from the knowledge base and click. The process for the page filling tool to perform automatic form filling: Obtain form data through methods such as multi-round extraction of user input information and extraction of user information from the database; recall the "Component Application Form" element data from the knowledge base and complete automatic input; recall button elements such as "Submit Purchase Application" and "Submit Deletion Application" from the knowledge base, and click according to the user's intention. When the user's intention is to "empty the component shopping cart", click button elements such as "Submit Deletion Application". Among them, the knowledge base includes the above-mentioned data set and knowledge data; the knowledge data is, for example, the page element data of the "Urban Digital Operating System".

[0207] Exemplarily, the core code for target tool calling is shown in Table 2. The intelligent assistant can match relevant data of the user, page element data, current page data, etc. from the data set, and execute the core code for target tool calling to call the target tool and complete the user's intention.

[0208] Table 2 Core Code for Target Tool Calling

[0209]

[0210] Next, based on the system 100 and method in the above embodiments, a page operation device provided in the embodiments of the present application will be introduced.

[0211] Exemplarily, Figure 6 is a schematic structural diagram of a page operation device provided in the embodiments of the present application. As Figure 6As shown in the figure, a page operation device 600 mainly includes the following modules:

[0212] An acquisition module 610, configured to acquire input information of a user. The input information represents information about a target operation performed by the user on a target page.

[0213] A processing module 620, configured to input the input information into an intention recognition model to output a user intention, where the user intention represents a category of the target operation; and, according to the user intention, match a target tool from multiple tools, where the target tool represents a functional module for performing the target operation; and, call the target tool to enable the target tool to perform the target operation.

[0214] In a possible implementation manner, the multiple tools include at least one of a page element click tool, a page jump tool, and a page filling tool; the page element click tool represents a module for performing a page element click operation, the page jump tool represents a module for performing a page jump operation, and the page filling tool represents a tool for performing a page filling operation.

[0215] In a possible implementation manner, the above-mentioned processing module 620 is specifically configured to: retrieve a data set according to the input information to obtain matching data, where the matching data includes parameter information related to performing the target operation; input the input information and the matching data into a text generation model to generate text, where the text includes target parameter information, and the target parameter information is obtained by the text generation model processing the parameter information in the matching data; and call the target tool according to the target parameter information.

[0216] In a possible implementation manner, there are multiple target tools, and the target parameter information includes multiple parameter sub-informations, and each parameter sub-information represents parameter information of one target tool among the multiple target tools. The above-mentioned processing module 620 is specifically configured to: determine the calling order of the multiple target tools according to the order of the multiple parameter sub-informations in the target parameter information; and input the multiple parameter sub-informations into the multiple target tools in sequence according to the calling order, so that the multiple target tools are called in sequence according to the calling order.

[0217] In a possible implementation manner, the above-mentioned processing module 620 is specifically configured to: use the data set and the input information as inputs of a data indexing model, so that the data indexing model matches the data of the input information and the data set to output matching data.

[0218] In a possible implementation manner, the data in the data set is vectorized data. The above-mentioned processing module 620 is specifically configured to: perform vectorization processing on the input information to obtain vectorized input information; and match the vectorized data in the data set with the vectorized input information to obtain matching data.

[0219] In a possible implementation, the data in the dataset includes at least one of page path data, page element data, and form element data.

[0220] In a possible implementation, the methods for collecting data in the dataset include at least one of collecting data through logging, accessing the application programming interface of the business database, manual entry, or collection through scripts.

[0221] In a possible implementation, the above-mentioned processing module 620 is specifically configured to: generate a prompt word based on the input information and the matching data; input the prompt word into the text generation model to generate text.

[0222] In a possible implementation, the above-mentioned processing module 620 is specifically configured to: compare the similarity between the user intention and the function information of each tool among multiple tools, and determine the tools with similarity exceeding the preset threshold among the multiple tools as target tools.

[0223] In a possible implementation, multiple tools are obtained through user definition and registration.

[0224] In a possible implementation, the above-mentioned processing module 620 is further configured to: when the data type of the input information is voice, input the voice data into the speech recognition model to convert the input information into text form.

[0225] In a possible implementation, the above-mentioned input information is obtained through multiple rounds of conversations and / or historical conversations between the user and the intelligent agent.

[0226] In a possible implementation, the above-mentioned processing module 620 is further configured to: when the user intention does not meet the preset requirements, input the input information and the user intention into the text generation model to output a reply text, and the reply text is used to generate multiple rounds of conversations between the user and the intelligent agent. The above-mentioned processing module 620 is specifically configured to: input the multiple rounds of conversation data between the user and the intelligent agent into the intention recognition model to output the user intention.

[0227] In a possible implementation, the above-mentioned processing module 620 is further configured to: obtain the execution result of the target tool, where the execution result represents the result of the target tool performing the target operation; input the execution result into the text generation model to generate output information; output the above-mentioned output information to the user.

[0228] In a possible implementation, Figure 6 The multiple functional modules shown in can be implemented by software or can be implemented by hardware. Exemplarily, taking the processing module 620 as an example, the implementation manner of the processing module 620 will be introduced below. Similarly, the implementation manners of other functional modules in the page operation device 600 can refer to the implementation manner of the processing module 620.

[0229] As an example of a software functional unit, the processing module 620 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the processing module 620 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0230] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0231] As an example of a hardware functional unit, the processing module 620 may include at least one computing device, such as a server, etc. Alternatively, the processing module 620 may also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0232] The multiple computing devices included in the processing module 620 may be distributed in the same region or in different regions. The multiple computing devices included in the processing module 620 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the processing module 620 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).

[0233] It should be noted that in other embodiments, the page operation device 600 is additionally provided with one or more modules for performing any of the steps included in the above implementation. The steps to be implemented by one or more modules in the page operation device 600 can be specified as needed. The functions of the page operation device 600 can also be implemented using more or fewer modules than those in the embodiments of the present application. One or more modules in the page operation device 600 are used to implement different steps in the above method respectively, so as to implement all the functions of the page operation device 600.

[0234] The present application also provides a computing device 700. As Figure 7 shown, the computing device 700 includes: a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other through the bus 702. The computing device 700 may be a server, such as a central server, an edge server, or a local server in a local data center, or may be an electronic device such as a desktop computer, a laptop computer, or a smart phone. It should be understood that the present application does not limit the number of processors and memories in the computing device 700.

[0235] The bus 702 can be a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. Among them, the unified bus is also known as the Lingqu bus. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 it is represented by only one line in Figure 7 , but it does not mean that there is only one bus or one type of bus. The bus 704 can include a path for transmitting information between various components of the computing device 700 (for example, the memory 706, the processor 704, the communication interface 708).

[0236] The processor 704 can include any one or more of computing devices such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP), an ASIC, an FPGA, a CPLD, an NPU, a SoC, an offload card, an acceleration card, etc.

[0237] The memory 706 can include volatile memory, such as Random Access Memory (RAM). The processor 704 can also include non-volatile memory, such as Read-Only Memory (ROM), flash memory, a Hard Disk Drive (HDD), or a Solid State Drive (SSD). In addition, the memory 706 can also be implemented through Storage Class Memory (SCM), Phase Change Memory (PCM), or other types of storage media.

[0238] It should be noted that in the same computing device, the same type of storage medium can be configured to implement the function of the memory 706, or two or more types of storage media can be configured to implement the function of the memory 706. This application does not make any limitations in this regard.

[0239] The memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement the functions of the foregoing one or more modules, thereby implementing the method described in the foregoing embodiments. That is to say, the memory 706 stores instructions for executing the method described in the foregoing embodiments.

[0240] Alternatively, the memory 706 stores executable codes, and the processor 704 executes the executable codes to respectively implement the Figure 6 functions of the page operation device 600 shown in the foregoing, thereby implementing the method described in the foregoing embodiments. That is to say, the memory 706 stores instructions for executing the method described in the foregoing embodiments.

[0241] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 700 and other devices or communication networks.

[0242] As a possible implementation manner, the computing device 700 may also include a chip system. The chip system includes a processor and a power supply circuit. The power supply circuit is used to supply power to the processor, and the processor is used to execute the operation steps corresponding to the method of the embodiments of the present application. For the sake of brevity, it will not be elaborated herein. Among them, the processor can be implemented by a GPU, or can be implemented by a computing device or an AI chip such as a DPU, an NPU, an XPU, an SoC, an offloading card, or an acceleration card.

[0243] As a possible implementation manner, the computing device 700 may include multiple types of processors 704, that is, the computing device 700 is a heterogeneous device. For example, the computing device 700 includes a CPU and a GPU, and at least one of the processors 704 can execute the operation steps corresponding to the method of the embodiments of the present application. For the sake of brevity, it will not be elaborated herein.

[0244] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be an electronic device such as a desktop computer, a laptop computer, or a smart phone.

[0245] Such as Figure 8As shown, the computing device cluster includes at least one computing device 700. Instructions for executing the methods described in the foregoing embodiments may be stored in the memories 706 of one or more of the computing devices 700 in the computing device cluster.

[0246] In some possible implementation manners, partial instructions for executing the methods described in the foregoing embodiments may also be stored separately in the memories 706 of one or more of the computing devices 700 in the computing device cluster. In other words, a combination of one or more computing devices 700 may jointly execute the instructions for executing the methods described in the foregoing embodiments.

[0247] It should be noted that the memories 706 in different computing devices 700 in the computing device cluster may store different instructions, respectively for executing partial functions of the page operation device 600 shown in the foregoing Figure 6 That is, the instructions stored in the memories 706 of different computing devices 700 may implement the functions of one or more modules in the page operation device 600.

[0248] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Wherein, the network may be a wide area network or a local area network, etc. Figure 9 This is a possible implementation manner. As Figure 9 shown, two computing devices 700A and 700B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manners, the memories 706 in the computing device 700A store instructions for the functions of one or more modules in the page operation device 600. At the same time, the memories 706 in the computing device 700B store instructions for the functions of other one or more modules in the page operation device 600.

[0249] It should be understood that Figure 9 the functions of the computing device 700A shown in

[0250] may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700. Figure 8 and Figure 9 The connection manner of the computing device cluster described above. The difference is that instructions for executing the methods in the foregoing embodiments may be stored in the memories 706 of one or more of the computing devices 700 in this computing device cluster.

[0251] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store part of the instructions for executing the foregoing data processing method respectively. In other words, the combination of one or more computing devices 700 may jointly execute the instructions for executing the foregoing method.

[0252] In addition to the foregoing methods and electronic devices, embodiments of the present application may further provide a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present application described in the "Methods" section of this specification. Among them, the computer program product may be written in any combination of one or more programming languages for computer program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. Among them, the computer program code may be in the form of source code, object code, executable file or some intermediate form, etc. The computer program code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In some examples, when the computer program instructions are executed by a computing device cluster including at least one computing device, at least one computing device in the computing device cluster is caused to execute the methods in the foregoing embodiments.

[0253] In addition, embodiments of the present application may further provide a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present application described in the "Methods" section of this specification. The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. In some examples, when the instructions are run by a cluster of computing devices including at least one computing device, at least one computing device in the cluster of computing devices is caused to execute the methods in the above embodiments.

[0254] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0255] The method steps in the embodiments of this application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0256] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.

[0257] It can be understood that the various numerical numbers involved in the embodiments of this application are only for convenience of description and are not used to limit the scope of the embodiments of this application.

[0258] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0259] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the specific details disclosed above are only for the purposes of illustration and facilitating understanding, rather than limitations. These details do not limit the present application to necessarily adopting the above specific details for implementation.

[0260] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0261] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.

[0262] The above description has been given for purposes of illustration and description. In addition, this description does not intend to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

[0263] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A page operation method, characterized in that, The method includes: Obtaining the input information of the user, where the input information includes information about the user's target operation on the target page; Inputting the input information into an intent recognition model to output a user intent, where the user intent represents the category of the target operation; According to the user intent, matching a target tool from multiple tools, where the target tool represents a functional module for performing the target operation; Invoking the target tool to enable the target tool to perform the target operation.

2. The method according to claim 1, wherein The multiple tools include at least one of a page element click tool, a page jump tool, and a page filling tool; the page element click tool represents a module for performing a page element click operation, the page jump tool represents a module for performing a page jump operation, and the page filling tool represents a tool for performing a page information filling operation.

3. The method according to claim 1 or 2, characterized in that, The invoking the target tool includes: Retrieving a dataset according to the input information to obtain matching data, where the matching data includes parameter information related to performing the target operation; Inputting the input information and the matching data into a text generation model to generate text, where the text includes target parameter information, and the target parameter information is obtained by the text generation model processing the parameter information in the matching data; Invoking the target tool according to the target parameter information.

4. The method according to claim 3, wherein There are multiple target tools, and the target parameter information includes multiple parameter sub-informations, and each parameter sub-information represents the parameter information of one target tool among the multiple target tools; The invoking the target tool according to the target parameter information includes: Determining the invocation order of the multiple target tools according to the order of the multiple parameter sub-informations in the target parameter information; According to the invocation order, sequentially inputting the multiple parameter sub-informations into the multiple target tools to enable the multiple target tools to be invoked sequentially according to the invocation order.

5. The method according to claim 3 or 4, characterized in that, The inputting the input information and the matching data into a text generation model to generate text includes: Generating a prompt word according to the input information and the matching data; Inputting the prompt word into the text generation model to generate the text.

6. The method according to any one of claims 1 to 5, characterized in that, The data in the dataset includes at least one of page path data, page element data, and form element data.

7. The method according to any one of claims 1-6, characterized in that, The matching a target tool from multiple tools according to the user intent includes: Comparing the similarity between the user intent and the function information of each tool among the multiple tools, and determining the tool with a similarity exceeding a preset threshold among the multiple tools as the target tool.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: When the user intent does not meet the preset requirements, inputting the input information and the user intent into a text generation model to output a reply text, where the reply text is used to generate the user's multi-round conversation; The inputting the input information into an intent recognition model to output a user intent includes: Inputting the multi-round conversation data of the user into the intent recognition model to output a user intent.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: Obtaining the execution result of the target tool, where the execution result represents the result of the target tool performing the target operation; Input the execution result into a text generation model to generate output information; Output the output information to the user.

10. A server, characterized in that, The server is used to execute the method according to any one of claims 1-9.