Script generation and execution method and device, storage medium and program product

Through the combination of simulated interactive operation, screenshots and control recognition models, the problems of low efficiency and low accuracy of automated script generation in the prior art are solved, and efficient and accurate automated script generation is achieved.

CN120010963APending Publication Date: 2025-05-16BEIJING 58 INFORMATION TTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510124188.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is inefficient and inaccurate when generating automated scripts, mainly due to the complexity of page control tree levels and the lack of control IDs.

Method used

By obtaining the interactive operation description information of the target application, the simulation performs these operations to generate an initial page, perform page screenshots, and use the control identification model to identify the target control based on the description information and screenshots, and finally generate efficient and accurate automated scripts.

Benefits of technology

Without relying on the page control tree, it can generate automated scripts more efficiently and accurately, improving the efficiency and accuracy of script generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010963A_ABST
    Figure CN120010963A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a script generation and execution method and device, a storage medium and a program product. In the method, description information of a plurality of interactive operations and a plurality of pages generated by controlling a target application to simulate and execute the plurality of interactive operations based on the description information of the plurality of interactive operations can be obtained; performing screenshot on the plurality of pages to obtain page screenshots corresponding to the plurality of interactive operations; and inputting the description information of the plurality of interactive operations and the page screenshots corresponding to the plurality of interactive operations into a control identification model according to the execution sequence of the plurality of interactive operations, so as to identify the page screenshots according to the action description information in the description information of any interactive operation. Identifying a target control for any interactive operation from a plurality of controls in the page screenshot corresponding to any interactive operation; and generating a target script according to the description information of the plurality of interactive operations, the execution sequence and the target control for the plurality of interactive operations. In this way, the automatic script can be efficiently and accurately generated without depending on a page control tree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automated control technology, and in particular to a script generation and execution method, device, storage medium and program product. Background Art

[0002] With the rapid development of mobile Internet, mobile applications on terminals have become an indispensable part of people's lives. Whether it is automated testing of mobile applications or using mobile applications to simulate user operations, automated scripts are needed to improve efficiency.

[0003] In the prior art, both the generation stage of the automation script and the application stage of the automation script rely on traversing the page control tree to locate the page control corresponding to the operation to be performed, so as to perform corresponding operations on the located page control.

[0004] On the one hand, the levels in the page control tree are usually numerous and complex, which results in the need to frequently traverse the page control tree each time an automation script is used for automation operations, which is inefficient. On the other hand, some pages have missing control IDs (Identity documents) in the page control tree, which makes it impossible to accurately traverse to the page control corresponding to the operation to be performed, resulting in low accuracy of the automation script finally generated.

[0005] Based on the above two aspects, the existing technology cannot generate automation scripts more efficiently, and the accuracy of the generated automation scripts is low. Summary of the invention

[0006] Various aspects of the present application provide a script generation and execution method, device, storage medium, and program product for generating automation scripts more efficiently and accurately.

[0007] An embodiment of the present application provides a script generation method, comprising: obtaining description information of multiple interactive operations for a target task of a target application, and, based on the description information of the multiple interactive operations, controlling the target application to simulate executing multiple initial pages generated by the multiple interactive operations; taking screenshots of the multiple initial pages to obtain page screenshots corresponding to each of the multiple interactive operations; inputting the description information of the multiple interactive operations and the page screenshots corresponding to each of the multiple interactive operations into a control recognition model in sequence according to the execution order of the multiple interactive operations, so as to identify a target control targeted by any interactive operation from multiple controls in the page screenshot corresponding to any interactive operation according to action description information in the description information of any interactive operation; generating a target script according to the description information of the multiple interactive operations, the execution order and the target controls targeted by each of the multiple interactive operations; the target script is used to: control the target application to execute the multiple interactive operations on the target controls targeted by each of the multiple interactive operations in sequence according to the execution order according to the description information of the multiple interactive operations to complete the target task.

[0008] Optionally, obtaining description information of multiple interactive operations for a target task of a target application includes: providing a task configuration interface of the target application to a user so that the user can configure the target task; in response to the user's configuration operation, determining description information of multiple interactive operations of the target task; or, using historical description information of multiple historical interactive operations in a historical target task as description information of the multiple interactive operations; the historical target task is a task performed for a historical version of a target application.

[0009] Optionally, the action description information of any interactive operation is used to describe the operation type of any interactive operation, including: action description text and / or action description icon; according to the action description information in the description information of any interactive operation, identifying the target control targeted by any interactive operation from multiple controls in the page screenshot corresponding to any interactive operation, including: extracting first feature information corresponding to the action description information of any interactive operation, and second feature information corresponding to each of multiple screenshot areas of the page screenshot corresponding to any interactive operation; the multiple screenshot areas respectively display the multiple controls; the second feature information corresponding to any screenshot area is used to characterize the control displayed in the screenshot area; calculating the feature similarity between the first feature information and the second feature information corresponding to each of the multiple screenshot areas; taking the screenshot area whose feature similarity among the multiple screenshot areas meets preset conditions as the target area; taking the control displayed in the target area as the target control targeted by any interactive operation.

[0010] Optionally, the control recognition model pre-learns the region segmentation capability for different page screenshots; extracts the second feature information corresponding to each of the multiple screenshot regions of the page screenshot corresponding to any interactive operation, including: utilizing the region segmentation capability to segment the page screenshot corresponding to any interactive operation into the multiple screenshot regions corresponding to the multiple controls; and extracts the second feature information corresponding to each of the multiple screenshot regions.

[0011] Optionally, a target script is generated based on the description information of the multiple interactive operations, the execution order, and the target controls targeted by the multiple interactive operations, including: generating multiple operation instructions based on the description information of the multiple interactive operations and the target controls targeted by the multiple interactive operations, any operation instruction being used to control the target application to perform a corresponding interactive operation on any target control; and integrating the multiple operation instructions into the target script according to the execution order.

[0012] Optionally, it also includes: obtaining a first sample page screenshot, sample action description information of a sample interactive operation, and a second sample page screenshot marked with a sample control; inputting the first sample page screenshot and the sample action description information into a preset control recognition network, and under the supervision of the second sample page screenshot, training the control recognition network with the first sample page screenshot and the sample action description information, with the goal of converging the loss function of the control recognition network to a target range, to obtain the control recognition model; wherein the loss function is used to calculate the error between the candidate control targeted by the sample interactive operation in the sample page screenshot identified by the control recognition network and the sample control.

[0013] An embodiment of the present application also provides a script execution method, applicable to a target application, comprising: obtaining a target script; generating the target script according to the script generation method; parsing the target script to obtain description information, execution order, and target controls of multiple interactive operations for a target task of the target application; and executing the multiple interactive operations on the target controls respectively targeted by the multiple interactive operations in sequence according to the execution order based on the description information of the multiple interactive operations to complete the target task.

[0014] An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; wherein the memory is used to: store one or more computer instructions; the processor is used to execute the one or more computer instructions to: execute the steps in the script generation method or the script execution method.

[0015] An embodiment of the present application also provides a computer-readable storage medium, which, when the computer program is executed by a processor, enables the processor to implement the steps in the script generation method or the script execution method.

[0016] An embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps of any one of the script generation method and the script execution method.

[0017] In this embodiment, description information of multiple interactive operations and multiple pages generated by simulating the execution of multiple interactive operations by controlling the target application based on the description information of multiple interactive operations can be obtained; screenshots of multiple pages are taken to obtain page screenshots corresponding to multiple interactive operations; the description information of multiple interactive operations and the page screenshots corresponding to multiple interactive operations are input into the control recognition model according to the execution order of multiple interactive operations, so as to identify the target control targeted by any interactive operation from the multiple controls in the page screenshot corresponding to any interactive operation according to the action description information in the description information of any interactive operation; and the target script is generated according to the description information of multiple interactive operations, the execution order and the target control targeted by multiple interactive operations. In this way, the automation script can be generated more efficiently and accurately without relying on the page control tree. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1 A flowchart of a script generation method provided by an exemplary embodiment of the present application;

[0020] Figure 2 A flowchart of a script generation method in an actual scenario provided by an exemplary embodiment of the present application;

[0021] Figure 3 A flowchart of a script execution method provided by an exemplary embodiment of the present application;

[0022] Figure 4 A schematic diagram of an electronic device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0024] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) are in compliance with relevant laws and standards.

[0025] In the prior art, both the generation stage of the automation script and the application stage of the automation script rely on traversing the page control tree to locate the page control corresponding to the operation to be performed, so as to perform corresponding operations on the located page control.

[0026] On the one hand, the levels in the page control tree are usually numerous and complex, which results in the need to frequently traverse the page control tree each time an automation script is used for automation operations, which is inefficient. On the other hand, there are cases where the control ID (Identity document) is missing in the page control tree of some pages, which results in the inability to accurately traverse the page control corresponding to the operation to be performed, resulting in a low accuracy of the automation script finally generated. Based on the above two aspects, the prior art cannot generate automation scripts more efficiently, and the accuracy of the generated automation scripts is low.

[0027] In response to the above technical problems, in an embodiment of the present application, a control recognition model can be used to generate an automation script more efficiently and accurately based on the description information of multiple interactive operations and their corresponding page screenshots without relying on a control tree.

[0028] Figure 1 The script generation method provided by the exemplary embodiment of the present application is as follows: Figure 1 As shown, the method may include the following steps:

[0029] Step 11: Obtain description information of multiple interactive operations for a target task of a target application, and control the target application to simulate executing multiple initial pages generated by the multiple interactive operations based on the description information of the multiple interactive operations.

[0030] Step 12: Take screenshots of multiple initial pages to obtain page screenshots corresponding to multiple interactive operations.

[0031] Step 13. According to the execution order of the multiple interactive operations, the description information of the multiple interactive operations and the page screenshots corresponding to the multiple interactive operations are input into the control recognition model in turn, so as to identify the target control targeted by any interactive operation from the multiple controls in the page screenshot corresponding to any interactive operation based on the action description information in the description information of any interactive operation.

[0032] Step 14: Generate a target script based on the description information of the multiple interactive operations, the execution order, and the target controls targeted by the multiple interactive operations; the target script is used to: control the target application to perform the multiple interactive operations on the target controls targeted by the multiple interactive operations in sequence according to the execution order according to the description information of the multiple interactive operations to complete the target task.

[0033] The execution subject of this embodiment can be any electronic device, such as a server or a terminal device. The terminal device can be a computer, a tablet computer or a mobile phone, etc., and this embodiment does not impose any limitation.

[0034] In this embodiment, the target application can be any type of application, such as a game application, a navigation application, or a housekeeping application, etc., which is not limited in this embodiment. The target task can be a test task or any other task for automating the control of the application, which is not limited in this embodiment. The target task may include multiple interactive operations on the target application, and any interactive operation may be any operation such as a search operation, a delete operation, or a return operation. Any interactive operation may have description information, which can be used to describe the details of the interactive operation, and may include: action description information and instruction description information. The action description information is used to describe the operation type of the interactive operation, and the instruction description information is used to describe the instruction details of the interactive operation. This embodiment does not limit the specific implementation of the operation type, and the operation type can be set to any type according to actual needs, such as a search type, a delete type, or a return type.

[0035] For example, the action description information of a certain interaction operation may be used to describe the interaction operation as a search-type action. For another example, the action description information of a certain interaction operation may be used to describe the interaction operation as a delete-type action.

[0036] In this embodiment, the target application can be controlled to simulate the execution of multiple interactive operations based on the description information of multiple interactive operations in advance, thereby generating multiple initial pages. When the target application is controlled to simulate the execution of any interactive operation based on the description information of any interactive operation, the target application can be controlled to simulate the execution of the interactive operation based on the instruction description information in the description information.

[0037] For example, the target task for the target application includes three interactive operations, namely interactive operations Z1-Z3. Then, based on the description information of interactive operation Z1, the target application can be controlled to simulate the execution of interactive operation Z1. After the execution of interactive operation Z1, the current page of the target application is page H1, so page H1 is an initial page; based on the description information of interactive operation Z2, the target application can be controlled to simulate the execution of interactive operation Z2. After the execution of interactive operation Z2, the target application will jump from page H1 to page H2, so page H2 is an initial page; based on the description information of interactive operation Z3, the target application can be controlled to simulate the execution of interactive operation Z3. After the execution of interactive operation Z3, the target application does not jump to the page, so page H2 after interactive operation Z3 is an initial page. It should be noted that whether the target application jumps to the page with the interactive operation depends on the actual page jump logic of the target application, and this embodiment does not limit this.

[0038] Based on this, description information of multiple interactive operations for the target task of the target application can be obtained, and based on the description information of the multiple interactive operations, the target application can be controlled to simulate the execution of multiple initial pages generated by the multiple interactive operations; and screenshots of the multiple initial pages are taken to obtain page screenshots corresponding to the multiple interactive operations. When taking screenshots of the multiple initial pages, the built-in screenshot function or screenshot API (Application Programming Interface) of the target application can be used to take screenshots, or the screenshot function in the operating system can be used to take screenshots, which is not limited in this embodiment.

[0039] Afterwards, the description information of the multiple interactive operations and the page screenshots corresponding to the multiple interactive operations may be input into the control recognition model in sequence according to the execution order of the multiple interactive operations.

[0040] The control recognition model can identify the target control for any interactive operation from multiple controls in the page screenshot corresponding to any interactive operation based on the action description information in the description information of any interactive operation. For example, the action description information in the description information of a certain interactive operation is used to describe the interactive operation as a search-type action. The control recognition model can identify the target control for any interactive operation from multiple controls in the page screenshot corresponding to the interactive operation based on the action description information. There is at least one control related to the search type among the multiple controls in the page screenshot corresponding to the interactive operation. Therefore, the at least one control can be identified from the multiple controls and determined as the target control for the interactive operation.

[0041] Afterwards, a target script may be generated according to the description information of the multiple interactive operations, the execution order, and the target controls for each of the multiple interactive operations. The target script may be used to control the target application to perform multiple interactive operations on the target controls for each of the multiple interactive operations in the execution order according to the description information of the multiple interactive operations to complete the target task.

[0042] In this method, description information of multiple interactive operations and multiple pages generated by simulating and executing multiple interactive operations by controlling a target application based on the description information of multiple interactive operations can be obtained; screenshots of multiple pages are taken to obtain page screenshots corresponding to multiple interactive operations; the description information of multiple interactive operations and the page screenshots corresponding to multiple interactive operations are input into a control recognition model according to the execution order of multiple interactive operations, so as to identify the target control targeted by any interactive operation from multiple controls in the page screenshot corresponding to any interactive operation according to the action description information in the description information of any interactive operation; and a target script is generated according to the description information of multiple interactive operations, the execution order and the target controls targeted by multiple interactive operations. In this way, the automation script can be generated more efficiently and accurately without relying on the page control tree.

[0043] In some optional embodiments, when obtaining the description information of multiple interactive operations for the target task of the target application, the task configuration interface of the target application may be provided to the user, and the user may configure the target task through the task configuration interface. Correspondingly, the description information of multiple interactive operations of the target task may be determined in response to the configuration operation of the user. Among them, the configuration operation may be implemented as any type of operation, for example, the configuration operation may be divided into two operations, one is an input operation, and the user may perform an input operation to input the description information of multiple interactive operations of the target task; the other is a drop-down list operation, and the user may click on the drop-down list and select multiple description information from the preset multiple description information as the description information of multiple interactive operations of the target task. In addition, the historical description information of multiple historical interactive operations in the historical target task may also be stored in the local memory or the cloud server, wherein the historical target task is a task executed for the target application of the historical version. In other words, before the target application is updated, the historical target task has been executed for the target application of the historical version, and there are multiple historical description information of historical interactive operations in the historical target task. Based on this, the historical description information of multiple historical interactive operations in the historical target task can be used as the description information of multiple interactive operations.

[0044] In this way, description information of multiple interactive operations for a target task of a target application can be obtained more accurately.

[0045] In some optional embodiments, the action description information of any interactive operation is used to describe the operation type of any interactive operation, and may include: at least one of an action description text and an action description icon. For example, if the action description information of any interactive operation is used to describe the operation type of any interactive operation as a search type, then the action description text included in the action description information of the interactive operation may be text related to the search type such as "search" or "search". For another example, if the action description information of any interactive operation is used to describe the operation type of any interactive operation as a delete type, then the action description icon included in the action description information of the interactive operation may be a trash can icon or other icons related to the delete type.

[0046] Based on this, step 13 in the aforementioned embodiment, "identifying a target control targeted by any interactive operation from multiple controls in a page screenshot corresponding to any interactive operation according to the action description information in the description information of any interactive operation", can be implemented by the following steps:

[0047] Step 131, extracting the first feature information corresponding to the action description information of any interactive operation, and the second feature information corresponding to each of the multiple screenshot areas of the page screenshot corresponding to any interactive operation. The multiple screenshot areas respectively display multiple controls. The second feature information corresponding to any screenshot area is used to characterize the control displayed in the screenshot area. The action description information of any interactive operation and the page screenshot corresponding to any interactive operation can be input into the feature extraction layer of the control recognition model to perform corresponding feature extraction.

[0048] Specifically, the control recognition model pre-learns the regional segmentation capability for different page screenshots, wherein the control recognition model can segment out the screenshot area showing the control in any page screenshot based on the regional segmentation capability. Based on this, when extracting the second feature information corresponding to each of the multiple screenshot areas of the page screenshot corresponding to any interactive operation, the regional segmentation capability can be used to segment the page screenshot corresponding to any interactive operation into multiple screenshot areas corresponding to multiple controls, and then the second feature information corresponding to each of the multiple screenshot areas can be extracted. Optionally, the control recognition model can learn the regional segmentation capability for different page screenshots in the following manner: first input multiple sample page screenshots and the sample segmentation results of each of the multiple sample page screenshots, input the multiple sample page screenshots and their respective sample segmentation results into the control recognition model, and under the supervision of the sample segmentation results of each of the multiple sample page screenshots, with the first loss function converging to the first target range as the goal, use the multiple sample page screenshots to train the control recognition model, so that it learns the regional segmentation capability for different page screenshots. Wherein, the loss function is used to calculate the error between the multiple segmentation results and the sample segmentation results generated and recognized by the control recognition model based on the multiple sample page screenshots.

[0049] Step 132: Calculate the feature similarity between the first feature information and the second feature information corresponding to each of the plurality of screenshot regions.

[0050] Step 133: The screenshot area whose feature similarity meets the preset condition among the multiple screenshot areas is used as the target area. The preset condition can be set according to the actual situation, for example, the preset condition can be that the feature similarity is greater than a certain threshold, or the feature similarity is the maximum.

[0051] Wherein, the first feature information and the second feature information are both semantic information, based on which, the feature similarity comparison between the first feature information and the plurality of second feature information is the comparison between the semantic information. Wherein, the feature similarity can be used to characterize: the degree of match between the action type described by the action description information and the function type of the control in the screenshot area, for example, if the action type described by the action description information of a certain interactive operation is a deletion type, then the function type of the control with the icon of "trash can" in the screenshot area is more matched.

[0052] Step 134: Use the control displayed in the target area as the target control for any interactive operation.

[0053] In this way, the control recognition model can more accurately identify the target control for any interactive operation from multiple controls in the page screenshot corresponding to the interactive operation.

[0054] In some optional embodiments, step 14 in the aforementioned embodiment, "generating a target script according to the description information and execution order of the multiple interactive operations and the target controls targeted by the multiple interactive operations", can be implemented based on the following steps:

[0055] Step 141: Generate multiple operation instructions based on the description information of multiple interactive operations and the target controls for each of the multiple interactive operations. Any operation instruction can be used to control the target application to perform a corresponding interactive operation on any target control. Specifically, an operation instruction corresponding to any interactive operation can be generated based on the instruction description information in the description information of any interactive operation and the target control for any interactive operation.

[0056] Step 142: Integrate multiple operation instructions into a target script according to the execution order. Here, a preset code generation rule can be used to integrate multiple operation instructions into any type of target script according to the execution order, such as Python type.

[0057] Based on any of the above embodiments, the control recognition model can be trained based on the following steps:

[0058] Step R1, obtain a first sample page screenshot, sample action description information of a sample interactive operation, and a second sample page screenshot marked with a sample control.

[0059] Step R2: input the first sample page screenshot and the sample action description information into a preset control recognition network. Under the supervision of the second sample page screenshot, with the loss function of the control recognition network converging to a target range, the control recognition network is trained using the first sample page screenshot and the sample action description information to obtain a control recognition model.

[0060] Among them, the loss function is used to calculate the error between the candidate control and the sample control targeted by the sample interaction operation in the sample page screenshot identified by the control recognition network. In this way, the control recognition type can be trained more accurately.

[0061] The following will explain the above script generation method in combination with actual scenarios:

[0062] like Figure 2As shown, the corresponding App (Application), i.e., the target application, can be started first. Then, the page screenshot can be taken in the following manner: based on the description information of the multiple interactive operations, the target application is controlled to simulate and execute the multiple interactive operations, thereby generating multiple initial pages, and the multiple initial pages are screenshoted to obtain the page screenshots corresponding to the multiple interactive operations.

[0063] After that, the big model (i.e., the control recognition model) can be told the relevant operations, that is, the description information of multiple interactive operations can be input into the control recognition model. Based on this, the big model can be waited for to return the operation target coordinates, that is, the coordinates of the target control targeted by any interactive operation in the page screenshot, so that the target control targeted by the interactive operation can be identified from the multiple controls in the page screenshot corresponding to the interactive operation based on the target coordinates.

[0064] In this way, on the one hand, scripts can be generated based on the description information of multiple interactive operations without determining the controls based on the page control tree. The large model can understand and identify the corresponding control coordinates based on the description information; on the other hand, Web-type pages, Rn (React Native, an open source cross-platform mobile application development framework) type pages, and Flutter (an application development framework) type pages can all support the above solution, where Web pages, RN pages, and Flutter pages are user interfaces built with three different technology stacks. In addition, when the version of the target application is upgraded, the above solution can also quickly adapt to the new version of the target application.

[0065] In accordance with the above embodiments, the present application also provides a script execution method, which can be applied to the target application. Figure 3 As shown, the method comprises the following steps:

[0066] Step 31, obtain the target script; the target script is generated according to the aforementioned script generation method.

[0067] Step 32: parse the target script to obtain description information, execution order, and respective target controls of multiple interactive operations for the target task of the target application.

[0068] Step 33: According to the description information of the multiple interactive operations, the multiple interactive operations are sequentially performed on the target controls targeted by the multiple interactive operations in the execution order to complete the target task. According to the instruction description information in the description information of the multiple interactive operations, the multiple interactive operations are sequentially performed on the target controls targeted by the multiple interactive operations in the execution order to complete the target task.

[0069] In this way, the target task can be executed more efficiently for the target application based on the target script generated by the script generation method.

[0070] It should be noted that the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 11 to 14 can be device A; for another example, the execution subject of steps 11 to 12 can be device A, and the execution subject of steps 13 to 14 can be device B; and so on.

[0071] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel, and the sequence numbers of the operations, such as 12, 13, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.

[0072] It should be noted that the descriptions such as “first” and “second” in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit “first” and “second” to different types.

[0073] Figure 4 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application, and the electronic device is applicable to the script generation method provided by the aforementioned embodiment, such as Figure 4 As shown, the electronic device may include: a memory 401 , a processor 402 , and a communication component 403 .

[0074] The memory 401 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0075] In some exemplary embodiments, the processor 402 is coupled to the memory 401 and is used to execute a computer program in the memory 401 to: obtain description information of multiple interactive operations for a target task of a target application, and, based on the description information of the multiple interactive operations, control the target application to simulate multiple initial pages generated by executing the multiple interactive operations; take screenshots of the multiple initial pages to obtain page screenshots corresponding to each of the multiple interactive operations; input the description information of the multiple interactive operations and the page screenshots corresponding to each of the multiple interactive operations into a control recognition model in sequence according to the execution order of the multiple interactive operations, so as to identify a target control targeted by any interactive operation from multiple controls in the page screenshot corresponding to any interactive operation according to the action description information in the description information of any interactive operation; generate a target script according to the description information of the multiple interactive operations, the execution order, and the target controls targeted by each of the multiple interactive operations; the target script is used to: control the target application to perform the multiple interactive operations on the target controls targeted by each of the multiple interactive operations in sequence according to the execution order according to the description information of the multiple interactive operations to complete the target task.

[0076] Optionally, when the processor 402 obtains the description information of multiple interactive operations for the target task of the target application, it is specifically used to: provide the user with a task configuration interface of the target application so that the user can configure the target task; in response to the user's configuration operation, determine the description information of the multiple interactive operations of the target task; or, use the historical description information of multiple historical interactive operations in the historical target task as the description information of the multiple interactive operations; the historical target task is a task performed for a historical version of the target application.

[0077] Optionally, the action description information of any interactive operation is used to describe the operation type of any interactive operation, including: action description text and / or action description icon; when the processor 402 identifies the target control targeted by any interactive operation from the multiple controls in the page screenshot corresponding to the any interactive operation based on the action description information in the description information of any interactive operation, it is specifically used to: extract the first feature information corresponding to the action description information of any interactive operation, and the second feature information corresponding to each of the multiple screenshot areas of the page screenshot corresponding to the any interactive operation; the multiple screenshot areas respectively display the multiple controls; the second feature information corresponding to any screenshot area is used to characterize the control displayed in the screenshot area; calculate the feature similarity between the first feature information and the second feature information corresponding to each of the multiple screenshot areas; use the screenshot area whose feature similarity among the multiple screenshot areas meets the preset conditions as the target area; use the control displayed in the target area as the target control targeted by any interactive operation.

[0078] Optionally, the control recognition model pre-learns the region segmentation capability for different page screenshots; when the processor 402 extracts the second feature information corresponding to each of the multiple screenshot regions of the page screenshot corresponding to any interactive operation, it is specifically used to: utilize the region segmentation capability to segment the page screenshot corresponding to any interactive operation into the multiple screenshot regions corresponding to the multiple controls; and extract the second feature information corresponding to each of the multiple screenshot regions.

[0079] Optionally, when the processor 402 generates a target script based on the description information of the multiple interactive operations, the execution order, and the target controls for which the multiple interactive operations are respectively targeted, it is specifically used to: generate multiple operation instructions based on the description information of the multiple interactive operations and the target controls for which the multiple interactive operations are respectively targeted, any operation instruction being used to control the target application to perform a corresponding interactive operation on any target control; and integrate the multiple operation instructions into the target script based on the execution order.

[0080] Optionally, the processor 402 is also used to: obtain a first sample page screenshot, sample action description information of a sample interactive operation, and a second sample page screenshot marked with a sample control; input the first sample page screenshot and the sample action description information into a preset control recognition network, and under the supervision of the second sample page screenshot, train the control recognition network using the first sample page screenshot and the sample action description information with the goal of converging the loss function of the control recognition network to a target range to obtain the control recognition model; wherein the loss function is used to calculate the error between the candidate control targeted by the sample interactive operation in the sample page screenshot identified by the control recognition network and the sample control.

[0081] In other exemplary embodiments, the processor 402 is coupled to the memory 401, and is used to execute a computer program in the memory 401, so as to: obtain a target script; the target script is generated according to the script generation method; the target script is parsed to obtain description information, execution order and target controls of multiple interactive operations for the target task of the target application; according to the description information of the multiple interactive operations, the multiple interactive operations are executed in sequence on the target controls respectively targeted by the multiple interactive operations in accordance with the execution order to complete the target task.

[0082] Further, if Figure 4 As shown, the electronic device also includes: a display 404, a power component 405, an audio component 406 and other components. Figure 4 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 4Components shown.

[0083] An embodiment of the present application also provides a computer-readable storage medium, which, when the computer program is executed by a processor, enables the processor to implement the steps in the script generation method or the script execution method.

[0084] An embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps of any one of the script generation method and the script execution method.

[0085] In this embodiment, description information of multiple interactive operations and multiple pages generated by simulating the execution of multiple interactive operations by controlling the target application based on the description information of multiple interactive operations can be obtained; screenshots of multiple pages are taken to obtain page screenshots corresponding to multiple interactive operations; the description information of multiple interactive operations and the page screenshots corresponding to multiple interactive operations are input into the control recognition model according to the execution order of multiple interactive operations, so as to identify the target control targeted by any interactive operation from the multiple controls in the page screenshot corresponding to any interactive operation according to the action description information in the description information of any interactive operation; and the target script is generated according to the description information of multiple interactive operations, the execution order and the target control targeted by multiple interactive operations. In this way, the automation script can be generated more efficiently and accurately without relying on the page control tree.

[0086] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0087] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.

[0088] The above-mentioned display includes a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0089] The power supply assembly provides power to various components of the device where the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply assembly is located.

[0090] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal.

[0091] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-readable storage media (including but not limited to disk storage, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage, etc.) containing computer-usable program code.

[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0093] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0095] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interface, network interface and memory.

[0096] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0097] Computer readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0098] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0099] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A script generation method, characterized in that: include: Acquire description information of a plurality of interactive operations for a target task of a target application, and control the target application to simulate executing a plurality of initial pages generated by the plurality of interactive operations based on the description information of the plurality of interactive operations; Taking screenshots of the multiple initial pages to obtain page screenshots corresponding to the multiple interactive operations; In accordance with the execution order of the multiple interactive operations, the description information of the multiple interactive operations and the page screenshots corresponding to the multiple interactive operations are sequentially input into the control recognition model, so as to identify the target control targeted by any interactive operation from the multiple controls in the page screenshot corresponding to the any interactive operation according to the action description information in the description information of any interactive operation; Generate a target script according to the description information of the multiple interactive operations, the execution order, and the target controls targeted by the multiple interactive operations; The target script is used to: control the target application to perform the multiple interactive operations on the target controls targeted by the multiple interactive operations in sequence according to the execution order according to the description information of the multiple interactive operations to complete the target task.

2. The method according to claim 1, characterized in that Get description information of multiple interactive operations for the target task of the target application, including: Providing a task configuration interface of the target application to the user so that the user can configure the target task; in response to the configuration operation of the user, determining description information of multiple interactive operations of the target task; or, The historical description information of the multiple historical interactive operations in the historical target task is used as the description information of the multiple interactive operations; the historical target task is a task executed for the target application of the historical version.

3. The method according to claim 1, characterized in that The action description information of any interactive operation is used to describe the operation type of any interactive operation, including: an action description text and / or an action description icon; According to the action description information in the description information of any interactive operation, identifying a target control targeted by the any interactive operation from a plurality of controls in the page screenshot corresponding to the any interactive operation includes: Extracting first characteristic information corresponding to the action description information of any interactive operation, and second characteristic information corresponding to each of multiple screenshot areas of the page screenshot corresponding to any interactive operation; the multiple screenshot areas respectively display the multiple controls; the second characteristic information corresponding to any screenshot area is used to characterize the control displayed in the screenshot area; Calculating feature similarity between the first feature information and second feature information corresponding to each of the plurality of screenshot areas; The screenshot regions whose feature similarities meet the preset conditions among the multiple screenshot regions are taken as target regions; The control displayed in the target area is used as the target control for any interactive operation.

4. The method according to claim 3, characterized in that The control recognition model pre-learns the region segmentation capability for different page screenshots; extracting second feature information corresponding to each of the plurality of screenshot regions of the page screenshot corresponding to any interactive operation, including: The page screenshot corresponding to any one of the interactive operations is segmented into the multiple screenshot regions corresponding to the multiple controls by utilizing the region segmentation capability; and the second feature information corresponding to each of the multiple screenshot regions is extracted.

5. The method according to any one of claims 1 to 4, characterized in that: Generating a target script according to the description information of the multiple interactive operations, the execution order, and the target controls targeted by the multiple interactive operations respectively, including: Generate multiple operation instructions according to the description information of the multiple interactive operations and the target controls targeted by the multiple interactive operations, any one of the operation instructions is used to control the target application to perform a corresponding interactive operation on any one of the target controls; According to the execution order, the multiple operation instructions are integrated into the target script.

6. The method according to any one of claims 1 to 4, characterized in that: Also includes: Obtain a first sample page screenshot, sample action description information of a sample interactive operation, and a second sample page screenshot annotated with a sample control; Inputting the first sample page screenshot and the sample action description information into a preset control recognition network, and under the supervision of the second sample page screenshot, training the control recognition network with the first sample page screenshot and the sample action description information with the goal of converging the loss function of the control recognition network to a target range, to obtain the control recognition model; The loss function is used to calculate the error between the candidate control targeted by the sample interaction operation identified by the control recognition network in the sample page screenshot and the sample control.

7. A script execution method, characterized in that: Suitable for target applications including: Obtain a target script; the target script is generated according to the method according to any one of claims 1-6; Parsing the target script to obtain description information, execution order, and respective target controls of multiple interactive operations for the target task of the target application; According to the description information of the multiple interactive operations, the multiple interactive operations are sequentially performed on the target controls targeted by the multiple interactive operations in the execution order to complete the target task.

8. An electronic device, characterized in that: include: A memory and a processor; wherein the memory is used to: store one or more computer instructions; and the processor is used to execute the one or more computer instructions to: execute the steps in the method described in any one of claims 1-6 or claim 7.

9. A computer-readable storage medium, characterized in that: When the computer program is executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 6 or claim 7.

10. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, enables the processor to implement the steps of the method according to any one of claims 1 to 6 or claim 7.