UI Automation Script Editing System and Method Based on Screenshot Recognition
Through the UI automated script editing system based on screenshot recognition, using XML format files and Base64 encoding, the problems of long development cycle, poor portability and cross-platform operation of the automated script generation system in the existing technology are solved, and efficient, flexible and cross-platform secondary writing of scripts is achieved.
Patent Information
- Application Number
- CN202510574426.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing automated script generation system has a long development cycle, high technical requirements, poor portability, and cannot flexibly handle software interface changes. The recording method is prone to running errors and cannot operate across platforms. The AI algorithm generates scripts requires a large amount of computing resources.
The UI automated script editing system based on screenshot recognition is adopted, script information is stored through XML format files, import and export functions are added, and screenshot image recognition software windows and operation information are used, and the Base64 data format is used for encoding to achieve cross-platform operations.
It improves the availability and convenience of scripts, can operate across platforms, simplifies the secondary writing of scripts, avoids absolute position limitations of recording methods, and enhances the operation capabilities of automated scripts.
Smart Images

Figure CN120085969B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of script generation, and particularly relates to a UI automation script editing system and method based on screenshot recognition. Background Art
[0002] In the current era of Internet informatization, the usage rate of computer software has doubled, and various software is being used in different fields to empower industries. However, the operation of software requires manual support. If relying solely on manual labor, the labor cost will be greatly increased. Moreover, the development of current software-based automation script methods is difficult. Therefore, developers need to write a large amount of development code according to project requirements. In addition, since most of the automation scripts are based on Slelenium, the portability is poor. Once the code has problems, modifying it will consume a lot of energy to first understand the code structure. It is very difficult for other developers to continue rewriting the previously developed automation scripts, and often need to develop separately again. In addition, there will be very complicated management and collaboration problems after the script development is completed, and the code needs to be continuously tested. Therefore, there is an urgent need for a simple and easy-to-use automation script editing method to help software operators meet the software operation requirements in different scenarios.
[0003] Existing automated script generation systems can monitor software's execution of work instructions and help people handle related software tasks well. However, the development cycle of automated scripts is long, the technical requirements are high, the portability is poor, and there are often problems such as being unable to re-edit other people's script codes. Moreover, current automated scripts are mostly developed based on the recording function to record mouse and keyboard functions, but they cannot record the icon information of the software. When the relative position of the software changes, it is extremely easy to cause running errors. For example, the invention patent application with the Chinese patent publication number CN115509910A, publication date of September 26, 2022, and patent name of "An Automated Script Generation Method and Tool Based on Recording Technology" proposed an automated script generation method based on recording. Although this method can generate automated scripts by recording the operation process, it is very limited to the overall recording process and lacks interaction with the software interface. Once the icon or window position of the software changes, it cannot be used again, and it cannot handle special situations such as prompt boxes or error boxes. At the same time, when dealing with operations that require long waiting times, the recording time cycle is long, and it does not have a secondary repeated editing function and needs to be re-recorded; the invention patent application with the Chinese patent publication number CN116340152A, publication date of June 27, 2023, and patent name of "UI Automated Script Generation Tool" improved the problem of inability to interact with the interface window in the above patent, but it is still limited to recording to generate scripts, so the production of scripts is still not flexible enough, time-consuming and inefficient; the invention patent application with the Chinese patent publication number CN119127721A, publication date of December 13, 2024, and patent name of "An AI-Algorithm-Based Selenium Automated Script Generation Method and System" uses AI algorithms to generate automated scripts, which is very novel. However, the training of the model requires a large amount of computing resources and requires relevant AI technologies for training as support, which is completely unnecessary for automated scripts with general low complexity. Summary of the Invention
[0004] In view of this, the present invention aims to provide a UI automated script editing system and method based on screenshot recognition. By storing script information in an XML format file and adding import and export functions on this basis, the script can be very conveniently re-edited, thereby improving the utilization rate of the script; the present invention also uses the method of intercepting relevant screenshot images, enabling the automated script to find information in the screenshot image rather than mouse information, improving the operation ability of the automated script.
[0005] To achieve the above object, the technical solution of the present invention is realized as follows:
[0006] A UI automated script editing system based on screenshot recognition, comprising:
[0007] A software window acquisition module, which is used to capture the current display desktop and obtain the window names of the software currently opened on the display desktop from the captured screenshot image;
[0008] An instruction information editing module, which receives the screenshot image and window names, as well as the operation information on the software input according to the screenshot image, and stores the operation information and the screenshot image information of the screenshot image in an internal Task dictionary;
[0009] A storage information editing module, which encodes the screenshot image information in the internal Task dictionary in the Base64 data format to obtain the corresponding screenshot image data, and jointly exports the screenshot image data and other information in the internal Task dictionary as an XML script file; and imports an external XML script file, decodes the external screenshot image data in the external XML script file in the Base64 data format to obtain the corresponding external screenshot image information, imports the external screenshot image information and other information in the external XML script file into the current internal Task dictionary, and after replacing the content of the current internal Task dictionary, delivers the current internal Task dictionary to the instruction information editing module for secondary editing;
[0010] A user interface display module, which is used to receive the operation information, window names and image information from the instruction information editing module.
[0011] Furthermore, the software window acquisition module includes:
[0012] A screenshot sub-module, which is used to capture the current display desktop and store the screenshot image;
[0013] A program window name acquisition sub-module, which receives the screenshot image, identifies the window names of all the software currently opened on the display desktop from the screenshot image, and loads all the software window names into the user interface display module for display.
[0014] Furthermore, the instruction information editing module includes:
[0015] An interface information sub-module, which obtains the interface information of the program in the currently captured screenshot image and displays the interface information in the user interface display module;
[0016] An operation information sub-module, which retrieves the interface information of the corresponding program, inputs the operation information of the corresponding program, and completes the operation on the corresponding program;
[0017] A double-layer information sub-module, which uses the operation information as the first-layer information, uses the image information of the current screenshot image as the second-layer information, and stores the two layers of information through an internal Task dictionary.
[0018] Further, the operation information includes: controlling the pointer to click on different positions of the partial screenshot in the interface, controlling the number of clicks and moving the pointer, and controlling the pointer to drag the partial screenshot, as well as typing operation instructions for the corresponding software and inputting text information into the corresponding software.
[0019] Further, the operation information also includes loop operation information, monitoring operation information, and input sequence number information; where:
[0020] The loop operation information includes the loop operation of operation instructions and the loop input of text information.
[0021] The monitoring operation information adopts the method of end monitoring to monitor whether the current software is operating normally.
[0022] The input sequence number information is to input numbers arranged in a certain order according to the loop order.
[0023] Further, the instruction information editing module also includes:
[0024] The email information sub-module is used to obtain the email information of the external XML script file, confirm the email information of the XML script file, and store the email information in the Task dictionary.
[0025] The warning information sub-module is used to generate warning information. The warning information includes the processing operations for error information. The processing operations include: indexing to the interface with error information according to the interface information of the error information, performing corresponding operation information on the interface containing the error information, and storing the corresponding operation information in the Task dictionary.
[0026] Further, the storage information editing module includes:
[0027] The export operation sub-module opens the screenshot image and the Task dictionary according to the parent class and subclass: if the subclass is a single element, add it directly; if the subclass is the Task dictionary, continue the inner loop; if the subclass is a list, open the subclass of the list and add all elements in the list; if the subclass is a screenshot image, save the screenshot image as a byte stream, and then encode the byte stream into the Base64 data format encoding to obtain the corresponding screenshot image data; export the screenshot image data and the Task dictionary as an XML script file.
[0028] The import operation sub-module receives the external XML script file, exports the external Task dictionary and external screenshot image data from the external XML script file, decodes the external screenshot image data in the Base64 data format to obtain the corresponding external screenshot image, and delivers the external screenshot image and the external Task dictionary to the instruction information editing module.
[0029] A UI automation script editing method based on screenshot recognition. According to the UI automation script editing system based on screenshot recognition provided by the present invention, it includes:
[0030] S1: Whether to import an external XML script file. If so, control the storage information editing module to read and decode the external XML script file, add and replace the content of the internal Task dictionary with the obtained data information, and then execute step S2; otherwise, directly proceed to step S3;
[0031] S2: Determine whether to perform secondary editing on the operation information and screenshot image information in the internal Task dictionary obtained in step S1. If so, proceed to step S3; otherwise, execute step S4;
[0032] S3: Control the software window acquisition module to receive the internal Task dictionary in step S2, and based on the screenshot image and window name in the internal Task dictionary, add the screenshot image and window name of the current desktop; or control the software window acquisition module to directly obtain the screenshot image and window name of the current desktop;
[0033] S4: Control the instruction information editing module to directly receive and modify the operation information in the internal Task dictionary in step S2; or control the instruction information editing module to modify the operation information in the internal Task dictionary according to the screenshot image and window name in step S3; store the current operation information and screenshot image information as a new internal Task dictionary;
[0034] S5: Control the storage information editing module to encode the screenshot image information in the new internal Task dictionary obtained in step S4 into Base64 data format to obtain the corresponding screenshot image data; export the new internal Task dictionary with the screenshot image data as an XML script file.
[0035] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0036] (1) In the UI automation script editing system and method based on screenshot recognition of the present invention, a method based on image search is adopted instead of the recording method, effectively preventing the absolute limitation of the recording method on the position of window content; at the same time, the present invention does not require additional pre-training, and script development can be achieved through the user-interactive window, with the characteristics of simplicity and high efficiency; in addition, the present invention stores script information with XML format files, and on this basis, import and export functions are added, enabling the script to be very conveniently rewritten on the UI interface, thereby improving the availability and convenience of the script. The present invention also uses the method of intercepting relevant screenshot images, enabling the automation script to find the information screenshot image instead of the mouse information, improving the operation ability of the automation script;
[0037] (2) In the UI automation script editing system and method based on screenshot recognition of the present invention, Base64 data format is used for encoding and decoding. Base64 encoding has good cross-platform reliability. If general ASCII code or UTF-8 is used for encoding, the encoding results are different on different platforms, resulting in inability to perform cross-platform operations. While Base64 has very good stability and is suitable for cross-platform operations. In addition, since the system provided by the present invention needs to save screenshots as data in XML script files, if it becomes binary data during this process, it will damage the text protocol of XML and cannot be stored. And Base64 can avoid such text parsing errors during the data conversion process. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0039] Figure 1 It is a schematic structural diagram of the UI automation script editing system based on screenshot recognition described in the embodiment of the present invention;
[0040] Figure 2 It is a UI interface diagram corresponding to the user interface display module described in the embodiment of the present invention;
[0041] Figure 3 It is a schematic diagram of the program window name acquisition sub-module described in the embodiment of the present invention;
[0042] Figure 4 It is a schematic diagram of the XML script file described in the embodiment of the present invention;
[0043] Figure 5 It is a schematic flow diagram of the UI automation script editing method based on screenshot recognition described in the embodiment of the present invention;
[0044] Figure 6 It is a flow chart of the UI automation script editing method based on screenshot recognition described in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0046] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0047] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0048] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0049] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0050] Such as Figure 1As shown in the figure, the UI automation script editing system based on screenshot recognition described in the embodiments of the present invention includes a software window acquisition module, an instruction information editing module, a storage information editing module, and a user interface display module. Among them, the software window acquisition module is used to take a screenshot of the current display desktop and obtain the window names of the software currently opened on the display desktop from the captured screenshot image. The instruction information editing module receives the screenshot image and the window name, and inputs the operation information of the software according to the screenshot image, and stores the operation information and the screenshot image information of the screenshot image in the internal Task dictionary. The storage information editing module encodes the screenshot image information in the internal Task dictionary in the Base64 data format to obtain the corresponding screenshot image data, and jointly exports the screenshot image data and other information in the internal Task dictionary as an XML script file; and imports an external XML script file, decodes the external screenshot image data in the external XML script file in the Base64 data format to obtain the corresponding external screenshot image information, imports the external screenshot image information and other information in the external XML script file into the current internal Task dictionary of the system provided by the present invention, and after replacing the content of the current internal Task dictionary, delivers the current internal Task dictionary to the instruction information editing module for secondary editing. The user interface display module is used to receive the operation information, window name, and image information from the instruction information editing module. The UI interface corresponding to the user interface display module is as Figure 2 shown.
[0051] In some embodiments, the software window acquisition module includes a screenshot sub-module and a program window name acquisition sub-module. Among them, the screenshot sub-module is used to take a screenshot of the current display desktop and store the screenshot image. The program window name acquisition sub-module receives the screenshot image, identifies the window names of all the software currently opened on the display desktop from the screenshot image (as Figure 3 shown), and loads all the software window names into the user interface display module for display.
[0052] In a certain embodiment, the screenshot sub-module uses the pyautogui.screenshot() function to take a screenshot of the entire screen, and by creating a new window on the root, then creating a new canvas, waiting for the user to make a selection on the canvas, and automatically saving the screenshot to the default path in the upper text box, generally the desktop. The program window name acquisition sub-module receives the screenshot image and uses the pygetwindow.getWindowsWithTitle() function to load and display the window names of all the currently opened programs line by line into the text box of the user interface display module. There may be blank lines between each program window. If no program is opened, the message “No windows found.” will appear.
[0053] In some embodiments, the instruction information editing module obtains and stores the operation information and screenshot image information of the software in the way of double-layer information. Specifically, the instruction information editing module includes an interface information sub-module, an operation information sub-module, and a double-layer information sub-module. Among them,
[0054] The interface information sub-module obtains the interface information of the program in the currently captured screenshot image and displays the interface information in the user interface display module. The operation information sub-module retrieves the interface information of the corresponding program, inputs the operation information of the corresponding program, and completes the operation of the corresponding program. Among them, the operation information includes: controlling the pointer to click on different positions of the local screenshot in the interface, controlling the number of clicks and moving the pointer, and controlling the pointer to drag the local screenshot in the interface, as well as typing the operation instructions for the corresponding software and inputting the text information to the corresponding software. The double-layer information sub-module stores the operation information as the first layer of information and the screenshot image information as the second layer of information through an internal Task dictionary.
[0055] In a certain embodiment, the interface information sub-module obtains the interface window name of the program in the current screenshot image, numbers the interface window name in the order of 1, 2, 3....., and loads it into the user interface display module for display, which is convenient for subsequent calling of the interface.
[0056] In a certain embodiment, the format of the screenshot image information is PIL image format, and the operation information is input into the operation information sub-module in the form of instructions. The operation information specifically includes:
[0057] The operation information for controlling the pointer to click on the local screenshot image in the interface is specifically: controlling the pointer to click on the local screenshot image in the interface. In this embodiment, specifically, the mouse is manipulated to click according to the set click orientation. The click orientations include east, south, west, north, northwest, northeast, southwest, and southeast. The click orientations and the corresponding click orientation symbols are shown in Table 1.
[0058] Table 1:
[0059]
[0060] The instruction format for clicking on the local screenshot image in the interface is: interface x - local screenshot - click orientation symbol - d, where x represents the interface serial number and d represents the double-click marker. If the click orientation symbol is not specifically added, it is defaulted to the center position. For the sake of convenience, in this embodiment, a single click is defaulted at the click orientation, and there is no double-click marker d in the corresponding instruction format, that is, the instruction format for a single click is interface x - local screenshot - click orientation symbol. Since this embodiment can locate a target point through eight orientations, the surrounding features will be more prominent and suitable for positioning.
[0061] Operation information for controlling the number of clicks of a pointer on a screenshot image in the interface. Specifically: Manipulate the pointer to continuously click on a partial screenshot image in the interface according to the set number of clicks. In this embodiment, specifically manipulate the mouse to continuously click according to the set number of clicks. The click positions of the mouse are divided into the left button, middle button, and right button, and the corresponding click symbols are left, middle, and right. The instruction format for the corresponding number of clicks is: m-left / middle / right - number of clicks, where m represents the identifier for determining the number of clicks of the pointer.
[0062] Operation information for controlling the movement of a pointer. Specifically: Manipulate the pointer to continuously move according to the set number of movements. In this embodiment, specifically manipulate the mouse to continuously move according to the set number of movements. The instruction format is: s-(relative movement distance in the x direction, relative movement distance in the y direction) - number of movements, where s represents the identifier for moving the pointer.
[0063] Operation information for controlling the dragging of a pointer on a partial screenshot in the interface. Specifically: Manipulate the pointer to drag the screenshot in the interface according to the set dragging direction. In this embodiment, specifically manipulate the mouse to drag the screenshot according to the set dragging direction. The dragging direction points from the pre-dragging position to the post-dragging position, and the corresponding instruction format is: interface x - screenshot - (xx, yy), where xx represents the symbol of the pre-dragging position and yy represents the symbol of the post-dragging position. The dragging positions also include eight different directions such as east, south, west, north, northwest, northeast, southwest, and southeast. The dragging directions and their corresponding symbols are shown in Table 2.
[0064] Table 2:
[0065]
[0066] Operation information for entering operation instructions for corresponding software. Specifically: Determine the operation instructions for the corresponding software by typing. In this embodiment, the tool for typing the operation instructions is the keyboard, and the corresponding operation information is to manipulate the keys on the keyboard to enter the operation instructions for the corresponding software, and thus, realize the operation of the corresponding software. The instruction format is: k - keyboard key content - number of operations, where k represents the identifier for entering the operation instructions. The keyboard key content and the corresponding operations are shown in Table 3.
[0067] Table 3:
[0068]
[0069] Operation information of the text information input to the corresponding software, specifically, manipulating the typing device to input text information. In this embodiment, specifically, manipulating the keyboard to achieve text input. However, in order to be able to input Chinese symbols, the operation here is essentially a copy-and-paste operation. This text input method is different from the automatic control method of the keyboard and avoids unnecessary character switching. The instruction format of the text information input to the corresponding software is: w - text content, where w represents the identifier for performing the text input operation.
[0070] In some embodiments, the operation information further includes loop operation information, monitoring operation information, and input sequence number information. Among them, the loop operation information includes the loop operation of the operation instruction and the loop input of the text information; the monitoring operation information adopts the end monitoring method to monitor whether the current software operates normally; the input sequence number information is to input numbers arranged in a certain order according to the loop order and perform related operations according to the numbers.
[0071] In a certain embodiment, the loop operation information realizes regular control of the operation instruction and text input. If the number of operations exceeds the given number of operations or text quantity, the loop operation or text input will be performed. In this embodiment, the instruction format corresponding to the loop operation is c - (operation times 1, operation times 2,..., operation times N) - operation instruction, and the instruction format for loop input of text is c - [text 1, text 2,.....]. Combining Figure 2 , when the software receiving the instruction reads the instruction for loop operation or loop input of text, it will recognize the prefix "c", indicating that a loop operation or loop input of text operation is to be performed, and then perform the corresponding number of operations on the operation instruction during each loop. The instruction format corresponding to the operation information of inputting sequence numbers is n - (aa, bb) - a, where n represents the identifier for the operation of inputting sequence numbers, (aa, bb) represents the array from number aa to number bb, and a represents the identifier for the select-all operation.
[0072] The monitoring operation information is used to monitor whether the current software is operating normally. Since the cost of manual monitoring is very high, adopting end - point monitoring can effectively reduce the monitoring cost of software operation. The monitoring operation information specifically is that the operating software runs until it reaches the end state. At this time, the end - interface of the software is obtained. When software - operation monitoring is required, after the system running the software receives the monitoring operation information, it retrieves the end - interface of the software and makes a real - time similarity comparison between the partial screenshot in the end - interface of the software and the partial picture in the current interface of the software. When the similarity reaches the preset threshold, the system running the software determines that the software has ended its operation and the software is working properly; if the similarity never reaches the preset threshold until the software ends its operation, the system determines that the software is not working properly. In this embodiment, the corresponding instruction format of the monitoring operation information is interface number - partial screenshot - z, where z represents the identifier of the end - interface, and the preset threshold is set at 95%. Exemplarily, when the partial screenshot in the end - interface of the software is that the progress bar reaches 100%, when monitoring the software operation, the partial picture where the progress bar is located in the current interface is extracted in real - time, and the similarity between the partial picture where the progress bar is located in the current interface and the partial screenshot of the progress bar reaching 100% is compared in real - time. When the similarity is above 95%, the system running the software will automatically determine that the software has ended its operation and the software is working properly.
[0073] It should be noted that in the embodiments of the present invention, for various operation information, the instruction format provided by the present invention adopts the form of adding key - instruction prefixes and suffixes. Exemplarily, for the operation information of controlling the pointer to click on interface 2 in the screenshot image cipan.png, the corresponding instruction format is: "2 - cipan.png - w - d". When the software receiving the instruction reads this instruction, it will first locate to interface 2, and then double - click on the position in the west of the screenshot cipan.png. Here, "2" is the interface number, which is the prefix, and the following "d" represents double - click, which is the suffix. The prefix and suffix reflect all the key words in the instruction information for the screenshot. By adding the method of key - instruction prefixes and suffixes, the present invention avoids the cumbersome recording process, can implement the relevant instruction functions, greatly reduces the data volume of the script information, and at the same time improves the understanding speed of the script execution system for the instructions.
[0074] In some embodiments, the instruction information editing module further includes an email information sub-module and a warning information sub-module. Among them, the email information sub-module is used to obtain the email information of the external XML script file, as well as confirm the email information of the exported XML script file, and store the email information in the internal Task dictionary. The warning information sub-module is used to generate warning information, and the warning information includes processing operations for error information. The processing operations include: indexing to the interface with error information according to the interface information of the error information, and the corresponding operation information for the interface containing the error information, and storing the corresponding operation information in the internal Task dictionary.
[0075] In a certain embodiment, in the email information sub-module, the email information of the XML script file generated by the system provided by the present invention, as well as the email information of the external XML script file, both have a format of four lines, including:
[0076] The first line represents the sender's email; the format is like: 1111@uu.com;
[0077] The second line represents the recipient's email; the format is like: 2222@pp.com;
[0078] The third line is default as the title of the email; the format is like: Hello,this is test!!;
[0079] The fourth line is default as the appendix of the email; the format is like: C: / Users / Administrator / Desktop / tttt.png.
[0080] If the email information is empty or only the first line is filled, the system provided by the present invention will not send emails after the operation ends; if the first line and the second line are filled in the email information, the system provided by the present invention will send an email from the sender to the recipient, and the email title is the default content; if the first line, the second line and the third line are filled in the email information, the system provided by the present invention will send an email from the sender to the recipient, and the email title is the specified content, and the appendix will not be sent; if the first line, the second line, the third line and the fourth line are filled in the email information, the system provided by the present invention will send an email from the sender to the recipient, and the email title is the specified content, with the specified appendix attached.
[0081] The warning information sub-module is for when error information needs to be processed. First, the interface information with problems needs to be added, and then the interface with error information is indexed according to the serial number to perform relevant operations. The warning information sub-module also needs to obtain the interface window name through the sub-module that obtains the program window name. The instruction format is as follows: interface serial number - screenshot - d, where the interface serial number is still the prefix, and the "d" behind is still the suffix.
[0082] In some embodiments, the storage information editing module includes an export operation sub-module and an import operation sub-module. Among them, the export operation sub-module opens the internal Task dictionary according to the parent class and sub-classes. Specifically, the internal Task dictionary is the parent class, and correspondingly, the information in the internal Task dictionary is saved in the form of sub-classes in the internal Task dictionary. The forms of sub-classes include second-level sub-classes, single elements, and lists. For the sub-classes in the form of lists in the internal Task dictionary, open the list and obtain all the elements in the list; for the single elements in the internal Task dictionary, directly obtain the single element; for the second-level sub-classes in the internal Task dictionary, repeat the above sub-class - parent class operation until all the elements in the second-level sub-classes are obtained. Finally, after compiling all the obtained elements using the lxml function, return them in the form of a string, store them in a formatted manner, and export them as an XML script file as shown in Figure 4 the following. It should be noted that for the screenshot image information in the internal Task dictionary, the screenshot image information in the PIL image format needs to be saved as a byte stream first, then the byte stream is Base64 encoded, and then the above sub-class - parent class operation is performed on the encoded screenshot image data.
[0083] In a certain embodiment, in the export operation sub-module, special characters in the stored dictionary information, such as the five symbols <, >, &, ’, and ", are respectively converted into symbols that XML can handle, namely <, >, &, &apos, and ". The spaces in front of the stored email information, interface information, warning information, and operation information are converted into underscores "_" for easy processing. Since the picture name cannot be loaded due to the presence of spaces, there is no need to replace the spaces.
[0084] In some embodiments, the import operation sub-module receives an external XML script file, imports the data in the external XML script file into the current internal Task dictionary of the system, and replaces the content of the current internal Task dictionary. Specifically, for the imported single element, directly import the single element into the current internal Task dictionary; for the imported list, open the list and import all the elements in the list into the current internal Task dictionary. Create a list of information for the strings in the external XML script file, but not the encoded screenshot image data; for the strings in the external XML script file that are encoded screenshot image data, perform Base64 decoding on the screenshot image data, and then create a list of information. Store and replace the content of the current internal Task dictionary with the data obtained above to implement the import operation of the external script information.
[0085] In one embodiment, in the import operation sub-module, an external XML script file is read, the number of stored tasks is confirmed, and then conversion is performed according to the task serial number; the formats saved in the XML script file such as <, >, &, &apos, " are converted into special characters in the dictionary information such as <, >, &, ’, "; the underlined email information, interface information, and information in the external XML script file are restored to the original format with spaces.
[0086] The present invention also provides a UI automation script editing method based on screenshot recognition. According to the UI automation script editing system provided by the present invention, combined with Figures 1 to 6 , the provided method includes:
[0087] S1: Whether to import an external XML script file: If so, control the storage information editing module to read and decode the external XML script file, add and replace the content of the internal Task dictionary with the obtained data information, and then execute step S2; otherwise, directly proceed to step S3;
[0088] S2: Determine whether to perform secondary editing on the operation information and screenshot image information in the internal Task dictionary obtained in step S1; if so, proceed to step S3; otherwise, execute step S4;
[0089] S3: Control the software window acquisition module to receive the internal Task dictionary in step S2, and based on the screenshot image and window name in the internal Task dictionary, add the screenshot image and window name of the current desktop; or control the software window acquisition module to directly obtain the screenshot image and window name of the current desktop;
[0090] S4: Control the instruction information editing module to directly receive and modify the operation information in the internal Task dictionary in step S2; or control the instruction information editing module to modify the operation information in the internal Task dictionary according to the screenshot image and window name in step S3; store the current operation information and screenshot image information as a new internal Task dictionary;
[0091] S5: Control the storage information editing module to encode the screenshot image information in the new internal Task dictionary obtained in step S4 into the Base64 data format to obtain the corresponding screenshot image data; export the new internal Task dictionary with the screenshot image data as an XML script file.
[0092] It should be noted that the modification operation on the operation information in the internal Task dictionary in step S4 includes changing the existing operation information in the internal Task dictionary and / or adding new operation information. In addition, since the instruction information editing module stores the operation information and the screenshot image information in an internal Task dictionary in a two-layer information manner, when step S1 does not import the external XML script file, it indicates that a new XML script file needs to be created at this time. At this time, the control software window acquisition module in step S3 is directly executed to directly obtain the screenshot image and window name of the current desktop. In this case, the modification operation on the operation information in the internal Task dictionary in step S4 can be understood as adding the corresponding operation information to the current internal Task dictionary according to the screenshot image and window name directly obtained in step S3.
[0093] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.
[0094] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A UI automation script editing system based on screenshot recognition, characterized in that Including: A software window acquisition module, which is used to capture a screenshot of the current display desktop and obtain the window names of the software currently opened on the display desktop from the captured screenshot image; An instruction information editing module, which receives the screenshot image and the window name, and inputs operation information on the software according to the screenshot image, and stores the operation information and the screenshot image information of the screenshot image in an internal Task dictionary; The instruction information editing module includes: an interface information sub-module, which obtains the interface information of the program from the currently captured screenshot image and displays the interface information in the user interface display module; an operation information sub-module, which retrieves the interface information of the corresponding program and inputs the operation information of the corresponding program to complete the operation of the corresponding program; a double-layer information sub-module, which uses the operation information as the first-layer information and the image information of the current screenshot image as the second-layer information, and stores the two layers of information in an internal Task dictionary; A storage information editing module, which encodes the screenshot image information in the internal Task dictionary in the Base64 data format to obtain the corresponding screenshot image data, and exports the screenshot image data and other information in the internal Task dictionary as an XML script file; and imports an external XML script file, decodes the external screenshot image data in the external XML script file in the Base64 data format to obtain the corresponding external screenshot image information, imports the external screenshot image information and other information in the external XML script file into the current internal Task dictionary, replaces the content of the current internal Task dictionary, and then delivers the current internal Task dictionary to the instruction information editing module for secondary editing; A user interface display module, which is used to receive the operation information, window name and image information from the instruction information editing module.
2. The UI automation script editing system based on screenshot recognition according to claim 1, wherein The software window acquisition module includes: A screenshot sub-module, which is used to capture a screenshot of the current display desktop and store the screenshot image; A program window name acquisition sub-module, which receives the screenshot image, identifies the window names of all the software currently opened on the display desktop from the screenshot image, and loads the window names of all the software into the user interface display module for display.
3. The UI automation script editing system based on screenshot recognition according to claim 1, wherein The operation information includes: controlling the pointer to click on different positions of the partial screenshot in the interface, controlling the number of clicks and moving the pointer, and controlling the pointer to drag the partial screenshot, as well as typing operation instructions for the corresponding software and inputting text information to the corresponding software.
4. The UI automation script editing system based on screenshot recognition according to claim 3, characterized in that, The operation information further includes loop operation information, monitoring operation information and input sequence number information; wherein: The loop operation information includes the loop operation of the operation instructions and the loop input of the text information; The monitoring operation information adopts the end monitoring method to monitor whether the current software operates normally; The input sequence number information is a number arranged in a certain order according to the loop order.
5. The UI automation script editing system based on screenshot recognition according to claim 1, wherein The instruction information editing module further includes: The email information sub-module is used to obtain the email information of the external XML script file, confirm the email information of the XML script file, and store the email information in the internal Task dictionary; The warning information sub-module is used to generate warning information, which includes handling operations for error information. The handling operations include: indexing to the interface where the error information exists according to the interface information of the error information, and the corresponding operation information for the interface containing the error information, and storing the corresponding operation information in the internal Task dictionary.
6. The UI automation script editing system based on screenshot recognition according to claim 1, wherein The storage information editing module includes: The export operation sub-module encodes the screenshot image information in the Base64 data format to obtain the corresponding screenshot image data, and after compiling the screenshot image data and other information in the internal Task dictionary using the lxml function, exports it as the XML script file; The import operation sub-module receives the external XML script file, imports and replaces the content of the current internal Task dictionary with the elements in the external XML script file, and delivers the current internal Task dictionary to the instruction information editing module; among them, the external screenshot image data is decoded in the Base64 data format to obtain the corresponding external screenshot image information.
7. A UI automation script editing method based on screenshot recognition, according to the UI automation script editing system based on screenshot recognition described in any one of claims 1 to 6, characterized in that, It includes: S1: Whether to import the external XML script file: If yes, control the storage information editing module to read and decode the external XML script file, add and replace the content of the internal Task dictionary with the obtained data information, and then execute step S2; otherwise, directly go to step S3; S2: Determine whether to perform secondary editing on the operation information and screenshot image information in the internal Task dictionary obtained in step S1; if yes, go to step S3; otherwise, execute step S4; S3: Control the software window acquisition module to receive the internal Task dictionary in step S2, and based on the screenshot image and window name in the internal Task dictionary, add the screenshot image and window name of the current desktop; or control the software window acquisition module to directly obtain the screenshot image and window name of the current desktop; S4: Control the instruction information editing module to directly receive and modify the operation information in the internal Task dictionary in step S2; or control the instruction information editing module to modify the operation information in the internal Task dictionary according to the screenshot image and window name in step S3; store the current operation information and screenshot image information as a new internal Task dictionary; S5: Control the storage information editing module to encode the screenshot image information in the new internal Task dictionary obtained in step S4 in the Base64 data format to obtain the corresponding screenshot image data; export the new internal Task dictionary with the screenshot image data as the XML script file.
Citation Information
Patent Citations
Automatic script generation method and tool based on recording technology
CN115509910A
Ui automated script generation tool
CN116340152A
AI algorithm-based Selenium automatic script generation method and system
CN119127721A
Automatic test method and device and computing device
CN105843734A
Editing system of script program for automated testing
CN113126983A