Automatic test script generation method and system based on multiple modes
Through a multimodal automated test script generation method, using large language models and visual analysis technology, the AIRTEST test suite is automatically generated, which solves the problems of low efficiency and difficult maintenance of existing tools and realizes efficient and accurate test script generation.
Patent Information
- Application Number
- CN202510795573.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-30
AI Technical Summary
Existing automated testing tools rely on manually written scripts, which are inefficient. Image recognition cannot generate complete test logic, and screenshots and annotations need to be re-taken when the interface is iterated. The maintenance cost is high, and there is a lack of automated solutions for semantic understanding and dynamic icon positioning.
A multimodal automated test script generation method is adopted to generate test scripts by calling a large language model through natural language descriptions and interface screenshots. Regular expressions and visual analysis are combined to locate elements, and the test suite required by the AIRTEST framework is automatically generated.
It realizes the automatic generation of test logic, improves efficiency, reduces maintenance costs, and allows non-technical personnel to quickly generate test cases. Icon positioning is accurate, solving the learning and maintenance difficulties of traditional tools.
Smart Images

Figure CN120723632A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of test script generation, and in particular to a multimodal-based automated test script generation method and system. Background Art
[0002] Existing automated testing tools, such as AIRTEST, rely on manually written scripts and require manual screenshots to locate UI elements, resulting in low efficiency. Existing technologies attempt to locate controls through image recognition, but are unable to generate complete test logic based on natural language descriptions, and icon coordinates require manual maintenance. Current solutions require re-screenshots and annotations for each interface iteration, resulting in high maintenance costs. The lack of an automated solution that combines semantic understanding with dynamic icon positioning results in long test script generation cycles and poor adaptability. Summary of the Invention
[0003] To overcome the above problems, the purpose of the present invention is to provide a multimodal-based automated test script generation method and system, which can solve the core pain points of traditional automated test tools, such as high learning cost and difficult maintenance.
[0004] The present invention is implemented by the following scheme: a method for generating an automated test script based on multimodality, the method comprising the following steps:
[0005] Step S1: The user uses natural language to describe the test requirement prompt words and uploads a screenshot of the interface for multimodal input. The large language model is called to generate the AIRTEST script. The large language model returns a test script containing image placeholders through visual understanding. The generated test script contains conditional judgment logic;
[0006] Step S2: Use regular expressions to perform pattern matching on the generated script, extract all image element names, and use a parsing function written in Go language to implement de-duplication sorting to form an element name list;
[0007] Step S3: Input the original interface screenshot and the element name list into the visual macro model for multimodal visual analysis and positioning, and require the visual macro model to return standardized coordinate data according to the output prompt word;
[0008] Step S4: perform an image cropping operation based on the coordinate data, and save the cropped icon to a specified directory according to the element name to provide actual image resources for the test script;
[0009] Step S5: Automatically generate a .air project folder named with a timestamp, create a Python script file with the same name in the directory of the .air project folder and write the test script code generated in step 1, save the cropped image resources into the project folder, and form a complete executable test suite that meets the requirements of the AIRTEST framework.
[0010] Furthermore, the de-duplication sorting process is implemented by a parsing function written in Go language in step S2 to form an element name list. The de-duplication sorting process is implemented by a parsing function written in Go language, first creating a temporary hash table to filter duplicate items, then storing the unique element names into slices and sorting them alphabetically, and forming the element name list with the sorting results.
[0011] Furthermore, the output prompt word in step S3 further includes: clearly specifying the JSON format output requirements, including the element name, starting coordinates (x, y) and width and height (w, h) parameters.
[0012] Furthermore, the image cropping operation based on the coordinate data in step S4 is further as follows: the cropping function implemented in the Go language first opens the original screenshot file, creates a rectangular cropping area according to the coordinate parameters, and performs precise pixel-level cutting through the imaging image processing library. The cropping process includes a boundary protection mechanism, and uses the clamp function to ensure that the coordinates do not exceed the image size range.
[0013] The present invention also provides a multimodal-based automated test script generation system, which includes a test script generation module, an element name extraction module, a multimodal visual analysis and positioning module, an image cropping module, and an executable test script module;
[0014] The test script generation module: The user describes the test requirement prompt words in natural language and uploads the interface screenshot for multimodal input, calls the large language model to generate the AIRTEST script, and the large language model returns the test script containing the image placeholder through visual understanding. The generated test script contains conditional judgment logic;
[0015] The element name extraction module uses regular expressions to perform pattern matching on the generated script, extracts all image element names, and forms an element name list by performing de-reordering processing using a parsing function written in Go language;
[0016] The multimodal visual analysis and positioning module inputs the original interface screenshot and the element name list into the visual large model for multimodal visual analysis and positioning, and requires the visual large model to return standardized coordinate data according to the output prompt word;
[0017] The image cropping module performs an image cropping operation based on the coordinate data and saves the cropped icon to a specified directory according to the element name to provide actual image resources for the test script;
[0018] The executable test script module automatically generates a .air project folder named with a timestamp, creates a Python script file with the same name in the directory of the .air project folder and writes the test script code generated by the test script generation module, and stores the cropped image resources in the project folder to form a complete executable test suite that meets the requirements of the AIRTEST framework.
[0019] Furthermore, the element name extraction module implements de-duplication sorting through a parsing function written in Go language to form an element name list. The de-duplication sorting is further implemented by: first, a temporary hash table is created to filter duplicate items, and then the unique element names are stored in slices and alphabetically sorted, and the sorting results are formed into an element name list.
[0020] Furthermore, the output prompt words in the multimodal visual analysis and positioning module further include: clearly specifying the JSON format output requirements, including the element name, starting coordinates (x, y) and width and height (w, h) parameters.
[0021] Furthermore, the image cropping module performs image cropping operations based on coordinate data as follows: the cropping function implemented in the Go language first opens the original screenshot file, creates a rectangular cropping area according to the coordinate parameters, and performs precise pixel-level cutting through the imaging image processing library. The cropping process includes a boundary protection mechanism, and uses the clamp function to ensure that the coordinates do not exceed the image size range.
[0022] The beneficial effects of the present invention are: the present invention provides a multimodal-based UI automated test script generation method and system, which realizes automatic generation of test logic through language models, greatly improving efficiency compared with traditional manual coding; establishing an icon library by dynamically intercepting interface elements, reducing maintenance costs; the coordinate mapping algorithm ensures high icon positioning accuracy, and the multimodal input method enables non-technical personnel to quickly generate test cases, solving the core pain points of traditional automated testing tools, such as high learning costs and difficult maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic flow chart of the method of the present invention.
[0024] Figure 2 It is a system principle block diagram of the present invention. DETAILED DESCRIPTION
[0025] The present invention will be further described below with reference to the accompanying drawings.
[0026] See also Figure 1 As shown, a multimodal automated test script generation method of the present invention comprises the following steps:
[0027] Step S1: The user uses natural language to describe the test requirement prompt words and uploads a screenshot of the interface for multimodal input. The large language model is called to generate the AIRTEST script. The large language model returns a test script containing image placeholders through visual understanding. The generated test script contains conditional judgment logic;
[0028] Step S2: Use regular expressions to perform pattern matching on the generated script, extract all image element names, and use a parsing function written in Go language to implement de-duplication sorting to form an element name list;
[0029] Step S3: Input the original interface screenshot and the element name list into the visual macro model for multimodal visual analysis and positioning, and require the visual macro model to return standardized coordinate data according to the output prompt word;
[0030] Step S4: perform an image cropping operation based on the coordinate data, and save the cropped icon to a specified directory according to the element name to provide actual image resources for the test script;
[0031] Step S5: Automatically generate a .air project folder named with a timestamp, create a Python script file with the same name in the directory of the .air project folder and write the test script code generated in step 1, save the cropped image resources into the project folder, and form a complete executable test suite that meets the requirements of the AIRTEST framework.
[0032] The present invention will be further described below with reference to a specific embodiment:
[0033] A method for generating automated test scripts based on multimodality.
[0034] 1. Users use natural language to describe test requirement prompts and upload interface screenshots for multimodal input. The large language model is called to generate the AIRTEST script. The large language model uses visual understanding to return a test script containing image placeholders. The generated test script contains conditional judgment logic.
[0035] For example: 1. When testing the announcement feature of the game "Hutong Moyu," the user-provided test requirement prompt "Check if there's an announcement button on the game interface. If so, open the announcement" and a game screenshot are input into the large model. The prompt is then formed into a complete form: "Check if there's an announcement button on the game interface. If so, open the announcement. Help me generate an airtest script. If images are used in the script, use the format buttonname.png. Just return the pure script code without any code block symbols." The large model uses visual understanding to return a test script containing image placeholders. Image references use the standardized naming convention "buttonname.png." The generated script contains conditional logic. If an announcement button is detected, a click is executed; otherwise, a prompt message is output.
[0036] The large language model returns a test script with the following code:
[0037]
[0038]
[0039] 2. Use regular expressions to perform pattern matching on the generated script to extract all image element names. Use the parsing function written in Go language to implement de-duplication sorting to form an element name list;
[0040] 2.1 The regular expression in the element name extraction module may further be: Template$r? "\{?([^"}]+)\.png\}? "$.
[0041] 2.2 De-duplication sorting is implemented by a parsing function written in Go language to form a list of element names. Further steps are as follows: De-duplication sorting is implemented by a parsing function written in Go language. First, a temporary hash table is created to filter out duplicates. Then, the unique element names are stored in a slice and sorted alphabetically. The sorted results are used to form a list of element names.
[0042] For example: 2. Parse the ["announcement"] array from the script to provide a list of target elements for subsequent visual analysis.
[0043] The code is as follows:
[0044]
[0045]
[0046] 3. Input the original interface screenshot and element name list into the visual model for multimodal visual analysis and positioning, and require the visual model to return standardized coordinate data based on the output prompt words;
[0047] 3.1 The output prompt further includes: clearly specifying the JSON format output requirements, including the element name, starting coordinates (x, y) and width and height (w, h) parameters.
[0048] For example: 3. Complete the prompt as follows: "Tell me the location information of the following icon in the picture: announcement. You need to tell me the starting coordinates of the icon and the width and height to be cropped. Directly return pure JSON format content without wrapping it with any code block symbols. Format example: [{\"name\":\"element1.png\",\"x\":1000,\"y\":500,\"h\":50,\"w\":50},{\"name\":\"element2.png\",\"x\":1000,\"y\":500,\"h\":50,\"w\":50}]". The large model uses image recognition technology to analyze the button position and return structured data such as [{"name":"announcement.png","x":1760,"y":105,"h":56,"w":56}], which accurately describes the pixel position and size information of each interface element in the screenshot.
[0049] 4. Perform image cropping based on coordinate data and save the cropped icon to the specified directory according to the element name to provide actual image resources for the test script;
[0050] The image cropping operation based on coordinate data in 4.1 is further as follows: the cropping function implemented in Go language first opens the original screenshot file, creates a rectangular cropping area based on the coordinate parameters, and performs precise pixel-level cutting through the imaging image processing library. The cropping process includes a boundary protection mechanism and uses the clamp function to ensure that the coordinates do not exceed the image size range.
[0051] For example: 4. The code is as follows:
[0052]
[0053]
[0054] 5. Automatically generate a timestamp-named .air project folder. Create a Python script file with the same name within the .air project folder and write the test script code generated in step S1. Store the cropped image resources in the project folder to form a complete executable test suite that meets the requirements of the AIRTEST framework. For example, the generated air project folder is ai_1747964395917.ai. The final packaged directory contains the executable script and associated image resources and can be directly imported into the AIRTEST IDE to run the test, achieving a one-click generation process from requirement description to executable script.
[0055] See also Figure 2 As shown, the present invention also provides a multimodal-based automated test script generation system, the system comprising a test script generation module, an element name extraction module, a multimodal visual analysis and positioning module, an image cropping module, and an executable test script module;
[0056] The test script generation module: The user describes the test requirement prompt words in natural language and uploads the interface screenshot for multimodal input, calls the large language model to generate the AIRTEST script, and the large language model returns the test script containing the image placeholder through visual understanding. The generated test script contains conditional judgment logic;
[0057] For example: 1. When testing the announcement feature of the game "Hutong Moyu," the user-provided test requirement prompt "Check if there's an announcement button on the game interface. If so, open the announcement" and a game screenshot are input into the large model. The prompt is then formed into a complete form: "Check if there's an announcement button on the game interface. If so, open the announcement. Help me generate an airtest script. If images are used in the script, use the format buttonname.png. Just return the pure script code without any code block symbols." The large model uses visual understanding to return a test script containing image placeholders. Image references use the standardized naming convention "buttonname.png." The generated script contains conditional logic. If an announcement button is detected, a click is executed; otherwise, a prompt message is output.
[0058] The large language model returns a test script with the following code:
[0059]
[0060] The element name extraction module uses regular expressions to perform pattern matching on the generated script, extracts all image element names, and forms an element name list by performing de-reordering processing using a parsing function written in Go language;
[0061] 2.1 The regular expression in the element name extraction module may further be: Template$r? "\{?([^"}]+)\.png\}? "$.
[0062] 2.2 The element name extraction module further forms an element name list by implementing de-duplication sorting processing through a parsing function written in Go language: first, a temporary hash table is created to filter duplicate items, and then the unique element names are stored in a slice and alphabetically sorted, and the sorting results are formed into an element name list.
[0063] For example: 2. Parse the ["announcement"] array from the script to provide a list of target elements for subsequent visual analysis.
[0064] The code is as follows:
[0065]
[0066]
[0067] The multimodal visual analysis and positioning module inputs the original interface screenshot and the element name list into the visual large model for multimodal visual analysis and positioning, and requires the visual large model to return standardized coordinate data according to the output prompt word;
[0068] 3.1 The output prompt words in the multimodal visual analysis and positioning module further include: clearly specifying the JSON format output requirements, including the element name, starting coordinates (x, y) and width and height (w, h) parameters.
[0069] For example: 3. Complete the prompt as follows: "Tell me the location information of the following icon in the picture: announcement. You need to tell me the starting coordinates of the icon and the width and height to be cropped. Directly return pure JSON format content without wrapping it with any code block symbols. Format example: [{\"name\":\"element1.png\",\"x\":1000,\"y\":500,\"h\":50,\"w\":50},{\"name\":\"element2.png\",\"x\":1000,\"y\":500,\"h\":50,\"w\":50}]". The large model uses image recognition technology to analyze the button position and return structured data such as [{"name":"announcement.png","x":1760,"y":105,"h":56,"w":56}], which accurately describes the pixel position and size information of each interface element in the screenshot.
[0070] The image cropping module performs an image cropping operation based on the coordinate data and saves the cropped icon to a specified directory according to the element name to provide actual image resources for the test script;
[0071] 4.1 The image cropping module further performs image cropping operations based on coordinate data as follows: the cropping function implemented in the Go language first opens the original screenshot file, creates a rectangular cropping area based on the coordinate parameters, and performs precise pixel-level cutting through the imaging image processing library. The cropping process includes a boundary protection mechanism, and uses the clamp function to ensure that the coordinates do not exceed the image size range.
[0072] For example: 4. The code is as follows:
[0073]
[0074]
[0075]
[0076] The executable test script module automatically generates a timestamp-named .air project folder, creates a Python script file with the same name in the directory of the .air project folder, writes the test script code generated by the test script generation module, and stores the cropped image resources in the project folder to form a complete executable test suite that meets the requirements of the AIRTEST framework. The air project folder is generated as ai_1747964395917.ai, for example. The final packaged directory contains executable scripts and associated image resources, which can be directly imported into the AIRTEST IDE to run the test, realizing a one-click generation process from requirement description to executable script.
[0077] In summary, the present invention provides a multimodal-based UI automated test script generation method and system, which realizes automatic generation of test logic through language models, greatly improving efficiency compared with traditional manual coding; establishes an icon library by dynamically intercepting interface elements, reducing maintenance costs; the coordinate mapping algorithm ensures high icon positioning accuracy, and the multimodal input method enables non-technical personnel to quickly generate test cases, solving the core pain points of traditional automated testing tools, such as high learning cost and difficult maintenance.
[0078] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.
Claims
1. A method for generating automated test scripts based on multimodality, characterized by: The method comprises the following steps: Step S1: The user uses natural language to describe the test requirement prompt words and uploads a screenshot of the interface for multimodal input. The large language model is called to generate the AIRTEST script. The large language model returns a test script containing image placeholders through visual understanding. The generated test script contains conditional judgment logic; Step S2: Use regular expressions to perform pattern matching on the generated script, extract all image element names, and use a parsing function written in Go language to implement de-duplication sorting to form an element name list; Step S3: Input the original interface screenshot and the element name list into the visual macro model for multimodal visual analysis and positioning, and require the visual macro model to return standardized coordinate data according to the output prompt word; Step S4: perform an image cropping operation based on the coordinate data, and save the cropped icon to a specified directory according to the element name to provide actual image resources for the test script; Step S5: Automatically generate a .air project folder named with a timestamp, create a Python script file with the same name in the directory of the .air project folder and write the test script code generated in step 1, save the cropped image resources into the project folder, and form a complete executable test suite that meets the requirements of the AIRTEST framework.
2. The method for generating automated test scripts based on multimodality according to claim 1, characterized in that: The step S2 further comprises: implementing de-duplication sorting by a parsing function written in Go language to form an element name list, first creating a temporary hash table to filter duplicates, then storing the unique element names into slices and sorting them alphabetically, and forming the element name list with the sorting results.
3. The method for generating automated test scripts based on multimodality according to claim 1, wherein: Outputting the prompt word in step S3 further includes: clearly specifying the JSON format output requirements, including the element name, starting coordinates (x, y) and width and height (w, h) parameters.
4. The method for generating automated test scripts based on multimodality according to claim 1, wherein: The image cropping operation based on the coordinate data in step S4 is further performed as follows: the cropping function implemented in the Go language first opens the original screenshot file, creates a rectangular cropping area according to the coordinate parameters, and performs precise pixel-level cutting through the imaging image processing library. The cropping process includes a boundary protection mechanism, and uses the clamp function to ensure that the coordinates do not exceed the image size range.
5. A multimodal automated test script generation system, characterized by: The system includes a test script generation module, an element name extraction module, a multimodal visual analysis and positioning module, an image cropping module, and an executable test script module; The test script generation module: The user describes the test requirement prompt words in natural language and uploads the interface screenshot for multimodal input, calls the large language model to generate the AIRTEST script, and the large language model returns the test script containing the image placeholder through visual understanding. The generated test script contains conditional judgment logic; The element name extraction module uses regular expressions to perform pattern matching on the generated script, extracts all image element names, and forms an element name list by performing de-reordering processing using a parsing function written in Go language; The multimodal visual analysis and positioning module inputs the original interface screenshot and the element name list into the visual large model for multimodal visual analysis and positioning, and requires the visual large model to return standardized coordinate data according to the output prompt word; The image cropping module performs an image cropping operation based on the coordinate data and saves the cropped icon to a specified directory according to the element name to provide actual image resources for the test script; The executable test script module automatically generates a .air project folder named with a timestamp, creates a Python script file with the same name in the directory of the .air project folder and writes the test script code generated by the test script generation module, and stores the cropped image resources in the project folder to form a complete executable test suite that meets the requirements of the AIRTEST framework.
6. The multimodal automated test script generation system according to claim 1, characterized in that: The element name extraction module further forms an element name list by implementing de-duplication sorting processing through a parsing function written in Go language: first, a temporary hash table is created to filter duplicate items, and then the unique element names are stored in slices and alphabetically sorted, and the sorting results are formed into an element name list.
7. The multimodal automated test script generation system according to claim 1, characterized in that: The output prompt words in the multimodal visual analysis and positioning module further include: clearly specifying the JSON format output requirements, including the element name, starting coordinates (x, y) and width and height (w, h) parameters.
8. The multimodal automated test script generation system according to claim 1, characterized in that: The image cropping module further performs image cropping operations based on coordinate data as follows: the cropping function implemented in the Go language first opens the original screenshot file, creates a rectangular cropping area based on the coordinate parameters, and performs precise pixel-level cutting through the imaging image processing library. The cropping process includes a boundary protection mechanism, and uses the clamp function to ensure that the coordinates do not exceed the image size range.
Citation Information
Cited By
Anti-illusion element positioning method and system based on RPA and visual semantic marking
CN122049915A