Automatic testing method and system based on retrieval enhancement generation and large language model
Through automated testing methods based on retrieval enhancement generation and large language models, traditional automated testing has solved the problems of high maintenance costs, poor flexibility and insufficient coverage, and automated testing of natural language use cases has been realized, which has improved testing efficiency and coverage.
Patent Information
- Application Number
- CN202510270061.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional automated tests have high maintenance costs, poor flexibility and insufficient coverage, making it difficult for the existing technology to adapt to dynamically changing test scenarios.
An automated testing method based on retrieval enhancement generation and large language models is adopted. A database is constructed in vector by obtaining historical test cases, and a large language model is used to simulate real people to execute test steps and judge the results. It combines OCR and YOLO object detection technology for screenshot recognition and exception processing.
Reduces the execution and maintenance cost of automated tests, improves test flexibility and coverage, saves testers time, and shortens the software development cycle.
Smart Images

Figure CN120295907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and software testing, and more particularly, to an automated testing method and system based on Retrieval-Augmented Generation (RAG) and Large Language Model (LLM). Background Art
[0002] Traditional automated testing and recording / playback technologies have been used in the software testing field for many years, but they have some inherent problems and limitations:
[0003] 1. High maintenance cost: Traditional automated test scripts are usually written manually or generated by recording. When the UI or functions of the application change, these scripts need to be updated and maintained frequently.
[0004] 2. Poor flexibility: Traditional automated test scripts usually contain hard-coded steps and data, lacking flexibility and being difficult to adapt to dynamically changing test scenarios.
[0005] 3. Insufficient test coverage: Manually written test scripts can only cover limited test scenarios due to cost issues, making it difficult to achieve comprehensive test coverage.
[0006] For example, a patent document with the publication number "CN119512941A" provides a testing method and device. The method includes: obtaining code submission records; determining code change files based on the code submission records; performing call relationship analysis on the code change files to determine the target interfaces affected by the code changes, where the target interfaces refer to the externally exposed interfaces affected by the code changes; searching for target test cases associated with the target interfaces from existing test cases; performing tests based on the target test cases and obtaining test results. This method searches for and tests target cases from existing test cases. On the one hand, it still requires manually writing test scripts for the target cases, resulting in low test efficiency, high maintenance cost, and poor flexibility. On the other hand, the target cases are selected based on "existing test cases", and there is still a problem of insufficient coverage. Summary of the Invention
[0007] To overcome the defects of high maintenance cost, poor flexibility, and insufficient coverage of traditional automated testing in the above-mentioned prior art, the present invention provides an automated testing method and system based on retrieval-enhanced generation and large language models. Through large model-driven UI automated testing technology, it simulates real people to automatically execute tests and judge test results, thereby realizing automated testing of natural language use cases, reducing the execution and maintenance costs of automated testing, and improving test flexibility and coverage.
[0008] To solve the above technical problems, the technical solution of the present invention is as follows:
[0009] An automated testing method based on retrieval-augmented generation and large language models, comprising the following steps:
[0010] Obtain a number of historical test cases from previous periods and vectorize them respectively to construct a vector database;
[0011] Obtain a new test case described in natural language, and use the retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test case and obtain the improved test case;
[0012] Convert the improved test case into a number of consecutive test steps, start the software under test, perform analysis and decision-making based on the large language model, and call the automated testing tool to automatically execute each test step;
[0013] After all test steps are executed, obtain the test results, and use the large language model to analyze the test results to complete the automated testing.
[0014] Preferably, during the automatic execution process, after each test step is executed, take a screenshot of the current page of the software under test, identify the text and icons in the screenshot, and obtain the screenshot recognition result; input the screenshot recognition result and the next test step into the large language model for analysis and decision-making, and output the first operation instruction; the automated testing tool executes the next test step according to the first operation instruction.
[0015] Preferably, during the automatic execution process, the automated testing tool calls the screenshot API of the software under test to take a screenshot of the current page.
[0016] Preferably, use OCR (Optical Character Recognition) technology to identify the text in the screenshot and the coordinates of each text, use the YOLO target detection series models to identify the icons in the screenshot and the coordinates of each icon, and perform holistic understanding of the content of each icon to obtain the screenshot recognition result.
[0017] Preferably, after all test steps are executed, take a screenshot of the final page of the software under test, identify the text and icons in the screenshot of the final page, and obtain the screenshot recognition result of the final page as the test result;
[0018] Input the test result into the large language model for analysis to determine whether the expected goal described in the new test case is achieved. If it is achieved, the new test case is successfully executed; otherwise, the new test case fails to execute.
[0019] Save the analysis results of the large language model to complete the automated test.
[0020] Preferably, the format of the new test case described in natural language is the BDD format.
[0021] Preferably, the automated test tool includes at least any one of UIAutomator and ADB (Android Debug Bridge).
[0022] Preferably, the large language model includes at least any one or more of GPT-4o, Claude-3.5, Gemini-1.5, GLM-4, QWen, ERNIE, DeepSeek, and Hunyuan.
[0023] Preferably, the method further includes:
[0024] During the automatic execution process, if the large language model analyzes that an abnormal situation appears in the screenshot recognition result, a second operation instruction for solving the abnormal situation is preferentially output, and the automated execution tool executes an operation according to the second operation instruction to solve the corresponding abnormal situation;
[0025] After the abnormal situation is solved, continue to automatically execute the next test step.
[0026] The present invention also provides an automated test system based on retrieval augmented generation and a large language model, applying the above-mentioned automated test method based on retrieval augmented generation and a large language model, including:
[0027] Knowledge base construction unit: used to obtain a number of historical test cases in the past and vectorize them respectively to construct a vector database;
[0028] Retrieval augmented generation unit: used to obtain a new test case described in natural language, and use retrieval augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test case and obtain the improved test case;
[0029] Automated test unit: used to convert the improved test case into a number of consecutive test steps, start the software under test, perform analysis and decision-making based on the large language model, and call the automated test tool to automatically execute each test step;
[0030] Result analysis unit: used to obtain the test result after all test steps are executed, and use the large language model to analyze the test result to complete the automated test.
[0031] Compared with the prior art, the beneficial effect of the technical solution of the present invention is:
[0032] The present invention provides an automated testing method and system based on retrieval-augmented generation and large language models. First, a number of historical test cases from previous periods are obtained and vectorized respectively to construct a vector database. Then, new test cases described in natural language are obtained, and the retrieval-augmented generation technology is used to index relevant content in the vector database as auxiliary information to improve the new test cases and obtain the improved test cases. After that, the improved test cases are converted into a number of coherent test steps, the software under test is started, and based on the large language model, analysis and decision-making are carried out, and an automated testing tool is called to automatically execute each test step. After all test steps are executed, the test results are obtained, and finally, the large language model is used to analyze the test results to complete the automated testing.
[0033] The present invention directly converts manual test cases into full-process automated execution without the need to write automated scripts, saving the time for writing automated test scripts. According to the actual use effect, the present invention can save 10% - 70% of the time of testers, which will help shorten the software development cycle and improve the overall R & D efficiency.
[0034] At the same time, the present invention uses large model technology to simulate real people to analyze and make decisions on the specific operations of each test step, calls an automated testing tool to automatically execute the test steps, and finally analyzes the test results. The automated execution based on the large model can reduce the time for writing automated test scripts, improve the stability of use case execution, enhance the comprehensiveness of test result judgment, and improve test flexibility and coverage.
[0035] In addition, the present invention realizes automated anomaly detection and handling based on large model technology. The large model can simulate people to judge abnormal situations occurring in each step and make correct operations, thus avoiding the problem that traditional automated execution cannot continue when encountering anomalies, and reducing the execution and maintenance costs of automated testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flowchart of an automated testing method based on retrieval-augmented generation and large language models provided in Embodiment 1.
[0037] Figure 2 It is a flowchart of an automated testing method based on retrieval-augmented generation and large language models provided in Embodiment 2.
[0038] Figure 3 It is an example diagram of text and icon recognition results provided in Embodiment 2.
[0039] Figure 4 It is a structural diagram of an automated testing system based on retrieval-augmented generation and large language models provided in Embodiment 3. Detailed implementation manners
[0040] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the present patent;
[0041] For better illustration of this embodiment, some components in the accompanying drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product;
[0042] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0043] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0044] Embodiment 1
[0045] As Figure 1 shown, this embodiment provides an automated testing method based on retrieval-augmented generation and large language models, including the following steps:
[0046] S1: Obtain a number of historical test cases from previous periods and vectorize them respectively to construct a vector database;
[0047] S2: Obtain a new test case described in natural language, and use the retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test case and obtain the improved test case;
[0048] S3: Convert the improved test case into a number of coherent test steps, start the software under test, perform analysis and decision-making based on the large language model, and call an automated testing tool to automatically execute each test step;
[0049] S4: After all test steps are executed, obtain the test results, and use the large language model to analyze the test results to complete the automated testing.
[0050] In the specific implementation process, with the breakthrough of large models in the fields of natural language processing and artificial intelligence, using large model-driven UI automated testing has gradually become a new trend; using a large model to understand test cases described in natural language, this method reflects anthropomorphic intelligence; therefore, this embodiment proposes an idea, that is, regarding the large model as a tester for use case regression, inputting logical manual use cases to it, adding some knowledge reserves and current screen information and endowing it with the ability to operate software, then it can directly execute use cases instead of people;
[0051] Specifically, first obtain a number of historical test cases from previous periods and vectorize them respectively to construct a vector database;
[0052] Next, obtain a new test case described in natural language. In this embodiment, testers can use natural language to describe test cases in any format (such as BDD format, etc.); this input method conforms to people's daily expression habits, reduces the threshold for writing test cases, and testers do not need to master complex programming syntax. They only need to clearly describe the test scenario, steps, and expected results, providing basic input for subsequent automated testing in an easy-to-understand way;
[0053] Secondly, adopt the retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test case and obtain the improved test case; in this embodiment, use the vector database storing past manual and automated test cases to improve the use case. After inputting a new test case, relevant auxiliary information is obtained by retrieving the vector database. After being retrieved by the system, it is used to supplement and improve the steps of the current test case, making the test case more complete and accurate in key information such as the execution path, providing detailed guidance for subsequent automated execution;
[0054] After that, convert the improved test case into several consecutive test steps that are more convenient for the large model to understand and can be automatically executed;
[0055] Start executing the test case. First, start the software under test, analyze and make decisions based on the large language model, and call the automated test tool to automatically execute each test step according to the decision result output by the large model;
[0056] After all test steps are executed, obtain the test results. Finally, use the large language model to analyze the test results to determine whether the execution of the test case is successful, thus completing the automated test;
[0057] This method uses the large model-driven UI automation testing technology to simulate real people to automatically execute tests and judge test results, thus realizing the automated testing of natural language use cases, reducing the execution and maintenance costs of automated testing, and improving test flexibility and coverage.
[0058] Embodiment 2
[0059] This embodiment provides an automated testing method based on retrieval-augmented generation and large language model, including the following steps:
[0060] S1: Obtain several historical test cases in the past and vectorize them respectively to construct a vector database;
[0061] S2: Obtain a new test case described in natural language, and adopt the retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test case and obtain the improved test case;
[0062] S3: Convert the improved test cases into several consecutive test steps, start the software under test, perform analysis and decision-making based on the large language model, and call the automated testing tool to automatically execute each test step;
[0063] During the automatic execution process, after each test step is executed, take a screenshot of the current page of the software under test, identify the text and icons in the screenshot, and obtain the screenshot recognition result; input the screenshot recognition result and the next test step into the large language model for analysis and decision-making, and output the first operation instruction; the automated testing tool executes the next test step according to the first operation instruction;
[0064] During the automatic execution process, if the large language model analyzes that an abnormal situation appears in the screenshot recognition result, it preferentially outputs a second operation instruction for solving the abnormal situation, and the automated execution tool performs operations according to the second operation instruction to solve the corresponding abnormal situation;
[0065] After the abnormal situation is solved, continue to automatically execute the next test step;
[0066] S4: After all test steps are executed, obtain the test result, and use the large language model to analyze the test result to complete the automated testing;
[0067] In this embodiment, during the automatic execution process, the automated testing tool calls the screenshot API of the software under test to take a screenshot of the current page;
[0068] In this embodiment, use OCR (Optical Character Recognition) optical character recognition technology to identify the text and the coordinates of each text in the screenshot, use the YOLO target detection series models to identify the icons and the coordinates of each icon in the screenshot, and perform comprehensive understanding of the content of each icon to obtain the screenshot recognition result;
[0069] In this embodiment, after all test steps are executed, take a screenshot of the final page of the software under test, and identify the text and icons in the screenshot of the final page to obtain the screenshot recognition result of the final page as the test result;
[0070] Input the test result into the large language model for analysis to determine whether the expected goal described in the new test case is achieved. If it is achieved, the new test case is executed successfully; otherwise, the new test case is executed failed;
[0071] Save the analysis result of the large language model to complete the automated testing;
[0072] The format of the new test case described in natural language is in BDD format;
[0073] The automated testing tool includes at least any one of UIAutomator and ADB;
[0074] The large language model includes at least any one or more of GPT-4o, Claude-3.5, Gemini-1.5, GLM-4, QWen, ERNIE, DeepSeek, and Hunyuan.
[0075] In the specific implementation process, this embodiment takes a certain software on an Android phone as an example for illustration. For example, Figure 2 As shown, first, obtain a number of historical test cases from previous periods and vectorize them respectively to construct a vector database;
[0076] Then, input a new test case described in natural language. Testers can use natural language to describe the test case, and the format is not limited (such as BDD format, etc.); this input method conforms to people's daily expression habits, reduces the threshold for writing test cases, and testers do not need to master complex programming syntax. They only need to clearly describe the test scenario, steps, and expected results, providing basic input for subsequent automated testing in this easy-to-understand way;
[0077] For example, the following is an example of a test case. Such a test case can be executed by testers familiar with the business for manual testing, but it is very difficult for an automated testing framework to execute and needs to be converted into a format convenient for automated testing execution;
[0078] Test case example:
[0079] "Title: 'Long press on text to select translation'";
[0080] Steps:
[0081] 1. Enter m.sohu.com and enter any news article body;
[0082] 2. Long press on the paragraph text;
[0083] 3. Click on translation;
[0084] Expected result: Pop up a translation dialog box and display the translated content of the selected text.";
[0085] After that, perform RAG (vector database) optimization. When a new test case is input, retrieve relevant auxiliary information by searching the vector database. After retrieving, use it to supplement and perfect the steps of the current test case, making the test case more complete and accurate in key information such as the execution path, providing detailed guidance for subsequent automated execution;
[0086] In this embodiment, the pre-set vector database stores manual and automated test cases accumulated in the past. Examples of historical test cases stored in the vector database are as follows:
[0087] "Path to enter a specified web page: Click on "Search the Whole Web", enter the website address, and click on "Enter";
[0088] Path to enter the toolbox: Click on "My", "Toolbox";
[0089] When entering any web page, if no website address is specified, it is recommended to use the qq.com website address;
[0090] Path to enter the message notification settings page: Click on "My", "Settings";
[0091] Path to enter the message center page: Click on "My", at the top right of the page button;
[0092] For the auxiliary information retrieved, taking entering a specific web page as an example, if the test case only mentions entering but does not specify the entry path, the vector database may store information such as "Path to enter a specified web page: Click on 'Search the Whole Web', enter the website address, and click on 'Enter'", which can be used as auxiliary information to complete the test case;
[0093] After that, with the help of prompt engineering, detailed specifications are provided to the large language model; clearly define the modules, titles, steps, expected results, etc. included in the use case, and at the same time give the auxiliary information, and put forward requirements such as ensuring that the precondition steps are not lost, the semantics are not lost, only one operation is performed in one step, and at least one assertion step is included, and also specify the specific output format;
[0094] For example, after giving the large language model the information related to the test case "Long press the text to select translation", the large language model converts it into the following 8 test steps according to the requirements, thus converting the natural language description into a format that conforms to the automated test operation logic:
[0095] 1. Click on "Search the Whole Web";
[0096] 2. Enter "m.sohu.com";
[0097] 3. Click on "Enter";
[0098] 4. Click on any news text;
[0099] 5. Long press the paragraph text;
[0100] 6. Click on "Translate";
[0101] 7. Assert that there is a "translation pop-up window";
[0102] 8. Assert that there is a selected text translation content;
[0103] After that, start the software under test and prepare to execute the above test steps; although there are test cases suitable for automated test execution, it is still necessary to execute the cases on the software under test and judge the assertion results; in this embodiment, it is necessary to select a suitable automated test framework (tool), and the automated test framework is to simulate human operations, such as clicking on a certain control or coordinate position, completing sliding or zooming, etc.
[0104] Common automated test frameworks include UIAutomator or ADB, and these two frameworks are less invasive to the program under test; among them, UIAutomator is a powerful test framework provided by Google, which is specifically used for automated testing of the user interface of Android applications. It allows developers to write scripts to simulate user interactions, verify the status of UI elements, and support cross-application and system-level operations; ADB is a multi-functional command-line tool that allows developers to communicate with and operate Android devices, including installing applications, debugging, transferring files, and executing device commands.
[0105] After selecting the automated test tool, you can start automatically executing the above test steps.
[0106] During the automatic execution process, after each test step is executed, the automated test tool will call the relevant screenshot API to take a screenshot of the current page of the software under test; then identify the text and icons in the screenshot to obtain the screenshot recognition result, as Figure 3 shown, Figure 3 is a screenshot of a mobile phone screen page, which shows an example of the text and icon recognition results.
[0107] In this embodiment, the OCR technology is used to recognize the text in the intercepted picture, and recognize the text information and its coordinate position in the picture. OCR is a text recognition technology used to extract and convert printed or handwritten text from images, scanned documents, or photos into editable and searchable machine-encoded text; after recognizing the text, the LLM analyzes whether the target is operable. If it is, perform automated operations. After the operations are completed, continue to take screenshots and perform OCR, repeat the operations, and execute the case steps until the assertion is completed.
[0108] For example, some of the OCR recognition results are:
[0109] [473, 956]; text: The positive economic development of cold resources is accelerating the transformation of "hot power";
[0110] [87, 355]; text: Video;
[0111] [142,199]; text: Search the entire network;
[0112] After performing OCR on the image containing the text "Search the entire network", the text "Search the entire network" and its coordinates [142,199] in the image can be obtained; if there is an executable target corresponding to the test case steps in the OCR result, such as "Search the entire network", the automation framework can directly perform a click operation based on the coordinates, thereby accurately operating the interface of the software under test;
[0113] For the icons existing in the screenshot, in this embodiment, the YOLO algorithm is used for icon recognition. The icon information and coordinates are recognized through the target detection algorithm, and the icon content is comprehensively understood. It is analyzed by the LLM and automated operations are performed. After the operation is completed, the screenshot is continued for icon recognition, and the operation is repeated until the assertion is completed;
[0114] For example, some results of icon recognition and comprehensive understanding are as follows:
[0115] [971,2269]; icon: This icon is a square with a black border and a black number "4" inside; the background is white;
[0116] [766,2269]; icon: This icon is in the shape of a simple black house, and the background is white; the icon lines are simple, presenting the basic outline of a roof and walls of a house;
[0117] [366,1290]; icon: This icon is a circle, the main color is yellow, and there is a cross symbol in the center; the background has some blurred yellow tones;
[0118] For the recognized icons, when the test steps require operating the icons, the automation framework performs corresponding operations based on the icon recognition results, providing accurate operation basis for automated testing;
[0119] If there is an operable target required by the symbol in the screenshot, the large model will directly call the automation testing framework to perform the corresponding operation; if no relevant execution target is recognized, it will be handed over to the LLM to make a decision on how to operate; for example, in the above test case, "Click on any news text", there is no relevant information in the OCR or icon description, and here the LLM needs to make a decision on how to operate; the LLM randomly selects and clicks on one of the news texts based on the content recognized by the OCR (multiple news texts on the current page), outputs the corresponding operation instructions, and then performs the relevant operations;
[0120] Additionally, if an exception occurs during the automatic execution process, such as an advertisement pop-up window suddenly appearing when opening the text of a piece of information, the LLM will prioritize the decision to click and close the advertisement pop-up window; another example is that if a phone call comes in during the execution of the steps, the LLM will also prioritize the decision to hang up the phone and continue to execute the relevant steps;
[0121] Specifically, taking "an advertisement pop-up window appears when entering the web page" as an example, an example of the decision result output by the large model:
[0122] Thought;
[0123] According to the user's instruction, we have completed the first three steps of the operation; there is an advertisement pop-up window on the current page, which hinders us from continuing to execute the subsequent operations; in order to continue to execute the subsequent operations, we need to close the advertisement pop-up window first; the advertisement pop-up window can usually be closed by clicking the "x" button in the upper right corner; according to the information provided, the "x" button of the advertisement pop-up window may be located at [366,1290];
[0124] Action
[0125] Tap(366,1290);
[0126] Operation
[0127] Click the "x" button in the upper right corner of the advertisement pop-up window to close the advertisement pop-up window;
[0128] After the abnormal situation is handled, continue to execute the above test steps. If the previous test step has not been completed, take a new screenshot and re-execute the previous test step;
[0129] After all the test steps are completed, the large language model will conduct result analysis. The analysis makes judgments from two dimensions. On the one hand, it is related keywords, such as "select the translated content of the text", and on the other hand, it will also make a graphical interface judgment to see if there are any abnormalities in the overall layout, such as the appearance of a "translation pop-up window" and whether there is overlap, etc.;
[0130] During the result analysis process, the large model utilizes its multi-modal processing ability to analyze the pictures captured during the execution process. By considering information such as the element layout, state changes, and text markings in the pictures, it determines the success or failure of the test case execution; for example, in the "long press on text to select translation" use case, the large model analyzes the screenshot to determine whether a translation pop-up window appears and whether the translated content of the selected text is displayed in the pop-up window; if a translation pop-up window appears on the page and the correct translated content is displayed, it is determined that the execution is successful; if there is no translation pop-up window after clicking the translation, it is determined that the execution fails, thus ensuring the integrity of the automated test process and the reliability of the results;
[0131] Example of successful execution of a large model:
[0132] According to the previous operation steps and the screenshot of the current page, we can see that the translation pop-up box has popped up on the page, and the translation content of the selected text "largest" is displayed; therefore, the assertion "call up the translation pop-up box and display the translation content of the selected text" is correct;
[0133] ###Action###
[0134] Assert(true);
[0135] ###Operation###
[0136] The current page has successfully called up the translation pop-up box and displayed the translation content of the selected text;
[0137] Example of a large model judgment execution failure:
[0138] According to the previous operation steps, the user has clicked the "Translate" button; the current page does not display any translation pop-up box content or prompts, so it can be concluded that the translation pop-up box does not exist;
[0139] ###Action###
[0140] Assert(false);
[0141] ###Operation###
[0142] The current page does not display the translation pop-up box;
[0143] Regarding the selection of large language models, this method is compatible with all existing large language models, such as GPT-4o, Claude-3.5, Gemini-1.5, GLM-4, QWen, ERNIE, DeepSeek, and Hunyuan, etc., without limitation here;
[0144] This method firstly improves the description of the test case through the auxiliary information of the vector database, converts it into a format that is easy for the large model LLM to understand, and then implements the automated test. In the process, screenshots are taken, and then the screenshots are OCRed and icons are recognized. The OCR text coordinate information and the icon description and coordinate information are sent to the LLM for decision-making, driving the automated framework to execute relevant steps, and finally the execution is completed to judge the assertion result and output the test result. This method simulates the automated execution of tests and the judgment of test results by real people based on the large model, thereby realizing the automated testing of natural language use cases, reducing the execution and maintenance costs of automated tests, and improving the test flexibility and coverage.
[0145] Example 3
[0146] As shown Figure 4 in the figure, this embodiment provides an automated testing system based on retrieval-augmented generation and large language models, which applies the automated testing method based on retrieval-augmented generation and large language models described in Embodiment 1 or 2, and includes:
[0147] Knowledge base construction unit 301: used to obtain several historical test cases of previous periods, vectorize them respectively, and construct a vector database;
[0148] Retrieval-augmented generation unit 302: used to obtain new test cases described in natural language, and use retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test cases and obtain improved test cases;
[0149] Automated testing unit 303: used to convert the improved test cases into several coherent test steps, start the software under test, perform analysis and decision-making based on the large language model, and call the automated testing tool to automatically execute each test step;
[0150] Result analysis unit 304: used to obtain the test results after all test steps are executed, and use the large language model to analyze the test results to complete the automated testing.
[0151] In the specific implementation process, first, the knowledge base construction unit 301 obtains several historical test cases of previous periods, vectorizes them respectively, and constructs a vector database;
[0152] Then, the retrieval-augmented generation unit 302 obtains new test cases described in natural language, and uses retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test cases and obtain improved test cases;
[0153] After that, the automated testing unit 303 converts the improved test cases into several coherent test steps, starts the software under test, performs analysis and decision-making based on the large language model, and calls the automated testing tool to automatically execute each test step;
[0154] Finally, after all test steps are executed, the result analysis unit 304 obtains the test results, uses the large language model to analyze the test results, and completes the automated testing;
[0155] This system uses large model-driven UI automated testing technology to simulate real people to automatically execute tests and judge test results, thereby realizing the automated testing of natural language use cases, reducing the execution and maintenance costs of automated testing, and improving test flexibility and coverage.
[0156] Like or similar reference numerals correspond to like or similar components;
[0157] The terms used in the drawings to describe positional relationships are for illustrative purposes only and should not be construed as limiting the patent;
[0158] Obviously, the above embodiments of the present invention are merely examples given to clearly illustrate the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. An automated testing method based on retrieval-augmented generation and large language models, characterized in that, Including the following steps: Obtain a number of historical test cases from previous periods, vectorize them respectively, and construct a vector database; Obtain a new test case described in natural language, and use retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test case and obtain the improved test case; Convert the improved test case into a number of coherent test steps, start the software under test, perform analysis and decision-making based on a large language model, and call an automated testing tool to automatically execute each test step; After all test steps are executed, obtain the test results, use the large language model to analyze the test results, and complete the automated testing.
2. The automated testing method based on retrieval-augmented generation and large language model according to claim 1, wherein, During the automatic execution process, after each test step is executed, take a screenshot of the current page of the software under test, identify the text and icons in the screenshot, and obtain the screenshot recognition result; input the screenshot recognition result and the next test step into the large language model for analysis and decision-making, and output the first operation instruction; The automated testing tool executes the next test step according to the first operation instruction.
3. The automated testing method based on retrieval-augmented generation and large language models according to claim 2, wherein, During the automatic execution process, the automated testing tool calls the screenshot API of the software under test to take a screenshot of the current page.
4. The automated testing method based on retrieval-augmented generation and large language model according to claim 2, wherein, Use OCR optical character recognition technology to identify the text and the coordinates of each text in the screenshot, use the YOLO object detection series model to identify the icons and the coordinates of each icon in the screenshot, and perform Hunyuan understanding on the content of each icon to obtain the screenshot recognition result.
5. The automated testing method based on retrieval-augmented generation and large language model according to claim 1, characterized in that After all test steps are executed, take a screenshot of the final page of the software under test, identify the text and icons in the screenshot of the final page, and obtain the screenshot recognition result of the final page as the test result; Input the test result into the large language model for analysis to determine whether the expected goal described in the new test case is achieved. If it is achieved, the new test case is executed successfully; otherwise, the new test case is executed failed; Save the analysis result of the large language model to complete the automated testing.
6. The automated testing method based on retrieval-augmented generation and large language model according to claim 1, characterized in that, The format of the new test case described in natural language is in BDD format.
7. The automated testing method based on retrieval-augmented generation and large language model according to claim 1, characterized in that The automated testing tool includes at least any one of UIAutomator and ADB.
8. The automated testing method based on retrieval-augmented generation and large language model according to claim 1, wherein The large language model includes at least any one or more of GPT-4o, Claude-3.5, Gemini-1.5, GLM-4, QWen, ERNIE, DeepSeek, and Hunyuan.
9. The automated testing method based on retrieval-augmented generation and large language model according to any one of claims 2 to 8, characterized in that The method further includes: During the automatic execution process, if the large language model analyzes that an abnormal situation appears in the screenshot recognition result, it preferentially outputs a second operation instruction for solving the abnormal situation, and the automated execution tool performs an operation according to the second operation instruction to solve the corresponding abnormal situation; After the abnormal situation is solved, continue to automatically execute the next test step.
10. An automated testing system based on retrieval-augmented generation and large language models, which applies the automated testing method based on retrieval-augmented generation and large language models described in any one of claims 1 to 9, characterized in that, Including: Knowledge base construction unit: used to obtain a number of historical test cases from previous periods, vectorize them respectively, and construct a vector database; Retrieval-Augmented Generation Unit: It is used to obtain new test cases described in natural language, and adopt retrieval-augmented generation technology to index relevant content in the vector database as auxiliary information to improve the new test cases and obtain the improved test cases; Automated Testing Unit: It is used to convert the improved test cases into several coherent test steps, start the software under test, perform analysis and decision-making based on the large language model, and call the automated testing tool to automatically execute each test step; Result Analysis Unit: It is used to obtain the test results after all test steps are executed, and use the large language model to analyze the test results to complete the automated testing.
Citation Information
Patent Citations
Test method and device
CN119512941A
Cited By
Intelligent test script generation and semantic maintenance method based on large language model
CN120892346A