Automatic UI function test method and system, electronic equipment and storage medium
Through the testing requirements described by natural language and multimodal recognition technology, combined with AI's understanding ability, generating and executing operation commands, the problem of insufficient testing accuracy in the existing technology is solved, and the efficiency and accuracy of automated UI functional testing is achieved.
Patent Information
- Application Number
- CN202510298306.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-10
AI Technical Summary
The script-based automated testing in the prior art is difficult to understand user intentions, resulting in insufficient test accuracy.
Using automated UI functional testing methods, through the testing requirements described in natural language, combined with multimodal recognition technology and AI's understanding ability, operation commands are generated, and the test results are executed and verified through the operating terminal.
It realizes an automated process from input of test requirements to verification of results, lowers the threshold for expressing test requirements, improves the accuracy and efficiency of tests, and is suitable for non-technical personnel to participate in testing.
Smart Images

Figure CN120123248A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of software development, and in particular to an automated UI function testing method, system, electronic device, and storage medium. Background Art
[0002] In the process of software development, front-end function testing is a key link to ensure product quality. However, there are many problems with script-based automated testing in the prior art. In particular, it is difficult to understand the user's intention because of the ambiguity and diversity of natural language. When users express test requirements, there are large differences in the words and expression structures used, and it is difficult for the prior art to accurately parse them, which easily leads to understanding deviations and thus insufficient test accuracy. Summary of the Invention
[0003] To solve the problem of insufficient test accuracy caused by understanding deviations, this application provides an automated UI function testing method, system, electronic device, and storage medium.
[0004] In a first aspect, this application provides an automated UI function testing method, adopting the following technical solution: An automated UI function testing method includes the following steps: A test requirement input step of inputting test requirements described in natural language; A screen capture and transmission step of capturing a screen image of the user interface and transmitting the captured screen image to the API; A request for multi-modal recognition step. After receiving the screen capture, the API initiates a multi-modal recognition request to the AI to identify the screen content, understand the user's intention, and generate corresponding operation commands; A return and forward operation command step. The AI returns the corresponding operation commands to the API, and the API forwards the operation commands obtained from the AI to the operation terminal; An execute command step. After receiving the command, the operation terminal sequentially executes the operation steps included in the operation command; A loop execution test step of looping the screen capture and transmission step, the request for multi-modal recognition step, the return and forward operation command step, and the execute command step; A test result verification step of testing the real-time test progress of the execute command step, capturing the state changes of interface elements and system logs, and verifying whether it meets the expected test requirements according to the operation results.
[0005] By adopting the above technical solutions, a series of steps of the automated UI function testing method form a complete and efficient testing process. Users input test requirements in natural language, which reduces the threshold for expressing test requirements and facilitates non-technical personnel to participate in testing. The loop screenshot can capture interface changes in real time, providing accurate data for subsequent multimodal recognition. Interact with AI through the API, utilize the multimodal recognition ability of AI to understand the screen content and user intentions, and generate reasonable operation commands. The operation terminal executes the commands and verifies the results to ensure the accuracy and reliability of the test results. At the same time, the entire process realizes the automation from test requirement input to result verification, improving the test efficiency and quality.
[0006] Preferably, in the test requirement input step, the natural language description input covers common front-end function test scenarios, including at least test descriptions of login, registration, search, and form submission.
[0007] By adopting the above technical solutions, the natural language description input covers common front-end function test scenarios, such as login, registration, search, form submission, etc. This enables the test system to comprehensively test various core functions, covering the main business processes of front-end applications, improving the comprehensiveness and integrity of testing, and ensuring the stability and correctness of the system in various common functions.
[0008] Preferably, in the loop execution test step, the operation terminal sets a loop period and periodically captures screen images of the user interface to achieve loop execution of the test.
[0009] By adopting the above technical solutions, a large amount of interface state information can be obtained within the period, and changes in interface elements can be captured more timely and accurately. For some UI interfaces with rapid dynamic changes, such as video playback and animation display scenarios, high-frequency screenshots can ensure that key information is not missed, improving the monitoring ability of the test for dynamic interfaces and ensuring the accuracy of the test.
[0010] Preferably, in the request multimodal recognition step, the AI adopts image recognition technology and natural language understanding technology to analyze the content on the screen and combine the input test requirements to identify the positions and text displays of interface elements.
[0011] By adopting the above technical solutions, the AI uses image recognition technology and natural language understanding technology in the request multimodal recognition step, can comprehensively analyze the screen content and user test requirements, and accurately identify information such as the positions and text displays of interface elements. This enables the system to more accurately understand user intentions, generate operation commands that more meet the actual requirements, improve the accuracy and pertinence of operation commands, and thus enhance the efficiency and reliability of the entire testing process.
[0012] Preferably, it further includes a test result recording and feedback step. The system records the detailed results of each test in a database, including test requirement descriptions, operation procedures, operation times, test results, and reasons for failures.
[0013] By adopting the above technical solution, the test result recording and feedback step records the test details in the database, facilitating testers to view, analyze, and statistically analyze the results, and intuitively understand the overall test situation and function passing rate.
[0014] In a second aspect, the present application provides an automated UI function test system as described in the first aspect above, adopting the following technical solution: An automated UI function test system includes: An LLM module that performs deep learning training using a semantic analysis engine based on the Transformer architecture, including natural language parsing and test instruction generation functions; the natural language parsing function uses a natural language parsing algorithm to identify keywords, operation instructions, and expected result information in the input and establish a semantic model; the test instruction generation function generates corresponding test instruction sequences according to the semantic model, including interface element positioning, operation actions, and data input content. An operation terminal module that realizes browser or mobile application automation and supports cross-platform testing by means of Selenium WebDriver or Appium, and realizes operations such as mouse control and keyboard control, simulating mouse click, drag, and scroll actions, as well as keyboard input and shortcut key operations. A test execution module that deploys a pre-trained CNN model for interface element recognition and generates operation commands by fusing information through an attention mechanism, and has a test process monitoring function to monitor the test progress in real time and capture the state changes of interface elements and system log information.
[0015] By adopting the above technical solution, the LLM module accurately understands natural language test requirements, generates accurate test instruction sequences, reduces the difficulty of requirement conversion, facilitates non-professional test engineers to participate, and improves the accuracy and efficiency of requirement communication. The operation terminal module realizes cross-platform automated testing, simulates various operations, and enhances the test coverage and compatibility. The test execution module accurately identifies interface elements, monitors the progress in real time, captures information, improves the test accuracy and stability, and timely discovers UI function problems. Through the collaborative work of the three modules, it saves software development costs and reduces the burden at multiple levels such as human resources, time, and technology.
[0016] Preferably, the semantic analysis engine of the LLM module based on the Transformer architecture is a BERT or GPT-3 model, and is trained and optimized by collecting a large amount of natural language description data related to front-end function testing. When the operation terminal module conducts web - based testing, it downloads and configures the driver program of the corresponding browser and configures the environment variables. The pre - trained CNN model deployed by the test execution module is a ResNet or Inception model, and the pre - trained model is fine - tuned based on actual test requirements.
[0017] By adopting the above - mentioned technical solutions, the LLM module adopts BERT or GPT - 3 optimized by front - end data to accurately identify requirements, transform instructions, reduce difficulty, improve efficiency, and ensure quality. The operation terminal control module configures the driver and environment variables to solve browser compatibility problems, realizes multi - system and multi - browser testing, and improves reliability and versatility. The test execution module deploys and fine - tunes the ResNet or Inception model to improve the element recognition ability, combines the attention mechanism to make commands accurate, and enhances the test accuracy and reliability.
[0018] Preferably, the natural language parsing function of the LLM module uses part - of - speech tagging and named - entity recognition technologies to assist in identifying keywords, operation instructions, and expected results in the input.
[0019] By adopting the above - mentioned technical solutions, the natural language parsing function of the LLM module uses part - of - speech tagging and named - entity recognition technologies, which can deeply analyze the grammar and semantic structure of natural language, improve the understanding ability of complex test requirement descriptions, enhance the accuracy and reliability of the test instructions generated by the LLM module, and ensure that the test instructions accurately reflect the true needs of users.
[0020] In a third aspect, the present application provides an electronic device, adopting the following technical solution: An electronic device, when the electronic device runs on a server, causes the server to execute the automated UI function test method as described above.
[0021] In a fourth aspect, the present application provides a storage medium, adopting the following technical solution: A storage medium includes instructions, when the instructions run on a server, causing the server to execute the automated UI function test method as described above.
[0022] In summary, the LLM module deeply understands the test requirements expressed by the user in natural language and converts them into instructions for simulating real user operations; the operation terminal control module realizes the simulation of various real operations such as mouse clicks and keyboard inputs, and can faithfully restore whether it is a quick click, a long press or a complex combination operation, effectively improving the intelligence level and efficiency of the test; the test execution module combines the attention mechanism and dynamically adjusts operations according to the interface changes like a real user. When testing complex scenarios such as social APPs, it can simulate diverse behaviors of users such as random swiping, clicking, and jumping, rather than executing according to a fixed script, effectively covering various complex scenarios, greatly improving the test efficiency, and ensuring the stability and accuracy of the software in real usage scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic flowchart of an embodiment of an automated UI function testing method of the present application; Figure 2 It is an interaction schematic diagram of an embodiment of an automated UI function testing method of the present application; Figure 3 It is a schematic structural diagram of an embodiment of an automated UI function testing system of the present application; Figure 4 It is a schematic structural diagram of an electronic device of the present application.
[0024] Reference numerals: 100, LLM module; 200, operation terminal module; 300, test execution module; 401, processor; 402, bus; 403, transceiver; 404, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Referring to the accompanying drawings and specific embodiments, the composition, characteristics, advantages, etc. of the automated UI function testing method, system, electronic device and storage medium according to the present application will be described by way of example below. However, all descriptions should not form any limitation to the present application.
[0026] In addition, for any single technical feature described or implied in the embodiments mentioned in this article, or any single technical feature shown or implied in the respective drawings, the present application still allows any combination or deletion to continue between these technical features (or their equivalents) without any technical obstacles. Therefore, it should be considered that these more embodiments according to the present application are also within the scope of the present article.
[0027] It should also be noted that terms such as "set" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected or indirectly connected through an intermediate medium. Unless otherwise clearly defined, those skilled in the art can understand the specific meaning of the above terms in this application according to the specific situation.
[0028] Referring to Figure 1 and Figure 2 An embodiment of an automated UI function testing method disclosed in an embodiment of the present application specifically includes the following steps: Step S100: Test requirement input step, where the user describes the test requirements in natural language.
[0029] The above natural language description includes various common front-end function test scenarios, such as operation steps like login, registration, search, form submission, etc. By inputting natural language, on the one hand, it solves the cost of requiring professional technical personnel, and on the other hand, it lowers the threshold for test requirement input.
[0030] Step S200: Screenshot transmission step, intercepting and transmitting the screen image of the user interface.
[0031] Take a screenshot of the screen image of the user interface and transmit the intercepted picture to the API. In order to accurately identify the screen image of the user interface, the tester can set a reasonable cycle period on the operation terminal to achieve regular screenshotting of the screen image, for the purpose of providing real-time interface information for subsequent multimodal recognition. In an embodiment of the present application, the cycle period can be set to 5 seconds. In actual operation, the tester can flexibly set the cycle period according to the tested function, without specific limitations. Here, the tester refers to the user of the test operation, and the users in subsequent embodiments of the present application all refer to the actual operating testers.
[0032] It should be understood that the above API refers to the Application Programming Interface, which is mainly used for interaction and communication between different software components, applications, or systems. In an embodiment of the present application, the API is mainly used to transmit images and operation instructions between the operation terminal and the AI; the operation terminal refers to the local program for implementing operation commands.
[0033] Step S300: Request multimodal recognition step, generating corresponding operation commands based on the multimodal recognition result.
[0034] Specifically, when the API receives the screen image, it will immediately initiate a multimodal recognition request to the AI to identify the content of the screen image, understand the user's intention, and generate corresponding operation commands.
[0035] AI will analyze the content on the screen image through image recognition technology and natural language understanding technology, and at the same time combine the test requirements input by the tester to accurately identify the position of interface elements and text display conditions, etc., so as to complete the operation of identifying the screen content and deeply understanding the user's intention, and finally generating corresponding operation commands.
[0036] By sending a multi-modal recognition request to AI and combining the screen image and the natural language test requirements input by the tester, AI can more accurately understand the user's intention and the actual situation of the interface; the screen image here is the image modality, and the natural language test requirements input by the tester are the text modality, and the combination of the two is multi-modal.
[0037] For example, when the tester inputs "Test the login function and verify whether it can successfully log in after entering the correct username and password", and at the same time captures the screen image of the login interface; then AI can use multi-modal recognition technology to identify the positions of the username input box, password input box and login button on the interface, and judge whether the input box is editable and the button is clickable, etc., so as to generate accurate operation commands, simulate the user to perform the login operation, and verify whether the login result meets the expectations. This can greatly improve the intelligence level and accuracy of testing, better simulate the operation behavior of real users, and cover more test scenarios.
[0038] Step S400: Return and forward the operation command step.
[0039] After AI generates the corresponding operation command, AI will return the generated operation command to the API, and the API will then forward the operation command to the operation terminal.
[0040] Step S500: Execute the command step.
[0041] After the operation terminal receives the operation command, it will execute the corresponding actions in sequence based on the operation steps included in the operation command, such as clicking the login button, entering the username and password, etc. This is the execute command step.
[0042] In step S200, after periodically capturing the screen image, continue the above steps S300 - step S500 to implement a loop step, which is to accurately identify the changes of interface elements in order to better complete the test.
[0043] During the above test process, monitor the progress of the execute command step in real time, capture the state changes of user interface elements and system logs, and verify whether it meets the expected test requirements according to the operation results. This is the test result verification step. For example, in the login test, verify whether it successfully jumps to the home page after logging in.
[0044] Meanwhile, during the testing process, the steps of recording and feedback of test results are carried out synchronously. Specifically, the detailed results of each test are recorded in the database, including information such as test requirement description, operation process, operation time, test result (pass or fail), and reason for failure (if any); the database here can be any memory connected to the operation terminal for storing data or the storage space of the terminal device installed on the operation terminal itself. In actual testing, testers can view these test result reports through the visual interface to comprehensively analyze and evaluate the test situation, discover problems in a timely manner and repair them.
[0045] The working principle of the embodiment of this application is as follows: testers can input test requirements described in natural language and, through the intercepted screen images, can recognize and understand the user's intention; then execute the corresponding operation commands and verify the results to ensure the accuracy and reliability of the test results. It realizes simulating the instructions of real user operations and dynamically adjusts the operations according to the changes of the user interface, effectively covering various complex scenarios, greatly improving the test efficiency, and ensuring the stability and accuracy of the software in the real usage scenario. This application integrates multi-modal recognition and AI intention understanding technology, overcomes the problems of high maintenance cost of traditional UI test scripts and low scenario coverage rate, realizes non-intrusive adaptive full-process testing, and in the actual testing process, the test efficiency is increased by more than 40%.
[0046] This application also provides an automated UI function testing system. Figure 3 It is a structural schematic diagram of an embodiment of an automated UI function testing system of this application. As can be seen from Figure 3 it, the above-mentioned automated UI function testing system can include an LLM module 100, an operation terminal module 200, a test execution module 300, etc.
[0047] Specifically, the LLM module 100 is used to understand user requirements and generate test instructions. In the embodiment of this application, the LLM module 100 can adopt the BERT model based on the Transformer architecture as the semantic analysis engine.
[0048] In the BERT model (Bidirectional Encoder Representations from Transformers) based on the Transformer architecture, the Transformer architecture is the foundation of BERT. The Transformer architecture mainly consists of an encoder and a decoder, while BERT only uses the encoder part. The encoder is stacked by multiple identical encoding layers. Each encoding layer has two sub-layers: the multi-head self-attention mechanism and the feed-forward neural network. Among them, the self-attention mechanism is the core innovation point, which enables the model to calculate the correlation weights between positions when processing sequences and capture long-distance dependencies. For example, when processing sentences, it can assign different weights to each word according to the context to understand the semantics. BERT adopts a bidirectional training method, which is its remarkable feature. Most traditional language models are unidirectional, while BERT can consider the context information of words before and after at the same time and perform pre-training through two tasks: the masked language model (MLM) and the next sentence prediction (NSP). The masked language model randomly masks some words in the input sequence and allows the model to predict these words based on the context. The next sentence prediction is used to judge the logical relationship between two sentences, enabling the model to learn the semantic coherence of sentences. BERT has strong semantic understanding ability. Through bidirectional training and large-scale unsupervised learning, it can deeply and accurately understand natural language, capture semantic relationships at the lexical, sentence, and discourse levels, and provide high-quality feature representations for natural language processing tasks. It also has generality and transferability. As a pre-trained model, it can be pre-trained on large-scale unlabeled text data and then applied to different downstream tasks through fine-tuning, reducing the training cost of each specific task. In addition, the self-attention mechanism of the Transformer architecture gives BERT obvious advantages in processing long texts, being able to effectively capture long-distance dependencies and avoid the problem of gradient disappearance or explosion when traditional recurrent neural networks process long sequences. In the embodiments of this application, BERT can be used to understand the test requirements described in the user's natural language. For example, when the user inputs the test requirements related to the search function, it can accurately identify information such as keywords, operation instructions, and expected results, and then generate a test instruction sequence to guide the test operation, improving the intelligence level and accuracy of the test.
[0049] To enable the model to accurately understand natural language descriptions related to front-end functional testing, a large amount of natural language data covering various common front-end functional testing scenarios was collected for training and optimization. These scenarios include, but are not limited to, login, registration, search, and form submission, etc. In natural language parsing, the LLM module uses advanced technologies such as part-of-speech tagging and named entity recognition to assist in identifying key information in the user input. For example, when the user inputs "Test the login function, log in using the username 'testuser' and password '123456', and verify if it successfully jumps to the home page", the system can accurately identify the keyword "login function", the operation instruction "log in using a specific username and password", and the expected result "successfully jump to the home page" with the help of these technologies, and then establish a corresponding semantic model. Based on this semantic model, the test instruction generation function will generate a detailed test instruction sequence, which includes interface element positioning information, such as the positions of the login button, username input box, and password input box; operation actions, such as click, input, etc.; and specific data input content, namely the username and password.
[0050] The main function of the operation terminal module 200 is to achieve automated operations on browsers or mobile applications and support cross-platform testing. Here, Selenium WebDriver is used to achieve this goal. When conducting web page testing, the operation terminal module 200 can download and configure the driver program of the Chrome browser and configure the corresponding environment variables to ensure that the system can communicate with the browser stably and accurately. Through the operation terminal module 200, the system can simulate a variety of user operations, including mouse click, drag, scroll actions, and keyboard input, shortcut key operations, etc. For example, when simulating a login operation, the system can precisely control the mouse to click on the username input box and then input the pre-set username through the keyboard, completely restoring the operation process of a real user.
[0051] The test execution module 300 is used to identify interface elements and generate corresponding operation commands, while monitoring the test process in real time. The test execution module 300 deploys a pre-trained ResNet model for interface element identification and fine-tunes this pre-trained model according to actual test requirements to improve its adaptability to specific test scenarios. When requesting multimodal recognition, the test execution module 300 uses image recognition technology and natural language understanding technology, combines with the test requirements input by the user, and deeply analyzes the content on the screen. In this way, the position of the interface elements and the text display situation can be accurately identified. For example, the test execution module 300 can accurately identify the position of the login button and the text "Login" displayed on the button. At the same time, by fusing information through the attention mechanism, the test execution module 300 can generate operation commands that conform to the actual situation. In addition, the test execution module 300 also has a powerful test process monitoring function, which can monitor the test progress in real time, capture the state changes of interface elements and system log information, and provide comprehensive and accurate data support for subsequent test result verification.
[0052] In the embodiments of the present application, an electronic device is provided, such as Figure 4 shown, Figure 4 The electronic device shown includes: a processor 401 and a memory 404. Among them, the processor 401 and the memory 404 are connected, such as through a bus 402. Optionally, the oil pressure detection device may further include a transceiver 403. It should be noted that in actual applications, the transceiver 403 is not limited to one, and the structure of the oil pressure detection device does not constitute a limitation to the embodiments of the present application.
[0053] The processor 401 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure of the present application. The processor 401 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0054] The bus 402 may include a path for transmitting information among the above components. The bus 402 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 402 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a thick line is used in Figure 4 , but it does not mean that there is only one bus or one type of bus.
[0055] The memory 404 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0056] The memory 404 is used to store the application program code for executing the solution of this application, and is controlled by the processor 401 for execution. The processor 401 is used to execute the application program code stored in the memory 404 to implement the content shown in the foregoing method embodiments.
[0057] Among them, Figure 4 the shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of this application.
[0058] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.
[0059] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0060] The above are all preferred embodiments of the present application, and do not limit the protection scope of the present application accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.
Claims
1. An automated UI function testing method, characterized in that: The following steps are involved: A test requirement input step is to input the test requirements described in natural language; A screenshot transmission step, capturing a screen image of the user interface, and transmitting the captured screen image to the API; Request multimodal recognition step. After receiving the screenshot, the API initiates a multimodal recognition request to the AI to identify the screen content, understand the user's intention, and generate corresponding operation commands; Return and forward the operation command step, AI returns the corresponding operation command to API, and API forwards the operation command obtained from AI to the operation terminal; Execute command steps. After receiving the command, the operation terminal executes the operation steps contained in the operation command in sequence; Loop through the test steps, loop through the screenshot transmission steps, loop through the multi-modal recognition request steps, loop through the operation command forwarding steps, and loop through the command execution steps; Test result verification steps, real-time test progress of test execution command steps, capture status changes of interface elements and system logs, and verify whether the operation results meet the expected test requirements.
2. The automated UI function testing method according to claim 1, characterized in that: In the test requirement input step, the input natural language description covers common front-end functional test scenarios, including at least test descriptions of login, registration, search, and form submission.
3. The automated UI function testing method according to claim 2, characterized in that: The cyclic execution test steps are performed by operating the terminal to set the cycle period and regularly capture the screen image of the user interface to implement the cyclic execution test.
4. The automated UI function testing method according to claim 3, characterized in that: In the step of requesting multimodal recognition, AI uses image recognition technology and natural language understanding technology to analyze the content on the screen image and the input test requirements to identify the location of interface elements and text display.
5. The automated UI function testing method according to claim 1, characterized in that: It also includes test result recording and feedback steps. The system records the detailed results of each test in the database, including test requirement description, operation process, operation time, test results, and reasons for failure.
6. An automated UI function testing system, characterized in that: include: The LLM module uses a semantic analysis engine based on the Transformer architecture for deep learning training, including natural language parsing and test instruction generation functions; the natural language parsing function uses a natural language parsing algorithm to identify keywords, operation instructions, and expected result information in the input, and establish a semantic model; the test instruction generation function generates the corresponding test instruction sequence according to the semantic model, including interface element positioning, operation actions, and data input content; The operation terminal module uses Selenium WebDriver or Appium to realize browser or mobile application automation, support cross-platform testing, realize mouse control and keyboard control operations, simulate mouse clicks, drags, scrolling actions, as well as keyboard input and shortcut key operations; The test execution module deploys a pre-trained CNN model to identify interface elements and generates operation commands by fusing information through the attention mechanism. It has the function of monitoring the test process, which can monitor the test progress in real time and capture the status changes of interface elements and system log information.
7. The automated UI function testing system according to claim 6, characterized in that: The semantic analysis engine of the LLM module based on the Transformer architecture is a BERT or GPT-3 model, which is trained and optimized by collecting a large amount of natural language description data related to front-end functional testing; When the operation terminal module performs web page testing, it downloads and configures the driver of the corresponding browser and configures the environment variables; The pre-trained CNN model deployed by the test execution module is a ResNet or Inception model, and the pre-trained model is fine-tuned based on actual test requirements.
8. The automated UI function testing system according to claim 6, characterized in that: The natural language parsing function of the LLM module uses part-of-speech tagging and named entity recognition technology to assist in identifying keywords, operation instructions and expected results in the input.
9. An electronic device, characterized in that: When the electronic device runs on the server, the server executes the method as described in any one of claims 1-5.
10. A storage medium comprising instructions, characterized in that: When the instructions are executed on a server, the server is caused to execute the method according to any one of claims 1 to 5.
Citation Information
Cited By
Interface operation instruction generation method, electronic equipment, storage medium and program product
CN120704792A
Automatic black box testing method and related device
CN120849303A
Smart phone automatic operation method and system based on sandbox and large language model
CN120850278A
Equipment control method and device, equipment, medium and product
CN120872278A