UI automatic operation platform construction method and system based on OCR recognition assistance
By adopting an OCR recognition-assisted platform in UI automation testing, combining lightweight object detection and text recognition models, the problem of UI interface structure recognition is solved, and the accuracy and accuracy of the test is improved.
Patent Information
- Application Number
- CN202510140339.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-07-08
AI Technical Summary
The existing OCR technology cannot accurately identify the arrangement structure of the UI interface in UI automation testing, resulting in inaccurate test results.
Using an UI automation operation platform based on OCR recognition assistance, we use tests in the target server, collect graphical interface images, and use OCR image recognition technology and lightweight object detection model to divide the pictures, identify key areas and match the corresponding text recognition model to obtain test results.
It improves the accuracy of UI automation testing, judges image similarity through structural similarity index, ensures character recognition when the similarity is higher than the threshold, reduces misrecognition, and improves the accuracy of the test.
Smart Images

Figure CN120276977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UI testing, and in particular to a method for constructing a UI automated operation platform assisted by OCR recognition, and also relates to a system for constructing a UI automated operation platform assisted by OCR recognition applying this method. Background Art
[0002] UI testing, short for User Interface testing, tests whether the layout of functional modules of the user interface is reasonable, whether the overall style is consistent, whether the placement positions of each control conform to the usage habits of customers. In addition, it also tests the convenience of interface operations, the simplicity and understandability of navigation, the usability of page elements, whether the text in the interface is correct, whether the naming is unified, whether the page is beautiful, and whether the combination of text and pictures is perfect, etc.
[0003] Automated testing technologies can be divided into intrusive and non-intrusive according to their mechanisms. The intrusive testing technology modifies the internal code of the test target software in a certain way or controls its running environment; while the non-intrusive testing technology monitors the user interface of the test target software, does not modify the internal structure or code of the software, but simulates testers to complete the test process.
[0004] OCR technology refers to the process in which an electronic device examines the characters printed on a picture, determines their shapes by detecting dark and bright patterns, and then translates the shapes into computer text using character recognition methods. Among them, OCR technology can be applied in the UI automated testing of software. However, OCR technology can only recognize the text on the UI interface of the software, and cannot recognize the arrangement structure of various types of regions on the UI interface, thus unable to accurately obtain the test results of UI automated testing. Summary of the Invention
[0005] In order to accurately judge the test results of UI automated testing, the present invention provides a method and a system for constructing a UI automated operation platform assisted by OCR recognition.
[0006] The present invention is implemented by the following technical solutions: A method for constructing a UI automated operation platform assisted by OCR recognition, which includes the following steps: Based on the test script of the target test software, test the target test software in the target server; wherein, the target server includes a UI automated testing environment; collect the to-be-processed pictures with information of interest in the graphical interface of the target test software during the test; if the structural similarity index between the to-be-processed pictures and the template pictures is greater than a preset threshold, then recognize the characters in the to-be-processed pictures based on OCR image recognition technology; compare the characters in the to-be-processed pictures with the expected values in the template pictures to obtain the test results.
[0007] As a further improvement of the above solution, identifying the characters in the to-be-processed picture based on the OCR image recognition technology includes the following steps: using a lightweight object detection model to divide the region of the to-be-processed picture to obtain a number of key regions, where the key regions include a text region, a title region, a table region, and a graphic region; matching a character recognition model corresponding to the type of the key region based on the type of the key region matched; and performing text recognition on the key region based on the character recognition model.
[0008] As a further improvement of the above solution, the character recognition model includes a text structure recognition model and a table structure recognition model.
[0009] As a further improvement of the above solution, before the step of, if the structural similarity index between the to-be-processed picture and the template picture is greater than a preset threshold, identifying the characters in the to-be-processed picture based on the OCR image recognition technology, includes the following steps: obtaining a brightness contrast value between the to-be-processed picture and the template picture based on the first pixel average value of the to-be-processed picture and the second pixel average value of the template picture; obtaining a variance contrast value between the to-be-processed picture and the template picture based on the first pixel variance of the to-be-processed picture and the second pixel variance of the template picture; obtaining a structural contrast ratio between the to-be-processed picture and the template picture based on the pixel covariance between the to-be-processed picture and the template picture; and obtaining the structural similarity index between the to-be-processed picture and the template picture based on the brightness contrast value, the variance contrast value, and the structural contrast ratio.
[0010] As a further improvement of the above solution, the step of real-time monitoring and collecting the to-be-processed pictures with test interesting information in the graphical interface of the target test software during the test process includes the following steps: automatically obtaining a graphical interface corresponding to the positioning interface site information based on the positioning interface site information in the test script; and obtaining the to-be-processed pictures with test interesting information in the graphical interface of the target test software based on the graphical interface.
[0011] As a further improvement of the above solution, identifying the characters in the to-be-processed picture based on the OCR image recognition technology includes the following steps: positioning a graphical feature position picture in the to-be-processed picture based on the OCR image recognition technology; performing binarization processing on the graphical feature position picture to obtain a binary image including a foreground color and a background color; and matching and identifying the characters in the binary image based on the binary image and a character feature library.
[0012] As a further improvement of the above solution, the step of matching and recognizing the characters in the binary image based on the binary image and the character feature library includes the following steps: performing row and column character segmentation on the binary image to obtain individual characters; extracting features of the individual characters, and matching the extracted features with the character feature library, and if the matching is successful, the characters are recognized.
[0013] As a further improvement of the above solution, the types of the test scripts include: python, selenium or pytest.
[0014] As a further improvement of the above solution, before testing the target test software in the target server based on the test script of the target test software, the following steps are included:
[0015] Allocating the test script to the target server in response to an execution instruction; or
[0016] Allocating the test script to the target server at a set time point.
[0017] The present invention further provides a system for constructing a UI automation operation platform based on OCR recognition assistance, which applies any of the above methods for constructing a UI automation operation platform based on OCR recognition assistance, and includes:
[0018] A test start module, which is used to test the target test software in the target server based on the test script of the target test software, wherein the target server includes a UI automation test environment;
[0019] A picture acquisition module, which is used to acquire a to-be-processed picture with information of interest in the graphical interface of the target test software during the test;
[0020] A character recognition module, which is used to recognize the characters in the to-be-processed picture based on the OCR image recognition technology if the structural similarity index between the to-be-processed picture and the template picture is greater than a preset threshold;
[0021] A test result acquisition module, which is used to compare the characters in the to-be-processed picture with the expected values in the template picture to obtain a test result.
[0022] Compared with the existing methods for constructing a UI automation operation platform based on OCR recognition assistance, the method and system for constructing a UI automation operation platform based on OCR recognition assistance of the present invention have the following beneficial effects:
[0023] The target server tests the target test software in the target server based on the test script of the target test software; wherein, a UI automation test environment is included in the target server; pictures to be processed with information of interest in the graphical interface of the target test software during the test are collected; if the structural similarity index between the pictures to be processed and the template pictures is greater than a preset threshold, characters in the pictures to be processed are recognized based on the OCR image recognition technology; the characters in the pictures to be processed are compared with the expected values in the template pictures to obtain a test result; through the above method, if the structural similarity index between the pictures to be processed and the template pictures is greater than the preset threshold, characters in the pictures to be processed are recognized based on the OCR image recognition technology. Before performing character recognition on the pictures to be processed, obtaining the structural similarity index between the pictures to be processed and the template pictures can improve the accuracy of UI automation testing.
[0024] The beneficial effects of the UI automation operation platform construction system assisted by OCR recognition are the same as those of the UI automation operation platform construction method assisted by OCR recognition described above, and will not be elaborated here. Brief Description of the Drawings
[0025] Figure 1 It is a flowchart of the method for constructing a UI automation operation platform assisted by OCR recognition according to an embodiment of the present invention.
[0026] Figure 2 It is a flowchart of the character recognition step according to an embodiment of the present invention.
[0027] Figure 3 It is a flowchart of the step for obtaining the structural similarity index according to an embodiment of the present invention.
[0028] Figure 4 It is a flowchart of the character recognition step according to an embodiment of an application of the present invention.
[0029] Figure 5 It is a flowchart of the character recognition step according to another embodiment of an application of the present invention. Detailed Description of the Embodiment
[0030] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] Please refer to Figure 1 , in an embodiment of the present application, a method for constructing a UI automation operation platform assisted by OCR recognition is proposed, and the method includes the following steps:
[0032] S101. Test the target test software on the target server based on the test script of the target test software;
[0033] Specifically, the user locally writes an automated test script, which includes the maintenance of environment variables, test data, test cases, and an automated execution script. After writing, in response to an execution instruction, the test script is assigned to the target server; or the test script is assigned to the target server at a set time point. The target server receives the test script and parses it. Among them, the types of test scripts include: python, selenium, or pytest.
[0034] S102. Collect the to-be-processed pictures with interesting information in the graphical interface of the target test software during the test process;
[0035] Based on the positioning interface site information in the test script, the target server automatically obtains the graphical interface corresponding to the positioning interface site information; based on the graphical interface, it obtains the to-be-processed pictures with test interesting information in the graphical interface of the target test software.
[0036] S103. If the structural similarity index between the to-be-processed picture and the template picture is greater than the preset threshold, then identify the characters in the to-be-processed picture based on the OCR image recognition technology;
[0037] Among them, on the basis that there are no major logical errors in the overall test process, there is a preset threshold for the matching degree between the UI process pictures of the target test software and the template pictures. Because different types of screenshots and different types of regions have different matching degrees, it needs to be set flexibly.
[0038] Specifically, if the structural similarity index between the to-be-processed picture and the template picture is greater than the preset threshold, based on the OCR image recognition technology, locate the picture of the graphical feature position in the to-be-processed picture; perform binarization processing on the picture of the graphical feature position to obtain a binary image containing foreground color and background color; based on the binary image and the character feature library, match and identify the characters in the binary image;
[0039] If the structural similarity index between the to-be-processed picture and the template picture is less than the preset threshold, then directly return the test result of the target test software, and no longer identify the characters in the to-be-processed picture using the OCR image recognition technology.
[0040] S104. Compare the characters in the to-be-processed picture with the expected values in the template picture to obtain the test result.
[0041] The target server tests the target test software in the target server based on the test script of the target test software; wherein, the target server includes a UI automated test environment; collect the to-be-processed pictures with information of interest in the graphical interface of the target test software during the test; if the structural similarity index between the to-be-processed pictures and the template pictures is greater than the preset threshold, then identify the characters in the to-be-processed pictures based on the OCR image recognition technology; compare the characters in the to-be-processed pictures with the expected values in the template pictures to obtain the test results; through the above method, if the structural similarity index between the to-be-processed pictures and the template pictures is greater than the preset threshold, then identify the characters in the to-be-processed pictures based on the OCR image recognition technology. Before performing character recognition on the to-be-processed pictures, obtaining the structural similarity index between the to-be-processed pictures and the template pictures can improve the accuracy of UI automated testing.
[0042] In an application embodiment, please refer to Figures 2 - 3 , S103 further includes the following steps:
[0043] S201. Obtain the brightness contrast value between the to-be-processed picture and the template picture based on the first pixel average value of the to-be-processed picture and the second pixel average value of the template picture;
[0044] It should be noted that the brightness contrast value is obtained by the following formula:
[0045]
[0046] where μ x , μ y respectively represent the pixel mean values of the to-be-processed picture and the template picture, c1 = (k1L) 2 , c2 = (k2L) 2 , L represents the range of image pixel values.
[0047] S202. Obtain the variance contrast value between the to-be-processed picture and the template picture based on the first pixel variance of the to-be-processed picture and the second pixel variance of the template picture;
[0048] It should be noted that the variance contrast value is obtained by the following formula:
[0049]
[0050] where σ x , σ y respectively represent the pixel variance values of the to-be-processed picture and the template picture, c2 = (k2L) 2 , L represents the range of image pixel values.
[0051] S203. Obtain the structural contrast between the picture to be processed and the template picture based on the pixel covariance between the picture to be processed and the template picture;
[0052]
[0053] where σ x , σ y respectively represent the pixel variance values of the picture to be processed and the template picture, and σ xy represents the pixel covariance between the picture to be processed and the template picture, and c3 = c2 / 2.
[0054] S204. Obtain the structural similarity index between the picture to be processed and the template picture based on the brightness contrast value, the variance contrast value, and the structural contrast.
[0055] The structural similarity index is obtained by the following formula:
[0056]
[0057] In the embodiments of the present application, before using the OCR image recognition technology to recognize the characters in the picture to be processed, the structural similarity index is used to measure the similarity between the picture to be processed and the template picture. After determining the high similarity between the picture to be processed and the template picture, the OCR image recognition technology is used to recognize the characters in the picture to be processed, which can improve the accuracy of UI automated testing.
[0058] In an embodiment of the application, please refer to Figures 4 - 5 , S103 includes the following steps:
[0059] S301. Use a lightweight object detection model to perform region division on the picture to be processed, and obtain a number of key regions, where the key regions include text regions, title regions, table regions, and graphic regions;
[0060] It should be noted that the lightweight object detection model is the PP-PicoDet model. This model uses a lightweight backbone network and optimizes the prediction size for the screenshot analysis scenario. It has the advantages of a small number of model parameters, fast inference speed, and high accuracy; this model is trained on the PubLayNet dataset, and the trained PP-PicoDet model has good generalization.
[0061] S302. Based on the types of the key regions, match the text recognition models corresponding to the types of the key regions;
[0062] where the text recognition models include a text structure recognition model and a table structure recognition model;
[0063] Specifically, the key areas include text areas, title areas, table areas, graphic areas, etc. Among them, the text areas, title areas, and graphic areas can match the text structure recognition model; the table areas can match the table structure recognition model.
[0064] S303. Perform text recognition on the key areas based on the character recognition model.
[0065] Specifically, the process of using the text structure recognition model to perform character recognition on the key areas includes: performing row and column character segmentation on the binary image of the key area to obtain individual characters; extracting features from the individual characters, and matching the extracted features with the character feature library. If the match is successful, the character is recognized, and a txt file containing the character is output.
[0066] In an application embodiment, in order to accurately restore the table structure information, the TableMaster table recognition model is used as the table structure recognition model, and the TableMaster table recognition model is used to perform character recognition on the key areas to obtain an excel file.
[0067] In the embodiment of the present application, according to the types of the key areas, the character recognition models corresponding to the types of the key areas are matched, which greatly improves the accuracy of the OCR image recognition technology in recognizing the characters in the to-be-processed picture.
[0068] In an application embodiment, a system for constructing a UI automation operation platform based on OCR recognition assistance is provided. The system includes:
[0069] A test start module, which is used to test the target test software in the target server based on the test script of the target test software, where the target server includes a UI automation test environment;
[0070] A picture acquisition module, which is used to acquire the to-be-processed pictures with the information of interest in the graphical interface of the target test software during the test process;
[0071] A character recognition module, which is used to, if the structural similarity index between the to-be-processed picture and the template picture is greater than a preset threshold, recognize the characters in the to-be-processed picture based on the OCR image recognition technology;
[0072] A test result acquisition module, which is used to compare the characters in the to-be-processed picture with the expected values in the template picture to obtain the test results.
[0073] Compared with the prior art, the advantages of the device for constructing a UI automation operation platform based on OCR recognition assistance are the same as those of the method for constructing a UI automation operation platform based on OCR recognition assistance, and will not be elaborated here.
[0074] In an application embodiment, a computer terminal is provided, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The steps of a method for constructing a UI automation operation platform assisted by OCR recognition are implemented when the processor executes the program. When applied, this method can be applied in the form of software, such as designed as an independently running program and installed on the computer terminal, which can be a computer, a smart phone, a control system, and other Internet of Things devices, etc. This method can also be designed as an embedded running program and installed on the computer terminal, such as installed on a single-chip microcomputer.
[0075] In an application embodiment, this embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by the processor, the steps of a method for constructing a UI automation operation platform assisted by OCR recognition are implemented. When applied, this method can be applied in the form of software, such as designed as an independently running program on the computer-readable storage medium. The computer-readable storage medium can be a USB flash drive, designed as a USB key, and designed as a program that triggers the start of the entire method through the USB flash drive.
[0076] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a UI automation operation platform assisted by OCR recognition, characterized in that The method includes the following steps: Based on the test script of the target test software, test the target test software in the target server; wherein, the target server includes a UI automated test environment; Collect the to-be-processed pictures with information of interest in the graphical interface of the target test software during the test; If the structural similarity index between the to-be-processed picture and the template picture is greater than a preset threshold, then identify the characters in the to-be-processed picture based on the OCR image recognition technology; Compare the characters in the to-be-processed picture with the expected values in the template picture to obtain the test result.
2. The method for constructing a UI automated operation platform based on OCR recognition assistance according to claim 1, wherein The identifying the characters in the to-be-processed picture based on the OCR image recognition technology includes the following steps: Use a lightweight object detection model to divide the to-be-processed picture into regions, and obtain a number of key regions, wherein the key regions include text regions, title regions, table regions, and graphic regions; Based on the types of the key regions, match the text recognition models corresponding to the types of the key regions; Based on the text recognition models, perform text recognition on the key regions.
3. The method for constructing a UI automation operation platform based on OCR recognition assistance according to claim 2, wherein The text recognition models include a text structure recognition model and a table structure recognition model.
4. The method for constructing a UI automation operation platform based on OCR recognition assistance according to claim 1, wherein, Before the step of if the structural similarity index between the to-be-processed picture and the template picture is greater than a preset threshold, then identify the characters in the to-be-processed picture based on the OCR image recognition technology, the following steps are included: Based on the first pixel average value of the to-be-processed picture and the second pixel average value of the template picture, obtain the brightness contrast value between the to-be-processed picture and the template picture; Based on the first pixel variance of the to-be-processed picture and the second pixel variance of the template picture, obtain the variance contrast value between the to-be-processed picture and the template picture; Based on the pixel covariance between the to-be-processed picture and the template picture, obtain the structural contrast between the to-be-processed picture and the template picture; Based on the brightness contrast value, the variance contrast value, and the structural contrast, obtain the structural similarity index between the to-be-processed picture and the template picture.
5. The method for constructing a UI automated operation platform assisted by OCR recognition as claimed in claim 1, wherein, The real-time monitoring and collecting the to-be-processed pictures with information of interest in the graphical interface of the target test software during the test includes the following steps: Based on the positioning interface site information in the test script, automatically obtain the graphical interface corresponding to the positioning interface site information; Based on the graphical interface, obtain the to-be-processed pictures with information of interest in the graphical interface of the target test software.
6. The method for constructing a UI automation operation platform based on OCR recognition assistance according to claim 1, wherein, The types of the test scripts include: python, selenium, or pytest.
7. The method for constructing a UI automation operation platform based on OCR recognition assistance according to claim 1, wherein, Before the step of based on the test script of the target test software, test the target test software in the target server, the following steps are included: Respond to the execution instruction to allocate the test script to the target server; or Allocate the test script to the target server at a set time point.
8. A system for constructing a UI automated operation platform assisted by OCR recognition, characterized in that, The system includes: A test startup module, which is used to test the target test software in the target server based on the test script of the target test software, wherein the target server includes a UI automated test environment; An image acquisition module, which is used to acquire the to-be-processed images with information of interest in the graphical interface of the target test software during the test; A character recognition module, which is used to recognize the characters in the to-be-processed images based on the OCR image recognition technology if the structural similarity index between the to-be-processed images and the template images is greater than a preset threshold; A test result acquisition module, which is used to compare the characters in the to-be-processed images with the expected values in the template images to obtain test results.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
User interface test method, system, equipment and medium
CN121255661A